How to scrape Capterra at scale in 2026 with proxies that work
Capterra is one of the better public sources for b2b software data. Category pages list hundreds of products, each product page carries pricing hints, feature lists and alternatives, and the review pages hold the text that sales teams, analysts and SaaS founders actually want. The problem is that the site sits behind bot protection, so a script that works for 50 pages from your home connection tends to fall over at 5,000. The failures are rarely clean: empty pages, or challenge pages that return HTTP 200.
This tutorial is for operators who need repeatable Capterra extraction: market researchers, lead-gen builders, competitive intel people, and anyone feeding a database of software vendors. I run my own scraping from Singapore, so some of the latency and geo notes come from there. You should be comfortable with Python and a terminal.
By the end you will have a working pipeline: a proxy pool, a Playwright-based fetcher with sane pacing, a parser that writes to SQLite, and a retry loop that survives blocks. I will also cover what changes when you go from a few hundred pages to hundreds of thousands.
what you need
- python 3.11 or newer, with pip
- playwright for python (
pip install playwrightthenplaywright install chromium), docs at playwright.dev - a residential proxy plan with country targeting and sticky sessions. billing is usually per GB, so check the vendor’s current pricing page before you commit. mobile proxies also work and tend to survive longer, but cost more per GB
- optional: a small pool of datacenter proxies for the cheap, low-risk requests such as sitemap fetches
- sqlite (ships with python) or postgres if you plan to go past a few million rows
- a VPS or home server that can run headless chromium. 2 vCPU and 4 GB RAM handles 3 to 5 parallel browser contexts without drama
- a written note of what you are allowed to collect. read the site’s terms and its robots.txt first. this is not legal advice, and if the data is going into a commercial product you should talk to a lawyer
- a budget for failure. expect some bandwidth lost to blocked requests while you tune
On proxy type: for Capterra I start with residential, sticky for the length of one product crawl, and rotate between products. If you want the longer comparison, I wrote it up in residential vs mobile proxies for scraping.
step by step
1. set up the project and check your baseline
Create a folder, a virtualenv and install dependencies.
mkdir capterra-scrape && cd capterra-scrape
python -m venv .venv
source .venv/bin/activate # on windows: .venv\Scripts\activate
pip install playwright selectolax tenacity
playwright install chromium
Before adding any proxy, load one category page with a plain request and see what you get. This is your baseline.
curl -s -o /dev/null -w "%{http_code}\n" https://www.capterra.com/project-management-software/
Expected output: a status code. A 200 does not prove success, since challenge pages can also return 200. A 403 or 429 tells you the datacenter or home IP is already flagged.
If it breaks: open the same URL in a normal browser from the same machine. If the browser works and curl does not, you are being fingerprinted on headers and TLS, which is why we use a real browser in step 3.
2. wire up the proxy and verify the exit
Most residential vendors give you a host, port, username and password, with session control encoded in the username. The exact syntax differs per vendor, so read yours. Store credentials in environment variables, never in the script.
export PROXY_HOST="gate.example-vendor.com:7000"
export PROXY_USER="user-country-us-session-abc123"
export PROXY_PASS="your-password"
curl -x "http://$PROXY_USER:$PROXY_PASS@$PROXY_HOST" https://api.ipify.org
Expected output: an IP address that is not yours. Run it twice with the same session id to confirm it is sticky, then change the session id and confirm it changes.
If it breaks: a 407 means bad credentials. A timeout usually means the port or protocol is wrong (http vs socks5). If the IP geolocates somewhere you did not ask for, fix the country flag before going further. Capterra shows different content and currency by region, so mixed geos will make your data inconsistent.
3. build a browser fetcher
Plain HTTP clients get flagged quickly on sites with bot protection. A real chromium via Playwright carries a believable TLS and JS fingerprint. Proxy config is per browser launch or per context, as described in the Playwright network docs.
import os
from playwright.sync_api import sync_playwright
def fetch(url, session_id):
proxy = {
"server": f"http://{os.environ['PROXY_HOST']}",
"username": f"user-country-us-session-{session_id}",
"password": os.environ["PROXY_PASS"],
}
with sync_playwright() as p:
browser = p.chromium.launch(headless=True, proxy=proxy)
ctx = browser.new_context(locale="en-US", viewport={"width": 1366, "height": 800})
page = ctx.new_page()
resp = page.goto(url, wait_until="domcontentloaded", timeout=45000)
page.wait_for_timeout(1500)
html = page.content()
status = resp.status if resp else 0
browser.close()
return status, html
Expected output: a status and an html string several hundred KB long for a category page.
If it breaks: print the first 300 characters of the html. If you see a challenge or “verify you are human” text, your exit IP is burned or the fingerprint is off. Try a fresh session id first. If that fails across several IPs, slow down (step 5) before touching anything else.
4. discover urls from category pages
Start from category pages, such as /project-management-software/, and collect product links. Product pages follow a pattern like /p/<id>/<name>/, and reviews hang off /p/<id>/<name>/reviews/. Check the live pattern before hard-coding it.
from selectolax.parser import HTMLParser
import re
def product_links(html):
tree = HTMLParser(html)
out = set()
for a in tree.css("a[href]"):
href = a.attributes.get("href", "")
m = re.match(r"^(/p/\d+/[^/]+)/?", href)
if m:
out.add("https://www.capterra.com" + m.group(1) + "/")
return sorted(out)
Expected output: a list of product urls, typically a few dozen per category page. Write them to a queue table with a status column so you can resume after a crash.
If it breaks: zero links with a 200 status means you got a challenge or a lazy-loaded shell. Add page.wait_for_selector("a[href^='/p/']") to the fetcher. If links appear in the browser but not in your html, scroll the page once with page.mouse.wheel(0, 4000) and wait a second.
5. pace requests like a person would
This is the step that decides whether you scale. Parallelism is cheap, but speed per exit IP is what gets you blocked. My starting rules:
- one request per 6 to 12 seconds per exit IP, with random jitter
- one product crawl (overview plus review pages) per sticky session, then rotate
- no more than 3 to 5 concurrent browser contexts per machine at the start
- a hard stop on any IP after two non-200 or challenge responses
import random, time
def polite_sleep():
time.sleep(random.uniform(6, 12))
Expected output: slower runs, with far fewer retries. Judge success by clean pages per GB, not pages per minute.
If it breaks: if block rates climb as you add workers, you are over the per-IP limit or sharing exits between workers. Give each worker its own session id.
6. parse the data you actually need
Decide your schema before you crawl. For product pages I keep: product id, name, vendor, category, rating, review count, starting price text, and the url. For reviews: review id, date, overall rating, title, pros, cons, reviewer role and company size bucket when shown. Capterra’s markup shifts, so prefer stable attributes and structured data in the page when present, and always keep the raw html for a few days so you can re-parse.
import sqlite3
db = sqlite3.connect("capterra.db")
db.execute("""CREATE TABLE IF NOT EXISTS products(
id INTEGER PRIMARY KEY, name TEXT, rating REAL, reviews INTEGER,
price_text TEXT, url TEXT UNIQUE, fetched_at TEXT)""")
db.execute("""CREATE TABLE IF NOT EXISTS queue(
url TEXT PRIMARY KEY, kind TEXT, status TEXT DEFAULT 'new', tries INTEGER DEFAULT 0)""")
db.commit()
Expected output: a populated products table after your first 100 pages, with no null names.
If it breaks: nulls in a column across many rows mean the selector changed or the page was a challenge in disguise. Add a sanity check that rejects any page with no product name and requeues it.
7. add retries and block detection
A 200 status is not proof of a good page. Check the content. If the html is shorter than a threshold, lacks the product title, or contains known challenge text, treat it as a block.
from tenacity import retry, stop_after_attempt, wait_exponential
class Blocked(Exception): pass
def looks_blocked(status, html):
return status in (403, 429) or len(html) < 20000 or "<h1" not in html
@retry(stop=stop_after_attempt(3), wait=wait_exponential(min=10, max=90))
def safe_fetch(url):
session = str(random.randint(10**6, 10**7))
status, html = fetch(url, session)
if looks_blocked(status, html):
raise Blocked(url)
return html
Expected output: blocked pages retry on a new session and succeed most of the time. Log each block with the proxy session id so you can spot a bad subnet.
If it breaks: if all three attempts fail for a url, mark it failed in the queue and move on. Retry it the next day.
8. run it and watch the numbers
Wire the queue, fetcher and parser into a loop, then run a pilot of 200 products. Track four numbers: success rate, average GB per 1,000 pages, blocks per 100 requests, and parse failures. Residential bandwidth is the main cost, so block images, fonts and media in Playwright with request routing to cut the bytes per page substantially.
page.route("**/*", lambda r: r.abort()
if r.request.resource_type in ("image", "font", "media") else r.continue_())
Expected output: a pilot that finishes with a success rate above roughly 90 percent. If you are far below that, do not scale yet.
If it breaks: if blocking only starts after blocking images, some protection scripts treat missing resources as a signal. Allow images on the first request of each session and block them afterwards.
common pitfalls
- trusting status codes: challenge pages often return 200. validate content, not just status
- rotating the IP on every request: a real visitor browses a few pages from one IP. rotate per product, not per page
- mixing geos: if your exit countries vary, prices and review ordering vary too, and your dataset becomes noisy
- skipping the pilot: running 50,000 urls before you know your block rate is the fastest way to waste a month of proxy budget
- ignoring robots.txt and terms: the Robots Exclusion Protocol (RFC 9309) is advisory, but it is the site’s stated preference. decide your position deliberately and document it
scaling this
From 10x to 100x (hundreds to tens of thousands of pages): the bottleneck is IP quality and pacing, not code. Move the queue to a proper database, add per-worker session ids, and log block rates per hour so you can spot time-of-day patterns.
From 100x to 1000x (hundreds of thousands of pages or more): you will want several small machines rather than one big one, each with its own proxy credentials and its own rate budget. Add a scheduler that spreads crawls across days instead of bursts, and do incremental refreshes (only product pages whose review count changed) instead of full recrawls. At this volume mobile proxies are worth testing for the hardest pages, since shared carrier IPs are costly for a site to block. If you run many accounts or profiles alongside the scraping, the browser-fingerprint side is covered over at antidetectreview.org, and the operational side of juggling many identities lives at multiaccountops.com.
Also expect maintenance. Selectors break, and block rates drift upward and downward through the year.
where to go next
- browse everything we have published in the blog index
- how to scrape G2 reviews with proxies covers a similar review-site target with different protection
- residential vs mobile proxies for scraping helps you pick the right pool before you spend on bandwidth
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-10.