How to scrape Google Shopping at scale in 2026 with proxies that work
Google Shopping is one of the most useful price datasets on the open web and one of the most annoying to collect. A plain requests call gets you a consent page or a captcha within a few dozen hits. Datacenter IPs get burned fast. The markup changes often enough that a parser you wrote in spring is quietly returning empty fields by autumn.
This is for people who already run scrapers and want a Google Shopping pipeline that survives past the first thousand queries. Price monitoring, assortment tracking, ad-hoc market research. I’m Xavier, I run proxy infrastructure out of Singapore, and what follows is the setup I’d build today, including the parts that break.
By the end you’ll have a Playwright-based collector that rotates mobile or residential proxies, backs off properly when Google pushes back, parses product cards into a clean schema, and writes to SQLite so you can resume after a crash. It’s a starting point, not a finished product.
what you need
- python 3.11 or newer and a machine with at least 4 GB of RAM per 4 to 5 concurrent browser contexts
- playwright for python, installed with
pip install playwrightthenplaywright install chromium. docs are at playwright.dev - a proxy pool with real-user IPs: residential or 4G/5G mobile. datacenter pools work for a few hundred queries a day at best
- geo-targeted exit nodes for each market you track. Google Shopping results and prices vary by country and sometimes by city
- a keyword or product list, ideally GTINs or model numbers rather than vague terms
- sqlite (ships with python) for state and results
- budget: proxy bandwidth is the main cost. check your provider’s current per-GB or per-port pricing before you commit, because it changes often and a browser-based scraper pulls far more bytes per query than an HTTP client
- a read of Google’s terms of service. this is not legal advice. whether scraping is acceptable depends on your jurisdiction, your use, and the terms you’re bound by
One more option before you build anything. If you only need data for your own products, the Merchant Center Content API gives it to you cleanly with no proxies at all. Scraping is for competitor and market data you can’t get any other way.
step by step
step 1: pick the proxy type for the job
Action: decide between residential and mobile before writing code.
Residential IPs are cheaper per GB and fine for moderate volume. Mobile (4G/5G) IPs share carrier-grade NAT with thousands of real phones, so Google is reluctant to block them hard. They cost more, and they’re the thing I reach for when residential pools start returning captchas. I compare the two in more detail in residential vs mobile proxies for SERP scraping.
Expected output: a proxy endpoint in host:port form with username and password, plus a way to request a fresh IP (a rotation URL or a new session ID).
If it breaks: test the proxy first with curl -x http://user:pass@host:port https://ipinfo.io/json. if the country or ASN is wrong, fix that with your provider before touching Google.
step 2: set up the project
Action: create a virtual environment and install dependencies.
python -m venv .venv
source .venv/bin/activate # on windows: .venv\Scripts\activate
pip install playwright selectolax
playwright install chromium
Expected output: playwright --version prints a version number and the chromium download completes without errors.
If it breaks: on a headless linux box, run playwright install-deps to pull the system libraries chromium needs.
step 3: build the job queue in SQLite
Action: create a table that holds every query, its country, and its status. This is what makes the scraper resumable, and it matters more than any anti-bot trick.
import sqlite3
db = sqlite3.connect("shopping.db")
db.executescript("""
CREATE TABLE IF NOT EXISTS jobs (
id INTEGER PRIMARY KEY,
query TEXT, country TEXT,
status TEXT DEFAULT 'pending',
attempts INTEGER DEFAULT 0,
UNIQUE(query, country)
);
CREATE TABLE IF NOT EXISTS offers (
job_id INTEGER, title TEXT, price TEXT,
merchant TEXT, url TEXT, scraped_at TEXT
);
""")
db.commit()
Expected output: a shopping.db file with two empty tables. Load your keywords with a simple INSERT OR IGNORE.
If it breaks: “database is locked” means two workers are writing at once. Use one writer process, or set PRAGMA journal_mode=WAL; once at startup.
step 4: fetch a results page through the proxy
Action: open the Shopping results URL in a Playwright context that uses your proxy, with a sensible locale and viewport.
from playwright.sync_api import sync_playwright
from urllib.parse import quote_plus
PROXY = {"server": "http://host:port", "username": "user", "password": "pass"}
def fetch(query, gl="us", hl="en"):
url = f"https://www.google.com/search?q={quote_plus(query)}&tbm=shop&gl={gl}&hl={hl}"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True, proxy=PROXY)
ctx = browser.new_context(locale="en-US", viewport={"width": 1366, "height": 900})
page = ctx.new_page()
resp = page.goto(url, wait_until="domcontentloaded", timeout=45000)
page.wait_for_timeout(1500)
html = page.content()
status = resp.status if resp else 0
browser.close()
return status, html
Expected output: status 200 and an HTML string a few hundred KB long that contains product titles and prices.
If it breaks: a redirect to a consent.google.com URL means you’re in a region that shows the cookie wall. Click through it once, save the cookies with ctx.storage_state(path="state.json"), and reuse that state. A page containing “unusual traffic” means that IP is flagged. Rotate it and move on, don’t retry on the same one.
step 5: parse product cards defensively
Action: extract title, price, merchant and link from the HTML.
Google’s class names are obfuscated and rotate. Don’t build your parser around them. Anchor on things that change less: the /shopping/product/ or /url? link patterns, aria-label text, and the presence of a currency symbol. Expect to revisit this step every few weeks regardless.
import re
from selectolax.parser import HTMLParser
PRICE = re.compile(r"[\$€£S]\$?\s?\d[\d,]*\.?\d*")
def parse(html):
tree = HTMLParser(html)
rows = []
for a in tree.css("a[href*='/shopping/product/']"):
text = a.text(separator=" ", strip=True)
m = PRICE.search(text)
if m:
rows.append({"title": text[:140], "price": m.group(0),
"url": a.attributes.get("href")})
return rows
Expected output: a list of dicts, usually 10 to 40 per page for a mainstream product query.
If it breaks: if you get zero rows with a 200 status, save the raw HTML to disk and open it. Nine times out of ten you got a degraded page, a captcha, or a layout variant. Log the page length alongside the row count so you can spot the pattern.
step 6: add retries, backoff and proxy rotation
Action: wrap the fetch and parse in a loop that treats a block as normal, not exceptional.
Google signals pressure with 429 responses, 503s, captcha pages, and redirects to /sorry/. The 429 status is defined in RFC 6585, and the polite response is to slow down. Mine is: rotate IP, wait with exponential backoff plus jitter, cap attempts at 4.
import random, time
def run_job(job_id, query, country):
for attempt in range(4):
status, html = fetch(query, gl=country)
blocked = status in (429, 503) or "/sorry/" in html or "unusual traffic" in html
if not blocked:
rows = parse(html)
if rows:
return rows
rotate_ip() # call your provider's rotation endpoint here
time.sleep((2 ** attempt) * 3 + random.uniform(0, 3))
return None
Expected output: most jobs succeed on the first or second attempt. A block rate of a few percent on mobile IPs is healthy.
If it breaks: if more than roughly a quarter of requests are blocked, the problem is usually not the code. Check proxy quality, then your request rate, then whether you’re re-using one browser fingerprint on every request.
step 7: pace yourself per IP and per market
Action: add a global rate limit and a per-IP cap. I run 1 request every 4 to 8 seconds per exit IP, with 3 to 5 concurrent workers per proxy port, and I rotate after 20 to 40 queries or the first sign of friction. Those are my numbers from my own setups, not a Google-published threshold, so tune them against your own block rate.
Expected output: a flat, boring request graph with no bursts.
If it breaks: bursts usually come from a worker pool that all wakes up at once after a backoff. Add jitter to the worker start times.
step 8: store results and make runs idempotent
Action: write offers in the same transaction that marks the job done, so a crash never leaves a half-finished job.
from datetime import datetime, timezone
def save(job_id, rows):
now = datetime.now(timezone.utc).isoformat()
with db:
db.executemany(
"INSERT INTO offers VALUES (?,?,?,?,?,?)",
[(job_id, r["title"], r["price"], None, r["url"], now) for r in rows])
db.execute("UPDATE jobs SET status='done' WHERE id=?", (job_id,))
Expected output: SELECT status, COUNT(*) FROM jobs GROUP BY status; shows the pending count falling and the done count rising.
If it breaks: duplicates on re-runs mean you’re inserting before checking status. Only pick jobs where status='pending', and mark a job failed after its attempts run out.
step 9: validate a sample by hand
Action: every run, pull 20 random rows and compare them to the live page in a normal browser from the same country.
Expected output: titles and prices match, give or take personalised ranking.
If it breaks: wrong currency or wrong prices almost always means the exit IP was in the wrong country. Verify geo per session, not per account.
common pitfalls
- trusting a 200 status. Google serves captcha and consent pages with a 200. always check the content, not just the code
- using datacenter proxies for volume. they work in a demo and die in production. budget for real-user IPs from the start
- parsing by obfuscated class names. it passes your test run and fails two weeks later without raising an error. anchor on structure and validate row counts
- retrying on the same IP. a flagged IP stays flagged for a while, so every retry on it just makes the pool worse
- ignoring geography. a US price scraped through a Singapore exit is a wrong price. check the location of every session
- not storing raw HTML for failures. you can’t debug a layout change you didn’t save
If you’re running many accounts or profiles alongside this, browser fingerprint consistency becomes its own problem. antidetectreview.org covers that side in more detail than I do here.
scaling this
From 10x to 100x to 1000x, different things become the bottleneck.
10x (roughly a few thousand queries a day): the single script above is fine. One machine, one proxy port, SQLite. Your main job is getting the block rate down and the parser stable.
100x (tens of thousands a day): move from one proxy port to a pool and track success rate per IP and per ASN. Split work across several processes, keep SQLite as the queue with a single writer, or move the queue to Postgres. Start logging cost per successful query, not per request, because blocked requests still burn bandwidth. Block images, fonts and media in Playwright with page.route to cut bytes dramatically.
1000x (hundreds of thousands a day and up): you’re now running a small distributed system. A real queue (Redis or similar), stateless workers on several machines, per-market proxy pools, and monitoring that alerts on parse-yield drops, not just errors. At this volume the proxy bill dominates everything, so a self-managed mobile modem pool can start to beat per-GB pricing. Check the maths against your own volume first. Also expect your parser to need weekly attention, and decide whether paid SERP APIs are cheaper than your engineer time. For many teams they are.
where to go next
- all tutorials on the blog
- build a proxy rotation layer in Python
- residential vs mobile proxies for SERP scraping
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-12.