How to scrape Best Buy at scale in 2026 with proxies that work
Best Buy is one of the more annoying big-box retailers to scrape. Prices and stock change by region and by store, the pages are heavy JavaScript, and the site sits behind bot detection that gets suspicious fast when the traffic looks like a datacenter. If you point a plain requests loop at it from a cloud VPS, you will get a wall of 403s or empty shells within a few hundred requests. I have watched this happen on my own jobs more than once.
This tutorial is for people who need price, availability or catalog data from Best Buy on a recurring basis: price trackers, deal sites, resellers doing sourcing research, small analytics shops. I run proxy infrastructure out of Singapore, so my view is the operator’s view. I care about cost per successful page, not about demos that work ten times.
The outcome is a working pipeline: a list of SKUs goes in, clean JSON comes out, traffic runs through residential or mobile proxies with sane session handling, and you know what changes when you push it from hundreds of pages to hundreds of thousands. One caveat first. Check the official route before you scrape. Best Buy publishes a developer API with product and store data, and for plenty of use cases it is cheaper and safer than scraping. Scraping is for what the API does not give you.
what you need
- python 3.11 or newer, with
playwrightandhttpxinstalled - a residential proxy pool with US geo-targeting and sticky sessions. mobile proxies work well too but cost more per GB
- a place to store results. sqlite is fine to start, Postgres once you go past a few million rows
- a SKU list. you can build one from Best Buy category pages or the sitemap, or from your own catalog
- a budget. residential bandwidth is billed per GB, and a full browser-rendered Best Buy product page can pull a lot of data unless you block images and fonts. I plan around my own measured bytes per page rather than anyone’s marketing numbers
- 1 to 2 hours for setup, then monitoring time every week
- a read through of the site’s robots.txt and terms, plus the Robots Exclusion Protocol spec (RFC 9309) so you know what the file does and does not mean
This is not legal advice. Scraping rules differ by country and by what you do with the data. Talk to a lawyer if you plan to resell it or build a product on it.
step by step
1. Decide if the official API covers you
Register for a key at the Best Buy developer portal and test a few SKUs. If name, price, sale price and availability are enough, stop here and use the API. It will be more stable than any scraper.
Expected output: a JSON response with the fields you need for 5 to 10 test SKUs.
If it breaks: a 403 or empty body usually means the key is not activated yet or the query is malformed. Check the portal docs. If the API simply lacks fields you need (store-level pickup details, some marketplace data, rendered page content), carry on with the steps below.
2. Pick the proxy type for the job
Datacenter IPs get flagged early on Best Buy. Residential proxies pass much more often because they look like home users. Mobile proxies (4G/5G carrier IPs) pass the most, because carrier-grade NAT means hundreds of real users share each IP, and sites are reluctant to block that.
My rule of thumb:
- up to roughly 50k pages a day: rotating residential, sticky sessions of 5 to 10 minutes
- bursty or high-value jobs where failures are costly: mobile proxies
- datacenter: only for the sitemap and static asset fetches that are not protected
Expected output: a proxy endpoint you can test with curl.
curl -x "http://USER:[email protected]:7000" https://api.ipify.org
If it breaks: you get your own IP back, so the proxy is being bypassed. Or you get a connection error, so check the port, auth format and that your IP is allowlisted if the provider requires that. The .example host above is a placeholder, use your provider’s real gateway.
3. Geo-target to the US, ideally the right metro
Best Buy shows different pickup and delivery options by ZIP code. If you need consistent numbers, pin your proxy to one US region and set the store or ZIP in the session before you read prices. Most residential providers let you add a country or city flag to the username. The syntax differs by provider, so read their docs.
Expected output: an IP-lookup check that confirms a US location.
If it breaks: you are being routed through the wrong country and Best Buy redirects you to the Canadian site or shows a region notice. Fix the geo flag before you waste bandwidth.
4. Set up Playwright with the proxy
Playwright handles JavaScript rendering and gives you control over resource blocking. The Playwright Python network docs cover proxy options in detail.
from playwright.sync_api import sync_playwright
PROXY = {
"server": "http://gate.yourprovider.example:7000",
"username": "USER-session-abc123",
"password": "PASS",
}
def fetch(url: str) -> str:
with sync_playwright() as p:
browser = p.chromium.launch(headless=True, proxy=PROXY)
ctx = browser.new_context(
locale="en-US",
timezone_id="America/New_York",
viewport={"width": 1366, "height": 800},
)
page = ctx.new_page()
page.route(
"**/*",
lambda r: r.abort()
if r.request.resource_type in ("image", "font", "media")
else r.continue_(),
)
page.goto(url, wait_until="domcontentloaded", timeout=45000)
page.wait_for_timeout(1500)
html = page.content()
browser.close()
return html
Expected output: HTML of a product page that contains the product title and a price.
If it breaks: if you get a short page with a “Access Denied” style message, the proxy IP or your browser fingerprint is flagged. Try a fresh sticky session first. If that does not help, the browser fingerprint is the problem. I cover that in the pitfalls section, and the people at antidetectreview.org test browser fingerprint tools in more depth than I do here.
5. Parse structured data, not CSS classes
Class names on retail sites change without notice. Many product pages embed JSON-LD (application/ld+json) with name, SKU, price and availability. Parse that first and fall back to selectors only if it is missing.
import json, re
def parse_product(html: str) -> dict | None:
for m in re.finditer(
r'<script type="application/ld\+json">(.*?)</script>', html, re.S
):
try:
data = json.loads(m.group(1))
except json.JSONDecodeError:
continue
items = data if isinstance(data, list) else [data]
for item in items:
if item.get("@type") == "Product":
offer = item.get("offers", {})
if isinstance(offer, list):
offer = offer[0] if offer else {}
return {
"sku": item.get("sku"),
"name": item.get("name"),
"price": offer.get("price"),
"currency": offer.get("priceCurrency"),
"availability": offer.get("availability"),
}
return None
Expected output: a dict like {"sku": "...", "name": "...", "price": ..., ...} for each good page.
If it breaks: None for pages that clearly loaded means the structure shifted. Save the raw HTML of the failing page to disk and inspect it. Do not guess selectors blind.
6. Add sessions, retries and backoff
One sticky session per small batch of pages, not one per request and not one for the whole job. Rotate the session on a block, wait with jittered backoff, and cap retries so a bad URL cannot eat your bandwidth.
import random, time, uuid
def run(urls, max_retries=3):
results, failed = [], []
for url in urls:
for attempt in range(max_retries):
PROXY["username"] = f"USER-session-{uuid.uuid4().hex[:8]}"
try:
html = fetch(url)
data = parse_product(html)
if data:
results.append(data)
break
except Exception:
pass
time.sleep(2 ** attempt + random.random() * 2)
else:
failed.append(url)
time.sleep(random.uniform(2, 6))
return results, failed
Expected output: a results list and a short failed list you can retry later.
If it breaks: if more than 20 to 30 percent of requests fail, stop. Do not just add retries. Something is wrong with the proxy type, the fingerprint or your speed, and retries multiply your bandwidth bill.
7. Store results with a timestamp and the region
Write one row per SKU per fetch, with the fetch time, proxy region and raw price string. Price history is the valuable part, and you cannot reconstruct it later.
CREATE TABLE prices (
sku TEXT, fetched_at TEXT, region TEXT,
price REAL, availability TEXT, raw TEXT
);
Expected output: rows accumulating after each run, no duplicates per run.
If it breaks: duplicates usually mean a retry wrote twice. Add a unique key on (sku, fetched_at, region) or dedupe per run.
8. Measure before you scale
Run 200 SKUs and record three numbers: success rate, megabytes per successful page, and seconds per page. Multiply by your target volume. That is your real cost, and it will differ from what any provider’s pricing page implies.
Expected output: a small report with those three figures.
If it breaks: if megabytes per page is high, check that your route blocking works and that you are not loading third-party trackers. Blocking images, fonts and media usually cuts transfer a lot.
common pitfalls
- using datacenter proxies to save money. the saving disappears when 70 percent of requests fail and you pay for every retry in time and compute.
- rotating the IP on every request. a real shopper keeps one IP for a few minutes. constant rotation within a single browser context looks odd and also breaks cookies and store selection.
- ignoring the fingerprint. a clean residential IP with a default headless Chromium can still get flagged. match locale, timezone and viewport to the proxy region, and test with a headed browser when debugging.
- scraping the whole catalog at the same speed all day. spread load, avoid hammering the same category, and keep concurrency modest per IP. being polite is also the cheapest way to avoid blocks.
- treating prices as universal. Best Buy shows different availability and sometimes different offers by location. if you do not log the region, your data is ambiguous.
scaling this
At 10x (a few thousand pages a day) nothing much changes. One machine, one proxy plan, sqlite, a cron job. Your main job is watching the failure rate and fixing the parser when the page changes.
At 100x (hundreds of thousands of pages a month) bandwidth becomes the dominant cost. You should be blocking all non-essential resources, caching category and sitemap fetches, and checking whether a lightweight HTTP request to a JSON endpoint can replace a full browser for part of your workload. Only do that where the endpoint is accessible without circumventing anything you should not be bypassing. Move to a queue (Redis or a hosted queue), run several workers, and give each worker its own sticky session pool. See my notes on choosing residential proxies for ecommerce scraping for how I compare providers.
At 1000x you are running infrastructure, not a script. You will want per-worker proxy budgets, automatic session rotation based on block signals, a dashboard for success rate by region and hour, and alerts when the parser breaks. At that volume it is worth asking whether a direct data agreement or a commercial feed is cheaper than keeping a scraping operation alive. For some of my own jobs it was. Mobile proxies earn their price here for the hardest targets, and a pool of your own modems can bring the per-GB cost down if you have the ops skills for it.
where to go next
- the full article index is at /blog/
- how to rotate proxies in Python with Playwright goes deeper on session handling than I did in step 6
- residential vs mobile proxies for retail scraping helps you decide which to pay for
- if you also run many accounts or profiles for the same kind of work, multiaccountops.com covers the account side of operations
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-09.