← all guides

How to scrape AliExpress at scale in 2026 with proxies that work

AliExpress is one of the harder marketplaces to pull data from. A script that fetches 50 product pages from your laptop on Monday will often hit a slider captcha or a “punish” redirect by Tuesday. The pages are heavy, much of the data is rendered client side, and the anti-bot layer scores your IP, your TLS fingerprint and your behaviour together. Throwing more threads at it makes things worse, not faster.

This tutorial is for operators who need price, stock and seller data across thousands to hundreds of thousands of listings. Think dropshipping research, price monitoring, or catalogue enrichment. I run proxy infrastructure out of Singapore, so the setup below comes from what I have seen working and failing on real runs, not from a vendor deck. Treat the specifics as a snapshot of 2026, because AliExpress changes its defences often.

By the end you will have a Playwright based scraper that routes through rotating residential or mobile proxies, keeps sessions sticky where it matters, backs off when it gets challenged, and writes clean JSON. You will also know which parts to change when you go from a few hundred pages a day to a few hundred thousand.

What you need

  • a machine with Python 3.11 or newer and about 4 GB of free RAM per 4 browser contexts
  • Playwright for Python (pip install playwright then playwright install chromium)
  • a proxy pool that supports sticky sessions: residential or mobile 4G/5G. datacenter IPs get challenged fast on AliExpress in my experience
  • a list of item IDs or product URLs (the numeric ID in /item/1005001234567890.html)
  • somewhere to store results: SQLite is fine up to a few million rows, Postgres after that
  • a budget for bandwidth: residential traffic is billed per GB and product pages are heavy, so block images and media (step 4) or the bill gets silly. check your provider’s current per GB rate before you plan, since prices move
  • about two hours for the first working version

A word on legality before we start. Read AliExpress’s terms and its robots.txt and decide what you are comfortable with. Robots.txt is a convention defined in RFC 9309, not a legal shield or a legal permission either way. Only collect public product data, never log into accounts that are not yours, and do not collect personal data about buyers. This is not legal advice, so talk to a lawyer if your use case is commercial and sizeable.

Step by step

1. Pick the right proxy type

Action: buy a small test plan from two providers, one residential and one mobile, and run 200 requests through each before you commit.

In my tests residential rotating pools work for most product pages. Mobile proxies (carrier IPs shared by many real users) survive longer on the stricter endpoints, but cost more per GB or are sold per port per month. Datacenter IPs are cheap and nearly always end up on the captcha page, so I only use them for things like checking whether a URL is alive.

Expected output: a simple table of success rate per proxy type, where success means you got a page containing the product title and price.

If it breaks: if both types fail above 30 percent, the problem is probably your browser fingerprint and not the IP. Jump to step 3 first and retest.

2. Set the region and currency up front

Action: AliExpress shows different prices, shipping and even stock depending on the country and currency tied to the session. Pick one target region and keep the proxy exit country matching it. If you are comparing Singapore prices, use Singapore exits and a SGD or USD currency, and do not mix.

AliExpress stores these preferences in cookies (the aep_usuc_f cookie carries site, region, currency and locale). Rather than guess the format, load the site once in a normal browser through your proxy, set the country and currency in the footer, then export the cookies from that session and reuse them in your scraper.

Expected output: product pages that show the same currency and ship-to country on every run.

If it breaks: if prices flip between currencies, your cookie is being overwritten. Make sure the proxy country, the browser locale and the cookie region all agree.

3. Launch Playwright with a clean, consistent profile

Action: use real Chromium, a normal desktop viewport and a timezone and locale that match the proxy exit. Mismatches (a Brazilian IP with a Singapore timezone) are an easy signal for the anti-bot layer.

from playwright.sync_api import sync_playwright

PROXY = {
    "server": "http://gate.example-proxy.com:7000",
    "username": "USER-session-abc123-country-sg",
    "password": "PASS",
}

def new_context(p):
    browser = p.chromium.launch(headless=True, proxy=PROXY)
    return browser.new_context(
        viewport={"width": 1366, "height": 850},
        locale="en-SG",
        timezone_id="Asia/Singapore",
        user_agent=(
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
            "(KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36"
        ),
    )

The proxy argument format is documented in the Playwright network guide. The session and country tokens in the username are provider specific, so check your provider’s docs for the exact syntax. Keep the user agent version close to the Chromium build you installed, because a mismatch there is also detectable.

Expected output: a context that can open https://www.aliexpress.com/ and show the normal homepage.

If it breaks: if the homepage loads but product pages redirect, add a short realistic path: homepage, then a category or search page, then the product. Direct hits on item URLs with no referrer are more likely to be challenged.

4. Block what you do not need

Action: product pages pull a lot of images, video and tracking scripts. Cut the images, media and fonts to save bandwidth, which matters a lot on per GB residential plans.

def block_heavy(route):
    if route.request.resource_type in ("image", "media", "font"):
        return route.abort()
    return route.continue_()

page = context.new_page()
page.route("**/*", block_heavy)

Do not block scripts or XHR. The price and SKU data often arrive through those calls, and removing them gives you empty pages.

Expected output: page weight dropping noticeably, usually by well over half in my runs, though it varies by listing.

If it breaks: if fields go missing after blocking, loosen the filter one resource type at a time until the data comes back.

5. Extract from embedded JSON, not from the DOM

Action: open a product page, view source and search for the price. In my experience AliExpress embeds structured data in the page (look for window.runParams or JSON inside script tags) and that survives layout changes better than CSS selectors.

import json, re

def parse_item(html: str) -> dict | None:
    m = re.search(r"window\.runParams\s*=\s*(\{.*?\});", html, re.S)
    if not m:
        return None
    data = json.loads(m.group(1))
    return data  # then pick title, sku prices, stock, seller id

The exact key names change, so inspect a saved page and write your mapping against it. Save the raw HTML of a handful of pages as fixtures so you can test the parser offline.

Expected output: one JSON record per item with id, title, price range, stock, rating and seller id.

If it breaks: if the regex returns nothing, the page is probably a challenge page. Check for the captcha markers (step 7) before assuming your parser is wrong.

6. Use sticky sessions and sane pacing

Action: one proxy session per browser context, held for a few minutes, with 3 to 8 seconds of jitter between pages. Rotating the IP on every request looks unnatural for a browser that keeps its cookies, and it burns through your pool.

import random, time

def polite_wait():
    time.sleep(random.uniform(3, 8))

Rotate the session after a set number of pages (I use 20 to 40) or immediately on a challenge.

Expected output: steady success rate over a 1,000 page run instead of a sharp drop after the first few dozen.

If it breaks: if you get challenged right on the first request of a fresh session, that exit IP is burnt. Ask your provider for a new one or switch pools.

7. Detect challenges and back off

Action: treat a challenge as a signal, not an error. Check for redirects to a punish or captcha URL, a slider element, or a missing product title, then rotate the session and requeue the item.

def is_blocked(page) -> bool:
    url = page.url.lower()
    if "punish" in url or "captcha" in url:
        return True
    return page.locator("h1").count() == 0

Do not try to solve the slider at volume. It is slow, adds cost, and most of the time a fresh IP and a calmer pace fix the problem for free.

Expected output: a retry queue with a count per item, and a log line for every rotation.

If it breaks: if the same item keeps failing across fresh sessions, it may be removed or region restricted. Cap retries at 3 and mark it dead.

8. Store results and make runs resumable

Action: write each result to SQLite as it arrives and keep a status column (pending, done, failed). A crash at item 40,000 should cost you minutes, not the whole run.

import sqlite3
db = sqlite3.connect("aliexpress.db")
db.execute("""CREATE TABLE IF NOT EXISTS items(
  id TEXT PRIMARY KEY, status TEXT, payload TEXT, tries INT DEFAULT 0)""")

Expected output: a database where you can rerun the script and it only picks up pending and failed rows with fewer than 3 tries.

If it breaks: if you see duplicates, check that you key on the numeric item ID and not the full URL, since AliExpress URLs carry tracking parameters.

Common pitfalls

  • using datacenter proxies to save money. you save on the plan and lose it again on retries and wasted engineering time
  • rotating the IP on every request while keeping the same cookies. real users do not change country between clicks
  • mismatching proxy country, browser locale and cookie region, which gives you mixed prices and a higher challenge rate
  • scraping with raw requests or httpx and expecting it to work. the httpx proxy docs show the plumbing is easy, but a plain HTTP client has no JavaScript and a very recognisable TLS fingerprint, so I only use it for endpoints I have already confirmed return JSON
  • ignoring bandwidth until the invoice arrives. measure GB per 1,000 pages on day one and multiply

Scaling this

At 10x (a few thousand pages a day) one machine and one small residential plan is enough. Run 2 to 4 browser contexts in parallel, keep the SQLite file, and watch your success rate daily.

At 100x (hundreds of thousands of pages a month) bandwidth becomes the main cost, so measure it per page and keep the image blocking tight. Move to a queue (Redis or a Postgres table with row locking) and run workers on a few small VPS instances so a single bad exit IP does not stall everything. Split proxies by purpose: a cheaper pool for category and search pages, a better pool for item pages. This is also where managing browser state starts to hurt, and I wrote about the point where hosted infrastructure starts to make sense in when to move to a managed browser farm.

At 1000x you are running a data pipeline, not a script. Expect to negotiate volume pricing with proxy vendors, monitor success rate per exit country and per ASN, and keep a rotating set of fingerprints. If you are juggling many browser identities, the tooling in the antidetect browser reviews is worth a look, since the same fingerprint consistency problems show up there. At this scale you should also consider whether an official or licensed data source covers part of what you need, because it is often cheaper than winning an arms race.

Where to go next

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-08.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →