How to scrape Expedia at scale in 2026 with proxies that work
Expedia is one of the harder travel targets to collect from. Prices are rendered client side, results change by location and device, and the site is quick to serve a challenge page or a quietly degraded result set when traffic looks automated. A script that returns 40 hotels on your laptop will often return empty pages or captchas within a few hundred requests from a datacenter range.
This guide is for people who need price and availability data for research, rate monitoring or market analysis, and who already know basic Python. I run mobile and residential proxy infrastructure out of Singapore, so what follows comes from operating this stuff, not from a vendor deck. I will not pretend there is a setting that makes blocking go away. What I can give you is a setup that keeps block rates low enough to be workable, and a way to tell when it stops being so.
By the end you will have a Playwright based collector that rotates sessions through proxies, saves structured hotel results to disk, and logs enough to show you where it fails. I will also cover what changes when you move from a few hundred pages to hundreds of thousands.
What you need
- python 3.11 or newer, with Playwright installed (see the Playwright Python docs)
- a proxy pool: residential or mobile IPs with sticky session support. datacenter IPs work for light testing and then fall over
- a place to store results: sqlite is fine to start, Postgres once you pass a few million rows
- a machine with 8 GB of RAM for around 4 to 6 concurrent browser contexts. headless Chromium is hungrier than people expect
- a budget. residential bandwidth is billed per GB and mobile is often billed per port or per GB, so check the current price sheet from your provider before you size anything. I am not quoting numbers here because they move every few months
- an hour to read Expedia’s robots.txt and terms of use before you start
On legality: scraping public pages sits in a grey area that varies by country and by what you do with the data. This is not legal advice. If the data feeds a commercial product, talk to a lawyer who knows your jurisdiction. If you want the sanctioned route for hotel inventory, Expedia Group has an official partner program at developers.expediagroup.com, and for some use cases that is the better answer.
Step by step
1. Pick the right proxy type
Do this first, because everything else depends on it.
- datacenter: cheap and fast, and flagged quickly on travel sites. use them to debug your parser, not to collect
- residential rotating: the usual workhorse. good diversity, reasonable cost per page
- mobile (4G/5G): the highest trust because carrier NAT puts many real users behind one IP. most expensive, but block rates are the lowest in my experience
For Expedia I start with residential and keep a small mobile pool for the pages that get challenged most. I compare the two in more detail in mobile vs residential proxies for scraping.
Expected output: a decision, written down, with a cost per 1,000 successful pages you will measure later.
If it breaks: if you cannot tell which type is failing, run the same 50 URLs through each and compare success rates. Do not guess.
2. Test the proxy before you test Expedia
Confirm the exit IP, country and latency with a plain request.
curl -x "http://USER:[email protected]:8000" https://ipinfo.io/json
Expected output: JSON with an IP, a city and an org name that matches the carrier or ISP you paid for.
If it breaks: a 407 means bad credentials, a timeout usually means a wrong port or an IP allowlist you forgot to update. If the org says a hosting company, your “residential” pool is not what was sold to you.
3. Match the proxy geography to the site version
Expedia serves different currencies, taxes and sometimes different inventory depending on the domain and the visitor’s location. If you want US prices from expedia.com, use US exits. If you want Singapore pricing from expedia.com.sg, use Singapore exits and set the locale to match. Mixing a German IP with a US locale and a UTC clock is an easy way to look fake.
Expected output: a country parameter in your proxy config that you set per job, not globally.
If it breaks: if prices come back in the wrong currency, you are being geolocated somewhere you did not intend. Print the exit country on every job.
4. Build the browser session
Plain HTTP requests to Expedia search pages return a shell, and the useful data arrives after JavaScript runs. Use a real browser. Here is the minimum I run:
from playwright.sync_api import sync_playwright
PROXY = {
"server": "http://gate.yourprovider.com:8000",
"username": "USER-session-abc123",
"password": "PASS",
}
def open_page(url):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True, proxy=PROXY)
ctx = browser.new_context(
locale="en-US",
timezone_id="America/New_York",
viewport={"width": 1366, "height": 800},
)
page = ctx.new_page()
page.goto(url, wait_until="domcontentloaded", timeout=45000)
page.wait_for_selector("[data-stid='lodging-card-responsive']", timeout=20000)
html = page.content()
browser.close()
return html
Selectors like data-stid change when Expedia ships front end updates, so treat the one above as an example to verify in devtools, not a promise.
Expected output: an HTML string containing the hotel cards.
If it breaks: if the selector times out, save a screenshot with page.screenshot(path="fail.png"). Nine times out of ten it is a challenge page and the screenshot tells you straight away.
5. Keep sessions sticky for the length of a search
A search is not one request. You land, set dates, scroll, paginate. If your IP changes halfway, the session looks broken. Most providers let you pin an IP with a session id in the username, as in the code above (session-abc123). Use one session id per search, then retire it.
Expected output: all pages of one search come from the same exit IP. Verify by logging the IP at the start and end.
If it breaks: if the IP changes mid search, your provider’s sticky window may be shorter than your job. Shorten the job or ask for a longer TTL.
6. Pace requests like a person who is slightly bored
Fixed intervals are a fingerprint. Use jittered waits between actions and cap concurrency per exit IP at one.
import random, time
def human_pause(low=1.8, high=5.5):
time.sleep(random.uniform(low, high))
Expected output: request timing that does not look like a metronome in your logs.
If it breaks: rising challenge rates after a speed up mean you went too fast. Back off by 30 percent and measure again. Do not respond to blocks by adding more threads.
7. Parse into a structured record
Pull out the fields you need and drop the rest. For a hotel card that is usually name, price, currency, rating, review count, the dates searched and a timestamp.
from selectolax.parser import HTMLParser
def parse(html, checkin, checkout):
rows = []
for card in HTMLParser(html).css("[data-stid='lodging-card-responsive']"):
name = card.css_first("h3")
price = card.css_first("[data-test-id='price-summary']")
rows.append({
"name": name.text(strip=True) if name else None,
"price_text": price.text(strip=True) if price else None,
"checkin": checkin,
"checkout": checkout,
})
return rows
Store the raw price text and parse the number in a second pass. Expedia shows nightly and total prices in different places, and you will want to reprocess later without refetching.
Expected output: a list of dicts, usually 20 to 40 per results page.
If it breaks: if names are filled but prices are None, the price element is probably lazy loaded. Scroll the page by a few hundred pixels before reading page.content().
8. Detect blocks and rotate instead of retrying blindly
Check for a block on every response. Signs include a short page, a captcha iframe, a redirect to a verification path, or zero hotel cards where you expected some.
def is_blocked(html):
return len(html) < 20000 or "captcha" in html.lower()
On a block: drop that session id, take a fresh IP, wait, and retry at most twice. Log the block with the IP, the time and the URL.
Expected output: a block rate figure per hour. Under 10 percent is healthy for me, over 25 percent means something upstream changed.
If it breaks: if every IP is blocked at once, the problem is your browser fingerprint or your request pattern, not your proxies. Look at the browser side before buying more IPs. I covered fingerprint tooling in the antidetect browser reviews on our sister site.
9. Store, dedupe and schedule
Write rows to sqlite with a unique key on property id, dates and capture hour. Run a scheduler (cron on Linux, Task Scheduler on Windows) that feeds a job queue instead of looping in one long script.
Expected output: a database that grows steadily and never holds duplicate rows for the same property and stay.
If it breaks: duplicates usually mean retries wrote partial results. Write a page’s rows in a single transaction after it parses cleanly.
Common pitfalls
- using one fast IP: operators test with a single clean IP, see success, then scale up the same pattern. the IP burns and the whole project looks broken
- ignoring locale and timezone: the exit says Tokyo, the browser says en-US and America/New_York. consistency matters more than people think
- retrying instantly: a blocked request retried within a second from the same IP teaches the site to block that IP for longer
- scraping dates you do not need: a 365 day calendar crawl for every property multiplies cost by hundreds. decide your date window first
- skipping robots.txt and the terms: read them, decide what you are comfortable with, and keep a record of that decision. RFC 9309 explains how robots.txt is meant to be interpreted, and it is advisory, not a permission slip
Scaling this
At 10x, so a few thousand pages a day, nothing structural changes. One machine, a residential pool, and a log you actually read. This is also when you find out your real block rate and your real cost per 1,000 pages.
At 100x, around a few hundred thousand pages a day, you need a job queue (Redis or a Postgres table with SELECT ... FOR UPDATE SKIP LOCKED), several worker machines, and per worker proxy accounting. Bandwidth becomes the main bill. Block images, fonts and media in Playwright with request routing and you can cut transfer per page a lot. Start tracking cost per successful row, not cost per request.
At 1000x you are running a data operation. Expect to split proxy sources across two or three providers so one provider’s bad week does not stop you, to keep a dedicated mobile pool for the hardest pages, and to run a canary job every ten minutes that fetches a known search and alerts you when the result shape changes. Selector breakage becomes weekly maintenance, not a one time fix. Also ask honestly whether the official Expedia Group partner APIs cover your need, because at this volume the engineering hours often cost more than a licensed feed.
If you run many accounts or sessions in parallel, the isolation habits from our friends at multiaccountops.com apply here too: one identity per browser profile, nothing shared.
Where to go next
- planning a crawl from the sitemap: how to size and order a crawl before you spend any proxy bandwidth
- rotating proxy session settings explained: sticky windows, session ids and how to test them
- the full article index: everything else we have written on proxies and scraping targets
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-11.