Crawlee Review 2026: Honest Pros, Cons and Pricing
first, a correction to the premise most people arrive with. crawlee is not a proxy provider. it is an open-source web scraping and browser automation library, maintained by Apify, that happens to have very good proxy handling built in. if you searched for “crawlee proxies” hoping to find an IP pool you can buy by the gigabyte, that product does not exist under this name. what exists is the tooling that decides how your requests use whatever proxies you bring.
i am reviewing it anyway because i get asked about it constantly by people building scrapers out of Singapore and elsewhere, and because the question “is crawlee any good for proxy-heavy work” has a real answer. i have run it against mobile, residential and datacenter pools, and the library is the part of my stack i touch least often because it mostly stays out of the way.
the headline verdict: crawlee is one of the best free ways to drive proxies at scale, and a poor choice if you want someone else to do the work. i score it 4 out of 5 as a framework, with the caveat that several of my evaluation axes (pool size, geo coverage, price per GB) simply do not apply to it. i will say where that happens rather than invent numbers.
what Crawlee actually does
crawlee is published at crawlee.dev with source on GitHub under apify/crawlee. the main version is Node.js and TypeScript, and there is a separate Python port. it gives you crawler classes that share one interface:
- CheerioCrawler and HttpCrawler: plain HTTP fetching with HTML parsing, fast and cheap on bandwidth
- PlaywrightCrawler and PuppeteerCrawler: full headless browsers for pages that need JavaScript
- a request queue, so a crawl can be paused, resumed and deduplicated
- a dataset store for results
the proxy part is where it matters for this site. the proxy management guide documents a ProxyConfiguration class. you hand it a list of proxy URLs, or a function that returns one, and the crawler rotates through them. you can also define tiers, so the crawler starts on cheap datacenter IPs and only escalates to residential or mobile when a site begins returning blocks. that single feature saves real money, which i will come back to.
alongside that sits a session pool. a session holds cookies and is bound to one proxy IP, so a “user” keeps the same exit address for its lifetime. when a session gets a 403 or a captcha, crawlee can retire it and start a fresh one on a new IP. this is the closest thing the library has to a sticky session control, and it works at the application layer rather than at the proxy gateway.
the third piece is autoscaling. crawlee watches CPU, memory and event loop lag and adjusts how many requests run in parallel, between a minimum and maximum you set. you can read how that behaves in the crawlee documentation, and it is worth understanding before you point it at a metered residential plan.
pricing
crawlee costs nothing. the library is Apache 2.0 licensed, so there is no licence fee, no seat price and no per-GB charge from crawlee itself. that is the real 2026 number and it has not moved.
the money goes to three other places:
- your proxy vendor: this is almost always the biggest line. rates vary hugely by type and vendor, and i am not going to quote a per-GB figure for a product crawlee does not sell. check your provider’s current price page before you plan a budget. in my own work, bandwidth-metered residential and mobile traffic costs far more per GB than datacenter, which is the whole argument for tiered proxies.
- compute: a headless Chromium crawl eats RAM. a small VPS handles a few concurrent browsers, and a heavier crawl needs a bigger box.
- optional Apify platform: crawlee is built to deploy onto the Apify platform, which sells hosting and its own proxy network. the Apify proxy documentation describes the datacenter, residential and Google SERP groups it offers. i have not re-verified its current rates for this review, so look at Apify’s pricing page directly rather than trust a number from me. you do not have to use it. crawlee runs fine on your own machine with a third-party proxy.
so the pricing per GB axis is not scorable here. what i can say is that crawlee reduces your effective cost per successful page by being careful about when it uses expensive IPs and by not retrying blindly.
what works
- one proxy abstraction across every crawler type. i swapped a CheerioCrawler job to PlaywrightCrawler for a site that moved to client-side rendering and the proxy configuration code did not change. that is the kind of boring consistency that saves an afternoon.
- tiered proxy escalation. on a retail catalogue crawl i started on datacenter IPs and let blocked requests climb to residential. most pages never left the cheap tier, so metered bandwidth was only spent where it was needed.
- session persistence that actually pins an IP. because the session pool ties cookies to a proxy URL, a login flow stays on one address. for mobile proxies with sticky ports this lines up nicely, since you give each session its own port URL.
- autoscaled concurrency. rather than guessing a thread count, i set a floor and a ceiling and let it find the level. when a target started throttling, the pool backed off instead of burning through IPs. concurrent connections end up bounded by your proxy plan’s limit, which you set as the ceiling.
- provider neutrality. it takes any HTTP, HTTPS or SOCKS-style proxy URL your vendor gives you, so you are not locked in. for a mobile setup i have pointed it at ports from singaporemobileproxy.com with no special adapter.
what doesn’t
- it is not a proxy network. there is no IP pool, no geo targeting and no connection success rate to measure on crawlee’s side. anyone who tells you otherwise has misread the project. every one of those axes depends entirely on the vendor you pick.
- you pay twice in effort. the library is free, but you write and maintain code. a stale selector or a changed page still breaks your job, and nobody emails you about it. if you want no-code, this is the wrong tool.
- rotation control is code, not a dashboard. you can do per-request rotation, per-session rotation or tiered logic, but you build that logic yourself. vendors with a control panel and API-based rotation give you a button. crawlee gives you a class and a docs page.
- support is community-based. issues go to GitHub and the Discord, and the maintainers are responsive in my experience, but there is no SLA. if a production crawl fails at 3am, you are the support team.
- it does not beat anti-bot systems by itself. crawlee ships fingerprint generation for its browser crawlers, and it helps, but a hard target that scores IP reputation will still reject a flagged datacenter address. the library cannot fix a poor pool. if you run into this on account-heavy work, the browser identity side is covered over at antidetectreview.org, which is a different problem from request rotation.
who should buy
crawlee is free, so “buy” here means adopt. these are the profiles where i would start with it:
- a developer already working in TypeScript or Python who needs a scraper that can survive blocks, and who is happy to bring a proxy plan from any vendor.
- a data team scraping product, price or listing pages at moderate to high volume, where tiered proxies can keep the residential bill down.
- an operator running a scraping business out of Singapore or Southeast Asia who wants local mobile or ISP exits and needs the framework to respect sticky sessions per port.
- someone who wants to start free and grow. you can prototype on a laptop with no proxies and add them later without a rewrite.
who should skip
- anyone who wants a finished proxy product with a dashboard and an invoice per GB. you want a vendor, not a library.
- non-developers. if you do not want to read a docs site and debug a stack trace, a hosted scraper or a no-code tool will serve you better.
- people hoping a library will make a heavily protected site easy. the hard part there is IP quality and fingerprint hygiene, and crawlee only helps with the plumbing.
- very small one-off jobs. for a single page fetch, plain curl or a simple HTTP client is quicker than setting up a crawler project.
one more note, because it comes up. check a site’s terms and its robots.txt before you crawl, and know that rules about scraping differ by country and by data type. this is not legal advice. the HTTP semantics specification, RFC 9110 is a useful reference for what status codes like 403 and 429 are supposed to mean, which helps when you are tuning retry behaviour politely.
alternatives to consider
- Scrapy: a mature Python framework with a large ecosystem and proxy middleware, a good pick if your team is Python-first and does not need headless browsers by default.
- Apify platform with its own proxy: if you like crawlee but want hosting and a proxy pool from one place, this is the same company’s managed route, at a price you should check yourself.
- a plain proxy vendor with a rotating gateway: if you only need IPs and have your own scraper, buy from a provider directly. our best residential proxies for scraping roundup compares vendors on the axes crawlee cannot cover, and the full index lives in the blog.
if your scraping runs alongside multi-account work, the operational side is covered at multiaccountops.com, and for phone-based setups there is cloudf.one.
verdict
crawlee is a very good free framework and a bad answer to “which proxy should i buy”. use it as the engine, choose a proxy vendor on pool size, geo coverage and success rate, and let tiered rotation keep the metered traffic to a minimum. i rate it 4 out of 5, and the missing point is only because it needs code and a separate proxy bill.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-06.