← all guides

Proxy authentication methods explained

proxies web-scraping proxy-authentication ip-whitelisting scraping-infrastructure

I found a working set of proxy credentials sitting in my own error tracker. Plain text, in the middle of a stack trace, on a hosted service, readable by everyone I had ever added to that project.

Nothing had been hacked. I had written the proxy as a single url with the password inside it, the way every example on every docs page writes it, and the HTTP client had logged the request it was about to make.

That is one half of this decision. The other half is the afternoon I lost to a scraper that had run untouched for eleven months and then refused every connection, because my home connection had picked up a new address overnight and the whitelist had never been told.

Both of those are authentication choices. Neither of them looks like an authentication choice at the point you make it.

The two methods

Username and password. Credentials ride on the proxy connection and the seller checks them. Standard HTTP proxy auth, supported everywhere, one line in whatever client you already use.

Address whitelisting. You register the source addresses allowed to use the proxy, and traffic from those addresses passes with no credentials at all.

Some sellers run both at once. Some make you pick one. A few support one only and never say which until you are already inside the dashboard, so it is worth asking before payment.

While you are asking, check the protocol too. It is common to find credentials working on the HTTP endpoint and whitelisting only on the SOCKS endpoint of the same account. Same host, one port apart, and the rules do not match. There is no warning anywhere.

Whitelisting is a bet that your address holds still

The appeal is real. There is no secret to store, so there is no secret to leak. Nothing lands in a config file, a repository, or a log line.

The bet underneath it is that the address you send from tomorrow is the address you registered today.

One box in one data centre on one static address: that bet is good for years, and I would take it. I run a whitelisted box myself.

Everywhere else it is quietly false. A laptop moves between the office, a cafe and a flat, and every network is another dashboard visit. A residential connection holds its address on a lease that renews, mine every few weeks, usually after the router reboots. An autoscaling worker gets a brand new address on every instance, so the pool that ran last night tells you nothing about the pool running tonight. A VPN or a corporate network presents whatever exit it feels like that morning.

Same shape every time. A static answer to a question your infrastructure keeps answering differently.

The failure lands in the wrong place

An address that is missing from the list gives the seller nothing to reject. No credential was presented, so no credential gets refused, and there is no natural status code to hand back.

What comes back is a dropped connection, a reset, or a timeout, depending on the seller. Your client reports a transport failure. That is the identical report you would get from a dead proxy, a broken route, or DNS falling over on the box.

So the debugging goes to the wrong systems. Status page, target site, local network, restart everything twice. Meanwhile the fault is one row in a dashboard nobody has opened since March.

The shortcut is comparing machines. Identical config, runs from your desktop, dies on the server. The code is fine. Go and read the list.

The list is shorter than you would like

Slots are capped. Five is normal, ten if you ask nicely. That is the ceiling on how many machines can reach the proxy without credentials, and it sits far below any worker pool.

Ranges are usually refused. Single addresses only. A cloud network whose address shape you do not control cannot be described to them at all.

Propagation is not instant either. A new entry takes a minute or two to reach whatever does the checking. I once added an address, watched it fail, added it a second time, and worked out much later that the first attempt had been correct and I was three minutes early.

Credentials go where the code goes

This is the whole argument for the other side and it is a strong one.

The same container deploys to three regions and authenticates from all three. A worker starts at 4am with nobody awake to edit anything, and it just works. Move a job to a different host and there is nothing to update anywhere.

The cost is that a secret travelling with the code ends up in the code. I have committed proxy credentials to a repository more than once. One of those repositories was public for about two days.

The url is where it actually leaks

Committing them is the obvious risk. The url is the one people miss, and it is a separate mistake.

Writing the proxy as http://user:pass@host:port is the compact form and every library accepts it. Then:

  • the HTTP client logs the request it is about to make, proxy included
  • the retry wrapper logs the connection that failed
  • the error reporter attaches the whole config object and ships it to a third party
  • a debug flag you turned on last week prints it a few thousand times an hour

None of that is a leak in the security sense. Every one of those lines is printing a url, and you put the password in the url.

Keep the host in one variable and the credentials in another, then join them where the client gets constructed. Dull fix, about ten minutes of work, and it takes the password out of every log surface at once.

Rotation is a different operation

Authentication decides whether you may use the proxy. Rotation decides which address you come out of. Two systems answering two questions, with no necessary relationship between them.

Sellers merged them anyway, because your connection string was already in the code and it made a convenient place to hang configuration.

Which is why your username is a config field

Look at how a sticky address gets requested. No API call. No setting to tick. Extra text appended to the username.

So the username stops being an identity and turns into a small configuration language, living in the one field HTTP proxy auth reserved for identity.

The practical consequence arrives late. A typo in that string does not fail as a bad username. It fails as an unrecognised option, or it drops back silently to whatever the default behaviour is, and you have two faults presenting one symptom.

It also welds routing to credentials. Changing how addresses get assigned means editing the same string that authenticates you. Read the seller’s format documentation before any of it reaches your code, because the format is theirs, it is not standardised, and the last seller’s format will not match.

Where I come down

Whitelisting suits exactly one situation. A single static server that will not move, owned by somebody who will remember to update the list on the day it does.

For anything else it is a trap, and the reason is the silence. It breaks at the worst possible moment, it points you at the wrong system, and hours spent on the wrong system are hours the job is not running.

If a seller supports whitelisting only and you are running anything that moves, buy elsewhere. I have walked away from cheaper lines over exactly this.

My default on anything new: credentials in one environment variable, host in another, joined at client construction, whitelisting switched off entirely so there is no second path in that can stop working while I am looking somewhere else.

What I cannot tell you

Whitelisting wins on revocation and I cannot argue that away. Change a password and every job using it stops at the same instant, including the four you had forgotten were using it. Delete a whitelist row and one machine loses access, and the list itself is a readable record of who has access.

There is a middle option I use more often than I let on. Whitelist one small box, put the credentials on that box alone, route everything else through it. One address to maintain, and the secret never leaves hardware you control. It costs you a hop, a few dollars a month, and one more service that can die at 3am.

I also do not know how any specific seller behaves with both methods enabled together. Some treat the whitelist as an override and admit credentialed traffic from anywhere. Some want both to pass. I have seen it go both ways inside the same week, there is no convention, so test it on your own account rather than assuming.

The checklist I run before wiring auth into a new scraper, and the settings that keep credentials out of your logs, are here.

Get new guides and videos first — join the Telegram channel.

proxies
Need proxies that survive the block wall?

Singapore Mobile Proxy runs real 4G/5G mobile IPs on rotating SIMs — the carrier-grade addresses most of these targets still trust.

see plans →
read on
More scraping guides

The rest of the field manual: target-site playbooks, library walkthroughs, provider reviews, and anti-bot troubleshooting.

browse all guides →