← back to blog

How to Route Scrapy Through a Mobile Proxy (Middleware Setup) 2026

mobile proxy scrapy web scraping python singapore mobile proxy

How to Route Scrapy Through a Mobile Proxy (Middleware Setup) 2026

Your scrapy spider runs clean for the first couple hundred requests, then every response comes back as the same captcha page and the data stops. You didn’t change the code. What changed is that the site finally noticed every request left from one datacenter ip and slammed the door. The fix isn’t a bigger retry count. It’s a small downloader middleware that sends each request out through a real mobile ip instead.

I run a hardware mobile proxy farm in Singapore. Real phones, real sim cards from Singtel, M1, and StarHub. I point scrapy at these proxies for actual scraping jobs, so I know where the middleware breaks and what the error messages really mean. Here’s the setup that holds up in production, and the traps that quietly kill a crawl.

Why scrapy’s built-in proxy support isn’t enough

Scrapy ships with an http proxy middleware that reads a proxy value from request meta, but it does not do authentication. This trips up almost everyone. If your proxy needs a username and password, and most mobile proxies do, scrapy will not send them. The request goes out, the proxy answers with a 407 proxy authentication required, and scrapy treats it like any other failed response. You get empty pages and no clear reason.

So you write your own middleware. It’s short. The whole job is to attach the proxy endpoint and the auth header that scrapy refuses to build for you.

Write the middleware

In your middlewares.py you make a class with a process_request method, and on every request you set the proxy and the authorization header. You take the string user:pass, base64 encode it, and set Proxy-Authorization to the word Basic, a space, and that encoded value. That header is the four lines of real work scrapy leaves to you.

import base64

class MobileProxyMiddleware:
    def __init__(self):
        self.proxy = "http://HOST:PORT"
        creds = "USER:PASS".encode("utf-8")
        self.token = base64.b64encode(creds).decode("utf-8")

    def process_request(self, request, spider):
        request.meta["proxy"] = self.proxy
        request.headers["Proxy-Authorization"] = "Basic " + self.token
        return None

Store the proxy url and the encoded credentials once in __init__, so you’re not re-encoding on every request. In process_request you set request.meta["proxy"] to the url and request.headers["Proxy-Authorization"] to Basic plus the stored token, then return None so the request keeps flowing down the chain. That’s the entire core. Everything else, retries, throttle, the ip check, lives around it.

Wire it in at the right order

In settings.py you add your class to DOWNLOADER_MIDDLEWARES with a number, and the number matters more than people expect. The built-in http proxy middleware sits at 750. You want your class to run before the retry logic and around the proxy stage, so a value like 740 works well.

DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.MobileProxyMiddleware": 740,
}

Set it too high and the retry middleware reacts to a failure before your auth header is even attached. Ordering is the single most common reason a correct middleware still fails.

Verify the exit ip before you trust it

Guessing is how you waste an afternoon, so verify before you scale anything. Write a throwaway spider with one start url pointed at an ip echo endpoint that just returns the ip that connected. Run it through the middleware and print the body. If the ip you see is your mobile exit ip, you’re tunneling. If it’s your own office or server ip, your meta["proxy"] key never took, usually because another middleware overwrote it. Check the order first.

Log the exit ip on the first request of every run so your logs prove which ip the job used. That’s gold when a client asks why something got blocked.

Handle the failures on purpose

A 407 means the auth header is wrong or missing, so decode your base64 and check for a stray newline, which is the classic bug. A connection reset mid-stream usually means the mobile tower handed you a new ip and the old tunnel dropped, so catch it and retry rather than crash.

Then there’s the soft block, and it’s the sneaky one. A captcha or a “please verify” page often comes back as a normal 200 status, so a status-only retry rule never fires. You have to look at the body. Write a retry condition that scans the response text for the captcha marker and forces a retry through a fresh ip when it sees one.

def process_response(self, request, response, spider):
    if b"captcha" in response.body.lower():
        new_request = request.copy()
        new_request.dont_filter = True
        return new_request
    return response

Pace one ip with autothrottle

Concurrency is where mobile proxies differ from a big datacenter pool. If you’re on one sticky session, that’s one phone, one ip, one real radio, so don’t blast it with 32 concurrent requests. Set concurrency to something modest, four or eight, and lean on autothrottle.

CONCURRENT_REQUESTS = 8
CONCURRENT_REQUESTS_PER_DOMAIN = 8
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 2.0
RETRY_TIMES = 4

Turn autothrottle on, set a start delay of one second and a target concurrency around two, and let scrapy back off when the site slows down. That pacing looks like a person on a phone, which is the whole point of using a mobile ip. If you’re pointed at a rotating endpoint instead, where each connection can get a different exit, you can run hotter, but you give up the steady session some sites want.

Scrapy keeps connections alive, and if you point a large crawl at one sticky endpoint, every alive connection rides the same radio. That’s fine until the radio drops and dozens of pooled connections all error at once. Keep concurrency low and the pool small. Slow and alive beats fast and dead. A job that finishes in two hours beats one that dies four times in twenty minutes.

Don’t leak dns

Dns is the leak nobody checks. By default your machine resolves the hostname locally, then connects through the proxy to that resolved address. So the site sees a mobile ip, but your real resolver, often your isp, saw the lookup. For plain http proxying this is usually fine because the connect target is a name, but if you ever switch to a socks layer, use socks5h. The h means resolve the name on the proxy side. The name lookup should happen where the traffic exits, not on your box. If the idea of a real carrier exit ip is new to you, here is what a mobile proxy actually is.

Match the rest of the request to the ip

The ip is only half the disguise. When you scrape from a real mobile ip, the site expects the rest of the request to look mobile too. If your scrapy user agent still says it’s a python client, or a desktop chrome on windows, the mismatch is glaring. A request from a Singtel mobile ip carrying a desktop user agent is a contradiction no real person produces.

Set a mobile user agent in your default headers, a recent chrome on android string, and let the cookies the site sets ride along on the same sticky session. Scrapy keeps cookies per spider by default, which is what you want when you’re holding one ip. The ip and the cookie jar should live and die together.

Pick sticky or rotating per site

There are two clean rotation models, and the right one depends entirely on the target. Sticky holds one ip for a whole crawl session, good for sites that tie a session cookie to an ip and get suspicious when the ip jumps mid-session. Rotating changes ip every request or every few minutes, good for wide shallow crawls where you touch each page once and never come back.

A logged-in target wants sticky. A price scrape across ten thousand product pages wants rotation. Scrapy lets you do either by what you put in that meta["proxy"] value: a fixed endpoint for sticky, a rotating gateway for rotation. If you want the longer reasoning on this with real jobs, I cover it in web scraping with mobile proxies across seven case studies.

Merge meta, never replace it

One trap with start_requests and your parse methods catches careful people. Those first requests flow through the middleware and get the proxy, good. But if you build a request somewhere and pass a custom meta, make sure you aren’t clobbering the proxy key by replacing meta wholesale. Update the dict, don’t replace it.

I’ve watched a crawl send its first page through the proxy and every follow-up link straight from the local ip, because a parse method rebuilt meta from scratch and dropped the proxy value. Always merge.

Where the proxy fits, and where it doesn’t

The middleware only touches the request and response layer. Your spider parse methods, your item pipelines, your feed exports, none of that changes when you add a proxy. That’s the nice part. You write the crawl logic once, and the proxy is a transport detail underneath.

So if a crawl works against a local copy of a page and fails live, the bug is almost always in the transport, the proxy, the auth header, or the block detection, not in your parsing. Knowing that split saves you from rewriting selectors when the real problem is a 407. And before you scale, test against rate limits: run the spider at low concurrency against fifty urls, watch the block rate, then climb. If blocks appear at concurrency eight but not at four, you found your ceiling for that site on that ip type. Write it down, because every site has a different tolerance and the only way to know is to measure it on the real exit ip you’ll use in production.

Put it on real Singapore mobile ips

If you want this to just work, the proxies under it have to be real. This is what I run: real Singapore mobile ips from actual Singtel, M1, and StarHub sims, sticky sessions you can hold for a full crawl, and dns that routes through the tunnel so you stop leaking your real resolver while you scrape. Point your scrapy middleware at it and the captcha pages stop.

A tiny middleware that injects the proxy and the auth header, the right ordering above retry, autothrottle so you don’t burn one ip, a body-based block check, and dns resolved at the exit. Get those five right and your spider runs for hours instead of dying at request two hundred. Start a free trial of Singapore Mobile Proxy, use the code YT30, and paste your host and port into the middleware above.

Get new guides and videos first — join the Telegram channel.

ready to try Singapore mobile proxies?

24-hour free trial. no credit card required.

start free trial
message me on telegram