
Retry Logic for Scrapers: Exponential Backoff, Idempotency, and Failure Recovery
We run scraping pipelines for teams processing 500K–10M URLs per day, and every new stack we audit has failed the same way at least once. A dev set retries=3 on every job. A target went down for an hour. The scraper hit the target three times per URL during the outage, concentrated at exactly the schedule's firing moments. Every URL in the queue fired three times. The IP pool burned. The target's anti-bot detected the retry pattern and perma-banned the fleet.
This is the Retry Amplifier: naive retry logic that turns a transient target failure into an outage for the scraper. The team added retries to make the system resilient. They made it brittle in the exact moments the target needed to be treated gently.
The retry logic itself isn't the hard part. The hard part is distinguishing which failures need retries, which need circuit breakers, and which need the request dropped entirely — and doing it at scale, without creating the amplifier that takes you down alongside the target.
Three Failures, Three Retry Strategies
The first mistake teams make is treating all failures the same. A 429 rate limit and a 500 server error both "failed" — but retrying them identically burns proxies and extends outages.
Transient failures — retry aggressively. Connection timeouts, DNS resolution failures, TCP reset, transient 5xx (502, 504 when intermittent). The target is reachable but a specific request didn't complete. Retry with short exponential backoff — 1s, 2s, 4s, 8s. Most transient failures either resolve within 3 attempts or reveal themselves as something worse.
For recoverable failures (429 Too Many Requests, 503 Service Unavailable with a Retry-After header, IP-specific blocks), the calculus changes: the target is actively asking for space. Retry with longer delays — 30s, 60s, 120s — and rotate the proxy if the block is IP-scoped. Honor Retry-After values literally. Overriding them with your own shorter backoff invites extended blocks.
Permanent failures — don't retry. 404 Not Found, 410 Gone, 403 Forbidden (CAPTCHA walls, permission denied), 401 Unauthorized, malformed responses that indicate a deprecated URL pattern. Retrying these classes wastes proxies and fills the queue with dead messages. Route them to the dead-letter queue for human review.
The error classifier lives in your HTTP middleware:
def classify_response(response, exception=None):
if exception:
if isinstance(exception, (ConnectionTimeout, DNSError, ConnectionReset)):
return RetryClass.TRANSIENT
return RetryClass.PERMANENT
if response.status_code == 429:
return RetryClass.RECOVERABLE
if response.status_code in (500, 502, 503, 504):
return RetryClass.TRANSIENT if response.status_code != 503 else RetryClass.RECOVERABLE
if response.status_code in (404, 410, 403, 401):
return RetryClass.PERMANENT
return RetryClass.PERMANENT # unknown is permanent — don't amplify
Default to "permanent" for unknown responses.
The alternative is amplifying every novel failure into a retry storm.
Exponential Backoff with Jitter
The canonical retry math: delay = 2^attempt * base_delay, capped at some maximum. Attempt 1 waits 2 seconds, attempt 2 waits 4, attempt 3 waits 8, attempt 4 waits 16. Both BullMQ and Celery implement this as their default backoff strategy.
Exponential alone is incomplete. If 1,000 scrapers all hit a 429 at the same moment, exponential backoff fires all 1,000 retries at exactly 2 seconds, 4 seconds, 8 seconds. The target sees a perfect retry spike — the thundering herd. The target's rate limiter responds by extending the block.
Jitter breaks the synchronization. Instead of retry at exactly 2^attempt * base_delay, retry at random(0, 2^attempt * base_delay). The 1,000 scrapers now spread their retries across the full backoff window. BullMQ implements this as the jitter option (0–1.0 value scaling the randomization range). AWS's architecture guidance calls this "full jitter" — retry at a uniform random point within the exponential envelope rather than exactly at the envelope boundary.
The minimum viable backoff policy:
import random
def compute_backoff(attempt: int, base_delay: float = 1.0, max_delay: float = 300.0) -> float:
"""Exponential backoff with full jitter. Returns delay in seconds."""
exponential = min(base_delay * (2 ** attempt), max_delay)
return random.uniform(0, exponential)
Cap the maximum delay at 5 minutes for recoverable failures, 30 seconds for transient failures. Longer backoffs mean URLs sit in retry limbo longer than they're worth.
Respect Retry-After headers when present. A target that returns Retry-After: 120 is giving you a literal instruction — wait 2 minutes. Overriding it with your own backoff policy invites longer blocks.
Idempotency: The Retry's Silent Partner
Retrying a request assumes the request is idempotent — running it twice produces the same result as running it once. GETs against read-only scraping targets pass this test trivially. Anything that writes state (login, cart, form submission) fails it.
Retries without idempotency don't fix failures. They multiply them.
Duplicate extraction writes. Scraper fetches URL, extracts data, writes to Postgres, then crashes before acknowledging the queue. Queue redelivers the message, new worker fetches the URL, extracts again, writes again — now there are two rows for one product. Fix: use idempotent upsert with a stable key (INSERT ... ON CONFLICT (url) DO UPDATE in Postgres, REPLACE INTO in MySQL, MERGE in Snowflake). The stable key is usually the target URL plus a timestamp bucket.
Duplicate cart/login actions. Less common in scraping, but scrapers that perform login or add-to-cart actions as part of a workflow need special handling. Retrying a "submit form" action submits twice. Fix: assign an idempotency key per workflow instance (UUID generated at workflow start), pass it as a form field or header, and have the target dedupe on it. If the target doesn't support idempotency keys, use exactly-once-ish semantics: check workflow state before the action, skip if already complete.
Duplicate invoicing from managed APIs. Most managed APIs bill per request. Naive retry logic at the scraper level plus retry logic inside the managed API client means your retry counter is multiplied by the API's internal retry counter. Bright Data Scraping Browser, ScrapFly, and Browserless all retry internally on certain failures. Your total retries become your_attempts × their_attempts — if both are 3, that's 9 billed requests per logical retry. Fix: disable retries in one layer, usually the managed API layer (set retries=0 or equivalent in the client), and manage retries yourself so the billing matches your logical attempts.
For queue-based scrapers, idempotency requires message acknowledgment discipline. Acknowledge the message after the write commits, not after the request succeeds. BullMQ's acks_late equivalent (moveToCompleted called from within the transaction that writes the result) and Celery's acks_late=True setting both do this. Without it, a worker crash between "request succeeded" and "result written" causes duplicate extractions.
Dead-Letter Queues: What to Stop Retrying
Every scraping queue needs a dead-letter target for messages that will never succeed. The failure mode without one: a poison message (malformed URL, deleted target page, unscrapeable CAPTCHA-protected endpoint) loops forever, consuming browser capacity, generating retry storms, and filling logs. Production pipelines typically have 0.5–3% of messages end up in the dead-letter queue; the absolute number grows with target count.
Three routing rules cover almost every scraping queue:
- Permanent failures (404, 410, 403, 401 classified above) go to the DLQ immediately. No retries.
- Recoverable failures exceeding max attempts go to the DLQ after 3–5 attempts. The message may genuinely succeed later, but blocking the queue on it is worse than reviewing it manually.
- Validation failures on extracted data (covered in our extraction post) go to a separate validation DLQ — the scraper succeeded, but the output failed schema checks. Often a signal of target layout change, not a transient error.
DLQ review is part of scraping operations, not an edge case. Wire a daily job that counts DLQ depth by target and alerts when any single target's DLQ grows above a threshold. Rising DLQ volume usually maps to one of three causes: the target changed its layout (extraction fix needed), the target banned the proxy fleet (rotation fix needed), or the target is genuinely down (no fix — move to lower-priority retry schedule).
BullMQ exposes this as the failed set; Celery uses a separate queue with acks_late + custom error handler; Temporal's failed workflow executions serve the same role. Scrapy's retry middleware dead-letters to a separate file/database rather than a queue, which is simpler but less observable.
Circuit Breakers: Stop Retrying When a Target Is Down
The Retry Amplifier's worst manifestation is during target outages. If a target is down for 30 minutes, every URL in the queue retries 3–5 times during that window — tripling or quintupling your request volume against a target that's already failing.
Circuit breakers prevent this. Track per-target success rate over a sliding window (last 100 requests or last 60 seconds). When success rate drops below a threshold (20%), open the circuit: reject new requests to that target immediately, without actually attempting them. After a cooldown (60–120 seconds), allow a single probe request — if it succeeds, close the circuit and resume; if it fails, extend the cooldown.
Python implementations: pybreaker, circuit-breaker-py, or roll your own in Redis with per-target counters. Node.js teams use opossum. For scraping, per-target granularity is non-negotiable — a global circuit breaker that opens when any one target is down is almost always wrong. Track state per-domain.
Per-target breakers let a cluster shed load on the failing target while continuing to serve every other target normally. Without them, one bad domain silently drags down the whole fleet.
The Tools
The retry + queue layer lives in one of four places for production scrapers:
BullMQ (Node.js). Redis-backed, native attempts and backoff options per job, built-in rate limiting, straightforward DLQ via the failed set. Standard in Node.js scraping stacks; what Crawlee uses under the hood. Documentation is thorough.
Celery (Python). Distributed task queue with Redis or RabbitMQ broker, autoretry_for decorator, retry_backoff with jitter, acks_late for exactly-once-ish semantics. The default Python choice.
Temporal. Durable workflow execution. Retries aren't a queue setting — they're a RetryPolicy on each activity, with InitialInterval, BackoffCoefficient, MaximumInterval, and MaximumAttempts parameters. Overkill for simple URL scraping; correct choice for multi-step workflows (login, search, extract) where each step must be retryable independently. Expensive to adopt; powerful once adopted.
Scrapy middleware (Python). scrapy.downloadermiddlewares.retry.RetryMiddleware handles retries inline with the crawler. Less flexible than Celery or BullMQ but simpler — everything stays in the Scrapy process, no external queue broker. Fine for pipelines under 1M pages/day with straightforward retry rules.
The choice doesn't matter as much as getting the error classification right. A scraper on Celery with bad retry rules is worse than a scraper on raw aiohttp with correct rules. Tools don't prevent the amplifier; rules do.
Need help with retry architecture? Talk to an engineer — we design retry, idempotency, and circuit-breaker patterns for scraping pipelines at the 500K–10M URL-per-day scale.
What This Framework Cannot Tell You
Whether your target's specific anti-bot system treats retries as signal. Cloudflare Bot Management and DataDome both factor retry patterns into their detection scoring — a client that retries with textbook exponential backoff and perfect jitter may still look more botlike than a human who reloads the page once and gives up. Randomize attempt counts per URL where it doesn't hurt completion: retry some URLs up to 5 times, others up to 2, based on priority and tolerance for failure.
Whether the target will lift your ban after you stop retrying. Some sites ban IPs on retry signatures; the ban persists even after you fix the retry logic. Proxy rotation helps; switching to residential proxies when datacenter IPs are banned helps; asking managed providers to cycle their pool helps. None of these is guaranteed.
The retry logic is only as good as the error classifier that feeds it.
Frequently Asked Questions
What is the best retry strategy for web scrapers?
Exponential backoff with jitter — delay = random(0, 2^attempt * base_delay), capped at 5 minutes — combined with per-error-class routing. Transient errors (timeouts, 5xx) retry aggressively with short backoffs. Recoverable errors (429, 503) retry with longer delays honoring Retry-After headers. Permanent errors (404, 403) don't retry — they go to a dead-letter queue.
Why do my scrapers make a banned site's outage worse?
Naive retry logic — fixed retries=3 on every request without error classification — amplifies target outages. During a 30-minute outage, every queued URL retries 3–5 times, tripling or quintupling your request volume against a target that's already failing. The fix is per-target circuit breakers: track success rate per domain, and reject new requests automatically when any target's rate drops below 20%.
How do I make web scraper retries idempotent?
Use upsert operations with stable keys when writing extracted data (INSERT ... ON CONFLICT DO UPDATE in Postgres, MERGE in Snowflake, or composite url + timestamp_bucket keys). Acknowledge queue messages after the write commits, not after the HTTP request succeeds — this is Celery's acks_late=True or BullMQ's equivalent pattern. Disable retries in managed API client layers to avoid multiplicative billing.
What's the difference between transient and permanent failures in scraping?
Transient failures (timeouts, DNS errors, 502/504) resolve quickly — retry with short exponential backoff. Recoverable failures (429 Too Many Requests, 503 with Retry-After) need longer backoffs and usually proxy rotation. Permanent failures (404, 410, 403, 401) indicate the URL is genuinely gone or access is denied — retrying wastes proxies. Route them to a dead-letter queue for human review.
How should I set max retry attempts for a web scraping job?
3 attempts for transient errors (covers network flakes without amplifying outages), 3–5 for recoverable errors with longer backoffs (handles slow rate-limiting recovery), 0 for permanent errors. Exceeding 5 attempts typically indicates a stuck job — the message either needs human review or target structure changed. Move to DLQ after max attempts rather than looping indefinitely.
Related posts

cURL vs. Playwright vs. LLM Scraper: A Decision Tree for 2026
Same scraping job costs $5, $150, or $6,000/month depending on the tool you pick. A 4-step decision tree for choosing between HTTP clients, headless browsers, and LLM-native scrapers in production.

Debugging Scrapers in Production: Logs, Screenshots, Video Replay, and Failure Forensics
Status codes aren't evidence. Screenshots, DOM snapshots, Playwright traces, and structured logs are. The Evidence Bundle pattern for turning scraper failures from hours of investigation into dashboard queries.

Playwright vs. Puppeteer vs. Selenium for Production Scraping: A 2026 Comparison
Most headless browser comparisons test on localhost. Production scraping at 100+ concurrent sessions reveals different winners — here's the data on memory, detection, and Kubernetes deployment.