ivinco
cURL vs. Playwright vs. LLM Scraper: A Decision Tree for 2026

cURL vs. Playwright vs. LLM Scraper: A Decision Tree for 2026

Ivinco Team·

We build scraping pipelines on Kubernetes for teams running 50K–5M pages a day. The first question we ask in every audit is which tool layer the target actually needs. Most teams answer wrong.

Take a common job: extract product titles and prices from 10,000 pages a day.

With curl_cffi and TLS impersonation, the fetch layer costs under $5/month in compute. With Playwright on Kubernetes, the same job runs $80–$150/month. With an LLM-native scraper like Stagehand extracting structured data on every page, the bill is $1,500–$6,000/month in LLM fees alone before infrastructure.

Same output. Three orders of magnitude cost spread. That's the Rendering Ladder: each rung up the stack costs 10–50x more than the one below it, and the anti-bot difficulty of the target is what forces you up.

We've seen the same failure pattern in every stack audit. Playwright gets picked because every tutorial uses it. LLM scrapers get picked because they're the newest thing. The AWS bill doubles, then doubles again, before anyone checks whether the data was already sitting in the initial HTML response. The right question isn't which tool is best. It's which rung of the ladder the target actually requires.

The Three Rungs of the Rendering Ladder

Rung 1 — HTTP client with TLS impersonation

curl_cffi, curl-impersonate, tls-client, httpx, aiohttp. No JavaScript engine. No DOM. Just raw HTTP requests with the right TLS fingerprint on the wire.

The critical capability isn't the HTTP library itself — requests and aiohttp have shipped that for years. It's TLS impersonation. Modern anti-bot systems look at the TLS ClientHello packet — cipher order, extensions, elliptic curves — and build a JA3 or JA4 fingerprint. Python's default TLS stack has a fingerprint that screams "not a real browser," and Cloudflare blocks it before your HTTP request is ever parsed.

curl_cffi is a Python wrapper around BoringSSL-patched curl that impersonates 40+ browser versions at the TLS layer — Chrome 99 through 146, Safari 153 through 260, Firefox, Edge, Tor. The usage is one line:

import curl_cffi
r = curl_cffi.get("https://example.com", impersonate="chrome")

That single impersonate="chrome" argument changes the JA3 fingerprint, the HTTP/2 frame settings, and the header ordering to match real Chrome. Against sites protected by Cloudflare Free tier or basic JA3 checks, this is often enough.

Per-request cost: 5–10 MB memory, milliseconds of CPU, one TCP round-trip. You can run 10,000 concurrent curl_cffi requests on a single 4 GB VM.

Fails against: JavaScript-rendered content (React, Vue, Next.js SPAs where the data is in a fetch call after page load), Cloudflare Managed Challenge (the 5-second JS execution check), DataDome behavioral analysis, and anywhere CAPTCHA triggers.

Rung 2 — Headless browser with CDP

Playwright, Puppeteer, Selenium. A real Chromium (or Firefox, WebKit) process that executes JavaScript, runs the target's frontend framework, and fires the same XHR/fetch calls the user's browser would.

The cost is steep. A single headless Chromium instance consumes 60–130 MB of memory on typical pages, 150–400 MB on heavy SPAs. Cold start takes 500ms–2s. Run 100 concurrent browsers on Kubernetes and you're already at 10 GB of RAM plus the "starved pods" failure mode detailed in our Playwright cluster comparison.

But the browser rung is what you need when:

  • The data is rendered client-side (check the HTML response — if the value isn't in there, it's a JS render)
  • You need to execute the page's anti-bot challenge (Cloudflare's 5-second check, Turnstile)
  • The session requires cookies set by JavaScript
  • You need to click through pagination or fill forms

Stealth-patched variants — patchright for Python Playwright, rebrowser-puppeteer for Node — close the Runtime.Enable CDP leak that vanilla Playwright/Puppeteer still carry. These get you past Cloudflare Business tier on most targets. Against Cloudflare Enterprise or DataDome, you need residential proxies and behavioral mimicry on top.

Per-page cost (infrastructure only): $0.005–$0.02 per page at scale, mostly CPU and memory. Per-page cost with managed browser API (Browserless, Bright Data Scraping Browser): $0.10–$1.00 per page.

Rung 3 — LLM-native scraper

Stagehand, Browser Use, Firecrawl, Crawl4AI, ScrapeGraphAI, skyvern, AgentQL. Two different architectures hide under the same marketing:

Pipeline LLM scrapers — Firecrawl, Crawl4AI. Fetch the page with a headless browser, convert it to clean Markdown, optionally pass through an LLM for structured extraction. No agent loop. The LLM is called once per page for the extraction step (or not at all if you just want Markdown).

Agent LLM scrapers — Stagehand, Browser Use, skyvern. An LLM agent drives the browser. You write natural language instructions — "click the login button, fill the email field, extract the price from the product card" — and the agent translates them to CDP commands by looking at the DOM (or screenshots, in skyvern's vision-first approach). Stagehand's three primitives — act(), extract(), observe() — make this explicit.

Pipeline scrapers cost browser-plus-a-little-extra-LLM. Agent scrapers cost browser-plus-an-LLM-call-per-action. The WebVoyager benchmark from He et al. (2024) is the shared evaluation task these tools report against — real-world web tasks across 15 major sites. Vendor-reported scores on comparable WebVoyager-style evals cluster around 72–78% task completion for LLM agents (Browser Use, Stagehand) versus roughly 98% for hand-written Playwright scripts against the same targets, per 2025 independent scraper benchmarks collated by NxCode.

The gap is the LLM's misinterpretation rate. Roughly one in four runs, the agent clicks the wrong button or extracts a plausible-but-incorrect value. Acceptable for one-off research agents. A production outage for structured data pipelines.

Per-page cost: $0.02–$0.30 for pipeline LLM scrapers (Firecrawl pricing at volume), $0.05–$2.00 for agent LLM scrapers depending on action count and model. Stagehand's documentation claims smart caching reduces LLM costs by up to 90% on repeated workflows by memoizing action sequences — but the cache hits only when the page structure actually repeats.

The Decision Tree

Run these checks in order. Stop at the first one that applies.

Step 1: Is the data in the initial HTML response?

curl -s "https://target.com/product/123" | grep -o 'data-price="[^"]*"' | head

If the value is there, you're on Rung 1. Everything above it is waste. The curl_cffi version of your scraper runs 100x cheaper and 10x faster than the Playwright version.

Counter-check: load the page in Chrome, open DevTools Network tab, reload, and filter to "XHR/Fetch." If the prices appear in responses here but not in the initial document, the site is rendering client-side. Move to Step 2.

Step 2: Does TLS fingerprinting block curl_cffi?

If the target is protected by Cloudflare, DataDome, Akamai, or PerimeterX, run a quick test: curl_cffi with impersonate="chrome" plus a residential proxy. If you get the HTML, you're done — Rung 1 wins. If you get a challenge page, a 403, or a JavaScript redirect, the target needs browser execution. Move to Step 3.

A sign that Rung 1 is failing even though the response looks okay: your scraper works for the first 10 requests and then rate-limits for an hour. That's usually Cloudflare's rate-limit rule triggered by a non-browser fingerprint — the response was technically successful but flagged you for follow-up.

Step 3: Is the page structure stable across runs?

This is the fork that decides between Rung 2 (headless browser with selectors) and Rung 3 (LLM scraper).

Stable structure = Playwright. Write selectors once, run them a million times. Per-page cost stays at a few cents. The maintenance tax is real: independent benchmarks (NxCode's 2025 scraper tracking) put 15–25% of Playwright selectors needing fixes over 30 days on actively-developed targets, versus under 5% prompt adjustments for LLM-driven scripts. On any single well-chosen target, paying the selector tax still beats paying the per-page LLM bill.

Unstable or heterogeneous structure is the LLM scraper's home turf — thousands of targets, each with its own layout. Ecommerce sites that A/B test their DOM weekly. Form flows with conditional fields. News aggregation across 500 publishers with different HTML.

The dividing line: if you can afford to write and maintain selectors for every target in your corpus, do it. If you can't — because the target count is too high or the layout changes too fast — pay the LLM tax.

Step 4: Pipeline LLM or agent LLM?

If you picked Rung 3, the last fork is architecture.

Firecrawl or Crawl4AI — the job is "fetch these URLs and give me clean Markdown or structured JSON." No decisions on the page. The LLM handles extraction, never navigation.

Stagehand or Browser Use is a different architecture. The LLM drives the browser through clicks, form fills, and multi-step workflows. Cost per run climbs. Completion rate drops. Agent scrapers are research tools wearing production clothes. A 22% error rate is fine when a human reviews every run. It's a catastrophe at 3 a.m.

The Cost Math at Volume

The Rendering Ladder priced for a 10,000-page/day product scraping job (structured output, US targets, published vendor pricing as of April 2026):

| Rung | Tool | Monthly cost (10K/day) | Maintenance tax | |------|------|------------------------|-----------------| | 1 | curl_cffi + residential proxy | $5 compute + $30–$100 proxy | Low (HTTP stable) | | 2 | Playwright on K8s | $80–$150 compute + $100–$300 proxy | 15–25%/month selector fixes | | 2-managed | Browserless / Bright Data Scraping Browser | $300–$3,000 | None (managed) | | 3-pipeline | Firecrawl | $500–$2,000 | Near-zero | | 3-agent | Stagehand + Claude Sonnet 4.6 | $1,500–$6,000 LLM only | under 5%/month |

Compute prices assume commodity cloud (AWS/Hetzner) and exclude proxy costs. Proxy costs assume residential rotation; datacenter proxies cut that by 5–10x but fail against Cloudflare more often. The LLM agent number comes from NxCode's published range of $0.005–$0.02 per extract() call × 10K × 30 days, matching Apify's own assessment that LLM scraping "quickly becomes expensive" at any meaningful volume.

The numbers shift with volume. At 1,000 pages/day, LLM scrapers are cheap enough that the reliability delta (under 5% maintenance vs 15–25%) can justify the premium outright. At 1,000,000 pages/day, curl_cffi with aggressive proxy rotation beats any managed option by 10–100x.

Per-page cost isn't the critical number. Engineering break-even is. Selector maintenance on a Playwright pipeline with 50 active targets costs roughly one engineer-week per month at US/EU senior rates. If an LLM scraper eliminates that work on a stable 10,000-page/day job, the $2,000–$6,000 monthly LLM bill may still be cheaper than the engineering time it replaces.

The Hybrid Pattern Teams Actually Ship

Pure-Rung systems are rare in production. Pipelines that survive contact with real targets route by rung:

  1. Rung 1 for the easy majority — curl_cffi hits APIs and simple HTML targets. Fast, cheap, stable.
  2. Rung 2 for Cloudflare-protected sites — Playwright with patchright and residential proxies, handling the sites that require JS execution.
  3. Rung 3 for the messy residual — LLM scraper wrapped around targets where layout varies or structure is ambiguous. Agent scrapers for workflows (login flows, search-and-extract). Firecrawl or Crawl4AI for training-data-style content extraction.

The time goes into orchestration. Route each URL to the right rung by domain, response code, or success history. Retry a failed Rung 1 request on Rung 2. Retry a failed Rung 2 request on Rung 3. Log the rung that succeeded so the router learns.

Where This Framework Breaks

Four failure modes catch teams repeatedly:

Silent Rung 1 failure. curl_cffi returns 200 OK with HTML that looks like the page. It isn't — it's Cloudflare's challenge page, rendered without the real content. The extractor runs against it, finds no prices, writes zero records. The pipeline logs success. The dashboard shows green. Someone notices the product count dropped three weeks later when a report breaks. Defense: validate every response against an expected-field assertion (price > 0 AND title != ""), not just HTTP status.

Rung 2 with wrong stealth. Vanilla Playwright or Puppeteer fails against Cloudflare Business. Patched variants (patchright for Python, rebrowser-puppeteer for Node) pass the same test. Teams deploy the vanilla version, spend weeks debugging, conclude "browsers don't work on this site," and escalate to Rung 3. The tool was fine. The configuration was wrong.

Rung 3 hallucinations. An LLM agent asked to "extract the product price" on a cart summary page can return $49.99 — the shipping line item — when the actual product price is $499.00. The agent is confident. The schema accepts the value. The pipeline writes a decimal-order-of-magnitude error into production data. Mitigation: structured output with Pydantic schemas, regex post-validation on numeric ranges, and cross-checks against sanity bounds per category. Stagehand and Browser Use both support schema enforcement. Use it.

Caching assumptions. Stagehand's 90% cost reduction from caching only materializes when workflows repeat. Hit every URL with a unique path (ecommerce with SKU-level URLs, search results, long-tail content), and the cache hit rate drops to near zero. Per-page cost stays at the uncached rate, and the pricing model you budgeted from the docs no longer describes your bill.


Need help choosing the right rung? Talk to an engineer — we design scraping pipelines that route each target to the cheapest tool that works.

What This Comparison Cannot Tell You

Whether your specific target will still work at Rung 1 next month. Cloudflare, DataDome, and PerimeterX update their detection continuously. The curl_cffi script that works today may need Playwright tomorrow — not because the tool got worse, but because the target upgraded from Cloudflare Free to Business.

The Rendering Ladder is a cost map, not a correctness map. A single target at a single hour can respond differently to different tools depending on its current fingerprint score, the time of day, or which CDN edge node you land on. Build for graceful fallback — Rung N fails, orchestrator routes to Rung N+1, monitoring fires on success-rate drops before the pipeline goes dark.

Pick the lowest rung that works today. Expect to climb it next quarter.

Frequently Asked Questions

When should I use cURL instead of Playwright for web scraping?

Use curl_cffi or a similar HTTP client with TLS impersonation whenever the data you need appears in the initial HTML response and the target doesn't require JavaScript execution. Check by running curl against the URL and grepping for the field. If it's there, HTTP wins — it's 10–100x cheaper than headless browsers at any scale above a few hundred requests per day.

Do LLM-powered scrapers like Stagehand replace Playwright?

No. LLM scrapers like Stagehand and Browser Use sit on top of Playwright, using it as the browser layer. They replace CSS selectors with natural-language instructions, trading reliability (98% task completion for hand-written Playwright vs 72–78% for LLM agents per WebVoyager) for lower selector maintenance. Best for heterogeneous targets or rapidly changing layouts, not high-volume stable pipelines.

What is TLS fingerprinting and why does it matter for scraping?

TLS fingerprinting analyzes the ClientHello packet sent during HTTPS handshake — cipher suite order, TLS extensions, and HTTP/2 settings — to identify the client. Python's default requests library produces a JA3/JA4 fingerprint that anti-bot systems instantly flag as non-browser. Tools like curl_cffi impersonate real Chrome or Firefox fingerprints at the TLS layer, bypassing this check without running a full browser.

How much does it cost to run an LLM-powered web scraper at scale?

At 10,000 extractions per day, an agent-based LLM scraper like Stagehand with Claude Sonnet 4.6 costs $1,500–$6,000/month in LLM fees alone. Pipeline LLM scrapers like Firecrawl run cheaper ($500–$2,000/month at the same volume). Compared to curl_cffi at $5–$100/month or Playwright at $80–$300/month, LLM scrapers are 10–50x more expensive per page.

When should I choose a managed web scraping API over building in-house?

Managed APIs like ScraperAPI, Bright Data, and ScrapFly make sense at low volume (under 50,000 pages/day) or when the team lacks scraping expertise. Above 100,000 pages/day, managed pricing often exceeds the fully-loaded cost of an in-house Playwright cluster on Kubernetes. The break-even point depends on target complexity — Cloudflare Enterprise targets justify managed APIs at any volume; simple HTML scraping rarely does.