
Scraping E-Commerce Data for AI: Architecture for Product, Pricing, and Review Extraction
A pricing intelligence team scrapes 1 million SKUs every hour to keep their dashboard fresh. Most of those SKUs change price once a week, not once an hour — so the scraper makes 24 million page loads per day that reproduce the same prices it fetched yesterday. The proxy and compute bill is real; the freshness value is mostly zero.
This is the central economic question in ecommerce scraping: how fresh does the data need to be at query time, and what is the cheapest crawl cadence that delivers that freshness? Over-scraping burns proxies and invites bans. Under-scraping means the dashboard lies. The Freshness Premium is what you pay for every page you fetch that didn't actually change.
The scraping pipelines we build for ecommerce clients span three data types — catalog, pricing, reviews — each with different volume, structure, and freshness requirements. We treat them as three pipelines, not one. Teams that try to scrape everything the same way get the economics of all three wrong.
Amazon's SP-API pricing change — the move from free API to $1,400/year subscription plus usage fees effective April 2026 — has pushed more teams toward scraping product pages directly. What used to be a "just use the API" decision is now a real build-vs-scrape calculus for anyone sourcing data from Amazon. And Amazon is only the highest-volume instance of a broader pattern: Walmart Connect, Target, Best Buy, and most mid-size retailers either don't expose public APIs at all or charge for them.
The Three Data Types
Ecommerce data splits into three layers with different scraping characteristics.
Catalog data. Product title, description, specs, images, categories, brand, SKU. Changes rarely — most products hold their catalog state for weeks or months between updates. Volume scales with SKU count (large marketplaces carry hundreds of millions of listings; mid-size retailers, hundreds of thousands; a single brand, often under 100K). Extraction is relatively stable because product pages follow templated layouts within a site.
Catalog pages ship Schema.org Product JSON-LD on most modern retailers (covered in our extraction post) — Shopify, Magento, WooCommerce, Salesforce Commerce Cloud, and BigCommerce all emit it by default. Running extruct against a heterogeneous US ecommerce corpus typically lifts half or more of the extraction load off custom selectors before you write a single line of per-target code.
Pricing data. Current price, list price, currency, discount flags, availability, stock quantity (when shown). Changes frequently — minutes for promotional items, hours for regular SKUs, days for stable categories. Volume is SKU count × crawl frequency. For a 1M SKU tracker scraping hourly, that's 24M page loads/day. Pricing extraction is usually embedded in the same JSON-LD as catalog, under offers.price.
Review data. Review text, rating, date, reviewer name, verified purchase flag, helpful votes. Growth-driven — new reviews add over time, old ones rarely change. Volume is total review count (which runs into the hundreds of millions on large marketplaces) plus net-new reviews per day. Reviews have also become a key input for ecommerce AI applications — product Q&A with retrieval, sentiment and feature extraction, recommendation tuning — which pushes demand for bulk review scraping independent of pricing needs.
Each layer has different crawl cadence, different storage, and different infrastructure cost profile. A pipeline architecture that treats them as one thing either over-scrapes the cheap layer or under-scrapes the valuable one.
The Freshness Premium
The cost structure of ecommerce scraping is dominated by how often you hit each page, not how cleverly you extract from it. A catalog scrape that runs once a week costs 168x less than one running hourly. The question is: what's the staleness cost of running weekly instead of hourly on your specific data?
For most catalog data, weekly is fine. Product titles and specs don't change often enough to justify hourly crawls. A customer looking at a slightly outdated description isn't a problem worth $10K/month to solve.
Pricing is the opposite. On a fast-moving category (electronics, fashion flash sales), hourly pricing data loses value within minutes. On a slow category (furniture, industrial supplies), daily is sufficient.
Adaptive cadence per SKU beats fixed cadence per crawl job. Sample a SKU every 5 minutes for an hour, measure the actual price change frequency, then set the production interval accordingly:
| Observed change frequency | Production crawl interval | |--------------------------|---------------------------| | Multiple times per hour | Every 5 minutes | | Multiple times per day | Every 30 minutes | | Daily | Every 6 hours | | Weekly | Daily | | Monthly or less | Weekly |
This reduces crawl volume on stable SKUs by 10–100x while maintaining the freshness guarantee on volatile ones. The adaptive cadence table needs to update monthly — categories that were stable in Q1 may be volatile in Q4 during holiday pricing.
SLA design: measure staleness at query time, not crawl frequency. A dashboard that shows prices scraped within the last hour is more useful than one that crawled every URL daily but shows 23-hour-old data. Instrument staleness per record: scraped_at timestamp on every price row, query-time warnings when staleness exceeds threshold, separate crawl priority for SKUs whose staleness is approaching the threshold.
The Architecture
Most scrapers treat catalog, pricing, and reviews as one pipeline with different cron intervals. That's the mistake.
Three concurrent pipelines, each tuned to its data type.
Catalog pipeline (weekly): Crawl the product list once per week. Parse JSON-LD for catalog fields. Write to a relational catalog table keyed on product URL or SKU. Detect new products, flag removed products, update changed fields. For 1M SKUs at 1MB per page, this is 1TB/week of bandwidth but runs during low-load windows.
Pricing pipeline (adaptive cadence): Per-SKU interval based on observed change frequency. Crawl the product page, parse the offer block, write a price history row with timestamp. Storage is time-series: TimescaleDB, ClickHouse, or a partitioned Postgres table. For a 1M SKU pricing tracker averaging 4 crawls/day per SKU, that's 4M price observations/day. At 10M rows/week retention, ClickHouse or TimescaleDB handles this trivially; raw Postgres starts struggling around 100M rows without partitioning.
Review pipeline (incremental): Crawl new reviews only — paginate from the newest backward, stop when you hit reviews already in the database. Don't rescrape entire review sets. Store reviews in a text-search-capable store: Postgres with tsvector columns, Elasticsearch, or OpenSearch. For LLM training data exports, denormalize reviews plus product context into Parquet files on S3 or GCS.
Queue architecture: separate queues per pipeline type. Catalog crawl can share worker capacity with pricing because it runs off-peak; reviews typically use different workers because the extraction logic is different (pagination-heavy, not single-page).
Change detection design: content hashing plus field-level diffs. Hash the extracted JSON object (excluding volatile fields like scraped_at). If the hash matches the previous scrape, no write. If the hash differs, write a new row and log which fields changed. This is the foundation of pricing alerts, catalog monitoring, and inventory tracking — your downstream consumers subscribe to the diff stream, not the raw crawl output.
Storage: The Three Tiers
Ecommerce scraped data needs three different storage patterns based on access patterns:
Hot catalog (Postgres). Current state of each SKU — the single row that represents what the product is right now. Read-heavy for product lookups, updated in-place on catalog crawl. Standard Postgres indexing; no special scale concerns below 10M SKUs.
Time-series pricing (ClickHouse or TimescaleDB). Every observed price with a timestamp. Append-only, query by SKU + time range. ClickHouse handles this at arbitrary scale with columnar compression — a billion price observations compresses to 2–5 GB. TimescaleDB offers similar scale with Postgres compatibility.
Bulk review text (S3 + Parquet or Elasticsearch). Reviews are too large for relational storage at scale (100M+ reviews across a catalog). Parquet files on S3 work for batch analytics and LLM training exports. Elasticsearch or OpenSearch work for review search. Don't put review text in your transactional database; queries against it will be slow and the data doesn't change often enough to justify the compute.
Data pipeline flow: scraper writes to a staging queue, a dispatcher splits records by type into the appropriate store, consumers subscribe to change events via CDC or direct query. This is standard ETL for ecommerce scraping — the scraper produces events, the pipeline distributes them.
Using the Data for AI
The reason ecommerce scraping demand has grown faster than other scraping verticals is AI. Three application patterns drive the demand:
Product Q&A models. Feed catalog data plus review text into a RAG system over product corpus. Customer asks "Does this jacket fit true to size?" — the model retrieves the relevant reviews and generates an answer. The scraping input is catalog plus review text, denormalized per product. Freshness requirement: weekly on catalog, daily on new reviews.
Price intelligence models. Competitive pricing prediction, dynamic pricing recommendation. Input is the pricing time-series across a market (your SKUs plus competitor SKUs), output is a suggested price. Freshness requirement: hourly or faster on pricing, weekly on catalog, irrelevant for reviews.
Sentiment and feature extraction. Parse reviews to identify product features customers complain about, extract comparative statements ("better than X"), classify sentiment per feature. Input is bulk review text, output is structured feature labels. Freshness requirement: weekly batch is fine; the models don't need near-real-time data.
Each application has different input requirements, so the scraping architecture has to serve all three without forcing the least-common-denominator. Over-scraping catalog to hit pricing freshness requirements is wasteful; under-scraping reviews to save proxy cost starves the Q&A model of data.
Scrape to the hottest requirement per data type, then let downstream consumers read snapshots. Pricing at hourly cadence, catalog weekly, reviews daily incremental — each consumer gets the freshness it actually needs by querying the store with time bounds, not by dictating the crawl cadence.
The Legal Posture
Ecommerce scraping sits in a more favorable legal position than most scraping verticals post the 2024 Meta v. Bright Data ruling — publicly accessible product data is not protected by CFAA. But two risk vectors remain:
Terms of Service clickwrap. Logging into Amazon Seller Central to scrape your own seller data binds you to the ToS; scraping public product pages without a login doesn't. The distinction matters for any scraper that touches account-gated pages (inventory data, seller metrics, restricted reports).
GDPR on review scraping. Reviews often contain personal identifiers (reviewer names, profile links). Scraping European-origin review data without GDPR Article 14 notification triggers EU compliance obligations. The €240,000 KASPR fine for LinkedIn scraping is the reference case — it extends to any personal data scraping without lawful basis. Strip reviewer PII at ingestion if you don't need it; aggregate-only analysis is safer than per-reviewer storage.
Amazon specifically has a history of sending cease-and-desist letters and filing suits against scrapers. The legal outcomes have been mixed, but the operational cost of litigation is real regardless of outcome. Rate-limiting, respecting robots.txt, avoiding logged-in sections, and not circumventing anti-bot specifically for account-gated data reduce the exposure.
Need help architecting an ecommerce scraping pipeline? Talk to an engineer — we design catalog, pricing, and review extraction systems with adaptive cadence and per-type storage for teams building AI features on product data.
What This Architecture Cannot Tell You
What the right SLA is for your specific use case. A price comparison dashboard used by procurement teams weekly has fundamentally different freshness requirements than a dynamic pricing engine updating retail stores in real time. Both are "ecommerce scraping." The architecture patterns scale; the crawl cadences that justify the infrastructure cost depend on what the data is being used for downstream.
What happens when a major retailer upgrades anti-bot. Amazon, Walmart, and Target all run sophisticated detection. A pipeline that works this quarter may need residential proxies or patched Playwright forks next quarter. The stealth tool ecosystem moves faster than any architectural decision can anticipate.
The cheapest crawl is the one you didn't need to run.
Frequently Asked Questions
How often should I scrape e-commerce product prices?
Use adaptive cadence, not fixed intervals. Sample each SKU at high frequency for one hour to measure its actual price change rate, then set the production interval accordingly — every 5 minutes for SKUs that change multiple times per hour, every 30 minutes for intra-day changes, daily for stable items, weekly for slow categories. This reduces crawl volume by 10–100x versus hourly crawls on all SKUs.
What's the cheapest architecture for large-scale e-commerce scraping?
Separate pipelines by data type: catalog (weekly), pricing (adaptive cadence per SKU), reviews (daily incremental — new reviews only). Store catalog in Postgres, pricing in ClickHouse or TimescaleDB for time-series, reviews in S3 Parquet or Elasticsearch. A pipeline that treats all three as one crawl type over-scrapes the cheap layer and under-scrapes the valuable one.
Is it legal to scrape Amazon product data?
Publicly accessible product pages are not protected by CFAA after the 2024 Meta v. Bright Data ruling. Scraping from a logged-in Seller Central account binds you to Amazon's ToS, which is a different legal vector. Amazon has sent cease-and-desist letters and filed suits, so operational litigation cost is real regardless of outcome. Respect robots.txt, rate-limit requests, and avoid logged-in sections to reduce exposure.
How do I store millions of e-commerce price observations efficiently?
Use ClickHouse or TimescaleDB for time-series pricing data. ClickHouse compresses 1 billion price observations to 2–5 GB with columnar encoding; TimescaleDB offers similar scale with Postgres compatibility. Append-only writes, partitioned by time range. Avoid putting time-series pricing data in transactional Postgres without partitioning — performance degrades past 100M rows.
What data do I need to scrape for an AI product Q&A system?
Catalog data (title, description, specs, images) plus review text plus rating distribution. Denormalize into a single record per product with associated reviews, store in a RAG-compatible format (Parquet, JSON Lines). Freshness requirement is weekly on catalog, daily on new reviews. Pricing data is usually not needed for product Q&A — it's the input for a different model.
Related posts

cURL vs. Playwright vs. LLM Scraper: A Decision Tree for 2026
Same scraping job costs $5, $150, or $6,000/month depending on the tool you pick. A 4-step decision tree for choosing between HTTP clients, headless browsers, and LLM-native scrapers in production.

Debugging Scrapers in Production: Logs, Screenshots, Video Replay, and Failure Forensics
Status codes aren't evidence. Screenshots, DOM snapshots, Playwright traces, and structured logs are. The Evidence Bundle pattern for turning scraper failures from hours of investigation into dashboard queries.

Playwright vs. Puppeteer vs. Selenium for Production Scraping: A 2026 Comparison
Most headless browser comparisons test on localhost. Production scraping at 100+ concurrent sessions reveals different winners — here's the data on memory, detection, and Kubernetes deployment.