ivinco
Web Scraping Legal Landscape 2026: CFAA, Terms of Service, GDPR in Practice

Web Scraping Legal Landscape 2026: CFAA, Terms of Service, GDPR in Practice

Ivinco Team·

This post is not legal advice. It's a map for engineering teams, built from public court rulings, regulator guidance, and enforcement actions through April 2026. Talk to a qualified attorney before making specific decisions about your pipeline. The goal here is to help you ask better questions of that attorney, not to answer them.

The 2024 Meta v. Bright Data ruling changed the default answer for "is web scraping legal in the US." Judge Edward Chen ruled that platforms cannot restrict non-users from accessing publicly available data, and that Bright Data did not breach contract by circumventing logged-in protections when those protections were themselves deceptive. Combined with the earlier hiQ v. LinkedIn precedent, US courts have now twice rejected the argument that public data scraping violates the Computer Fraud and Abuse Act.

That's the good news. The other three layers aren't so clean. Contract law (clickwrap ToS), privacy regulation (GDPR Article 14, CCPA), and a newer DMCA Section 1201 vector opened by Reddit v. Perplexity in 2025 create three additional layers of legal risk that survive the CFAA ruling. Teams that read the Bright Data ruling as "scraping is now legal" and build production pipelines on that assumption are reading half the story.

The useful framing is the Authorization Line: the point where "publicly accessible" stops protecting the scraper, and contract, privacy, or circumvention law starts applying. Where your scraper sits relative to that line determines which legal regime governs you.

Layer 1: The CFAA (US Federal)

The Computer Fraud and Abuse Act (18 U.S.C. § 1030) criminalizes accessing a computer system "without authorization or exceeding authorized access." For two decades, platforms argued that ToS violations or IP bans constituted "without authorization" — scraping in defiance of either triggered CFAA liability.

Two cases undid that reading:

hiQ Labs v. LinkedIn (9th Circuit, 2019, remanded and finalized 2022). LinkedIn sent cease-and-desist letters to hiQ, blocked its IPs, and sought injunction. The Ninth Circuit ruled that scraping publicly accessible data — data that doesn't require login — does not violate the CFAA. hiQ eventually settled with damages and destroyed scraped data, but the CFAA principle held: public data is not "protected" within the meaning of the statute.

Meta Platforms v. Bright Data (Northern District of California, 2024). Judge Edward Chen went further: Meta could not prevent Bright Data from scraping non-user-facing public Facebook and Instagram data. The ruling explicitly rejected Meta's contract-breach claims where the underlying protections were found to be in service of restricting public data. This is the strongest legal precedent for commercial scraping as of 2026.

What CFAA still covers: scraping authenticated sections after creating a login, circumventing technical access controls that protect genuinely private data, accessing data that's not published on a public URL. A scraper that creates sock-puppet accounts to access gated pages is well outside the fact pattern Bright Data addressed — and arguably within CFAA's remaining reach.

Layer 2: Contract Law (Clickwrap Terms of Service)

CFAA is federal criminal law. Contract law is separate and survives the CFAA ruling intact.

The distinction that matters: browsewrap vs. clickwrap.

  • Browsewrap ToS — terms shown in a site's footer link, which users are theoretically bound by through use of the site. Courts have ruled browsewrap generally unenforceable against users who didn't actively assent. A scraper hitting public product pages on a site with footer-only ToS has weak contractual exposure.

  • Clickwrap ToS — terms accepted by clicking "I agree" during account creation. These are binding contracts. A scraper operating through a created account is bound by the ToS of that account, regardless of what the CFAA says about public data.

This is where most teams quietly assume more risk than they understand. Scraping Amazon product pages without logging in is the fact pattern the Bright Data ruling addresses, and the only ToS that applies is browsewrap (generally weak against non-assenting users). Scraping from Amazon Seller Central after logging in with a seller account is bound by the Seller Central ToS, which prohibits automated access. The same scraper, same target, different legal vector depending on whether the login happens.

Practical contract risk reduction:

  • Don't log in unless the data genuinely requires it
  • Don't create accounts to scrape (this converts browsewrap exposure into clickwrap)
  • Don't scrape sections explicitly requiring agreement to ToS as a condition of access
  • If logged-in scraping is required, have counsel review the target's specific clickwrap terms against your business model

Amazon, LinkedIn, Facebook, Glassdoor, and Reddit all have aggressive legal teams that monitor scraping at scale. A cease-and-desist letter is a material operational cost even if you ultimately prevail. Companies scraping at volume should budget for legal response time the way they budget for proxy costs.

Layer 3: Privacy (GDPR, CCPA)

Privacy regulation is the layer most likely to surprise US-based teams. The CFAA ruling doesn't touch it. Scraping publicly available data that includes personal information — names, email addresses, LinkedIn profiles, review authorship — triggers GDPR obligations when the data subject is in the EU, regardless of where the scraper operates.

GDPR Article 14 requires data controllers who collect personal data indirectly (from sources other than the data subject) to notify the data subject within one month, providing the identity of the controller, purpose of processing, and rights the data subject can exercise. For scraped data, this means sending notices to the individuals whose data was scraped. The operational cost is significant for any scraper collecting personal data at scale.

The €240,000 KASPR fine imposed by France's CNIL in 2024 is the reference enforcement action. CNIL found that KASPR scraped LinkedIn profiles, enriched them with contact information, and sold the dataset — without providing Article 14 notification, without valid lawful basis, and without respecting data subject rights. The fine signals that EU regulators will pursue scraping-driven business models when personal data is involved.

CNIL's guidance on web scraping documents the legitimate interest balancing test that determines whether scraping personal data is permissible:

  1. Is the scraping genuinely needed for a legitimate purpose?
  2. Does the scraping respect the reasonable expectations of data subjects?
  3. Are technical safeguards (robots.txt compliance, rate limiting) in place?
  4. Do the commercial benefits to the scraper outweigh the privacy impact on data subjects?

Ignoring robots.txt now weakens the legitimate interest defense in EU jurisdictions. What was once an operational nicety has regulatory weight.

CCPA (California) has narrower scope than GDPR but covers similar ground — consumers have rights to know, delete, and opt out of sale of personal information, and scrapers selling California resident data are regulated entities. The enforcement track is less aggressive than GDPR but real.

The practical pattern:

  • Strip personal identifiers at ingestion if you don't need them for the downstream use case
  • Aggregate-only analysis is safer than per-person storage
  • If personal data is genuinely required (contact enrichment, reviewer identification), implement Article 14 notification or obtain explicit consent
  • Apply the same standards to EU-origin and non-EU-origin data — the economics don't justify maintaining two pipelines

Layer 4: The DMCA 1201 Vector (New in 2025)

The newest legal layer opened in 2025 when Reddit filed suit against Perplexity AI under DMCA Section 1201, alleging circumvention of technical protection measures — specifically, rate limits and anti-bot systems. DMCA 1201 prohibits circumventing "technological measures that effectively control access" to copyrighted works.

The argument: rate limits and Cloudflare challenges are technological measures; bypassing them to access copyrighted content (Reddit posts) is circumvention under 1201. The case is unresolved as of April 2026, but the theory has implications beyond Reddit:

  • Every use of patchright, undetected-chromedriver, or residential proxies specifically designed to bypass Cloudflare could be argued to be DMCA 1201 circumvention
  • Scraping any site whose content is copyrighted (news articles, forum posts, user reviews) while bypassing anti-bot protection is in scope

Whether the Reddit case succeeds matters less than the existence of the theory. Plaintiffs now have a new vector beyond CFAA for pursuing scrapers at scale. The legal calculus for high-volume scraping of copyrighted content has gotten more complex, not simpler, despite the favorable CFAA rulings.

The Compliance Floor

The operational practices that don't guarantee anything but reduce exposure across all four layers:

Respect robots.txt. Not legally binding in the US after the Bright Data ruling, but CNIL weighs robots.txt compliance in the GDPR legitimate interest test. A scraper that disrespects robots.txt has a weaker position in EU jurisdictions and in any contract dispute. The cost of compliance is near zero.

Rate-limit deliberately. Excessive request volume risks being characterized as abusive access or denial-of-service under various computer crime statutes. The defensive pattern is target-aware rate limiting — adaptive delays based on observed response time, not fixed fast intervals. Retry patterns apply here; rate limiting is a reliability feature as much as a legal one.

Don't circumvent authentication. The bright line is clear: scraping data behind a login you created is different from scraping public URLs. Authentication circumvention is where CFAA, ToS, and DMCA 1201 all intersect. Stay public.

Strip PII at ingestion. If your downstream use case doesn't need personal identifiers, remove them from the scraping pipeline before storage. Aggregate-only analysis sidesteps most privacy regulation.

Don't publish scraped training data. Reddit v. Perplexity is partly driven by scraped content flowing directly into consumer products. Teams using scraped data for internal analysis, model training, or market research have a different exposure profile than teams republishing scraped content.

Keep an audit trail. Log what you scraped, when, from where, under what authorization theory. The audit trail is defensive evidence — it shows good-faith operation if disputes arise.


Need help with a scraping compliance review? Talk to an engineer — we help teams design scraping architectures that reduce legal exposure through technical controls. For specific legal guidance, we'll recommend qualified counsel.

What This Framework Cannot Tell You

Whether your specific pipeline is legal. That depends on jurisdiction, target data, business model, user agreement structure, and case law in the relevant venue. Qualified counsel familiar with both data protection law and your industry is the answer. This map should help you ask better questions, not replace that conversation.

Whether the law will stay where it is. The Bright Data ruling was issued in 2024. The Perplexity case is unresolved. GDPR enforcement patterns shift with each DPA action. EU AI Act implementations are still evolving in 2026. A posture that's defensible today may not be in two years. Plan for periodic review, the same way you plan for stealth tool turnover in the detection arms race.

Public access earns CFAA safety. Everything else earns a lawyer.

Frequently Asked Questions

Scraping publicly accessible data is generally legal under US federal law after the 2024 Meta v. Bright Data ruling and the earlier hiQ v. LinkedIn precedent. Both courts held that public data isn't protected by the Computer Fraud and Abuse Act. However, contract law (clickwrap Terms of Service for logged-in accounts), privacy regulation (GDPR for EU data subjects), and DMCA Section 1201 (Reddit's 2025 theory) create separate legal risks that survive the CFAA ruling.

What is the difference between browsewrap and clickwrap ToS for scrapers?

Browsewrap Terms of Service appear as footer links that users theoretically accept through use of the site — courts have generally ruled these unenforceable against non-assenting users. Clickwrap ToS, accepted by clicking "I agree" during account creation, are binding contracts. A scraper hitting public URLs without logging in has weak contractual exposure; a scraper operating through a created account is bound by clickwrap terms regardless of public-data CFAA protections.

Does GDPR apply to scraping publicly available data?

Yes, when the data includes personal information about EU residents. GDPR Article 14 requires data controllers who collect personal data indirectly to notify data subjects within one month. France's CNIL fined KASPR €240,000 in 2024 for scraping LinkedIn profiles without Article 14 notification. Public availability doesn't exempt scraped personal data from GDPR obligations, regardless of where the scraper operates.

What does the Reddit v. Perplexity DMCA 1201 case mean for scrapers?

Reddit's 2025 lawsuit alleges Perplexity circumvented technological protection measures (rate limits, anti-bot systems) to access copyrighted content. If successful, the theory extends DMCA Section 1201 liability to any scraper using tools like patchright, undetected-chromedriver, or residential proxies to bypass anti-bot protection on copyrighted content. The case is unresolved, but the legal theory is already influencing compliance planning at high-volume scrapers.

Respect robots.txt (weighed in GDPR legitimate interest tests), rate-limit requests to avoid abusive-access characterizations, don't scrape logged-in sections where clickwrap ToS applies, strip personal identifiers at ingestion when possible, and maintain an audit trail of scraped sources and purposes. These reduce exposure across CFAA, contract, privacy, and DMCA vectors. Pair technical controls with qualified legal review before launching commercial pipelines.