
Is Web Scraping Legal? US, EU and UK Law in 2026
Yes. Scraping publicly available web pages is generally legal in the US, the EU and the UK. A project becomes illegal through how and what you collect: getting past a login, breaking terms you accepted, collecting personal data without a lawful basis, copying protected content, or getting around anti-bot systems. The lowest-risk case is public pages, collected while logged out, with no personal data, and facts rather than articles or photos. The same rules apply whether you call it web scraping, screen scraping or data scraping.
New in 2026: two federal courts ruled on getting around Google's bot detection under the DMCA (Google v. SerpApi and Reddit v. SerpApi); the Third Circuit found that copying Westlaw headnotes to train an AI search tool was not fair use (Thomson Reuters v. Ross); and the Ninth Circuit, at an early stage, treated an AI shopping agent as its user's tool (Amazon v. Perplexity). Each is covered below.
This article is general information, not legal advice. The rules differ by country, several of the cases below are still running, and a project with personal data or real money at stake should get a lawyer's review before it starts.
Five ways a scraping project becomes illegal
1. Getting past a login or another access barrier (CFAA)
In the US, the main anti-hacking law is the Computer Fraud and Abuse Act (CFAA). In Van Buren v. United States (2021), the Supreme Court described its test as "gates-up-or-down": either you are allowed into a system, or a part of it, or you are not.
The Ninth Circuit applied that test to scraping in hiQ Labs v. LinkedIn (2022). Public pages have no gates to lift or lower, so the court, affirming a preliminary injunction, concluded that scraping them is likely not access "without authorization". The same opinion says that parts of a site that require authorization are still protected, and that site owners still have other claims, such as breach of contract, trespass to chattels and copyright infringement.
Two later cases show where the edges are. In Ryanair v. Booking.com, a jury found Booking.com liable under the CFAA in 2024, but the judge set the verdict aside in January 2025 because Ryanair had not proved at least $5,000 in loss, which a civil CFAA claim requires. In Amazon v. Perplexity, the Ninth Circuit vacated a preliminary injunction against Perplexity's shopping agent in August 2026, finding Amazon unlikely to prove, on the record before it, that Perplexity rather than the user accessed Amazon. The court said it was not creating "a new legal regime governing agentic AI" and noted that Amazon can still control access through its terms.
In the UK, the Computer Misuse Act 1990 makes it an offence to make a computer perform a function to get access you know is unauthorized. The clear risk is logging in with credentials you were not given or breaking through an access control.
2. Breaking terms you accepted (contract)
A site's terms of service are a contract, and they bind the people who accept them, usually by creating an account. LinkedIn could not use the CFAA to stop hiQ from scraping public profiles, but it won on contract. In November 2022 the court found that hiQ had breached LinkedIn's User Agreement, both by scraping and through contractors who registered fake profiles, though hiQ's waiver and estoppel defenses to the scraping claim were left for trial. A month later the case ended in a consent judgment: $500,000 against hiQ, a ban on scraping LinkedIn in violation of its User Agreement and on fake accounts, and an order to delete the scraped data.
Scraping while logged out is different. In Meta v. Bright Data (January 2024), the court found that Bright Data did not "use" Facebook and Instagram when it scraped public pages while logged out, so Meta's terms did not cover that scraping. Meta dropped the case in February 2024 instead of appealing. In X Corp. v. Bright Data, the court held in May 2024 that X's claims that scraping breached its terms conflicted with the Copyright Act, because X held only a non-exclusive license to its users' posts. X later revived claims based on harm to its servers, but not claims based on the scraping itself, and the companies settled in June 2025.
Not every court has accepted that copyright argument. In Reddit v. Anthropic, a federal judge found in March 2026 that each of Reddit's contract and tort claims over scraping for AI training had an "extra element" beyond copying, so copyright law overrode none of them, and sent the case back to California state court.
In the EU, the Court of Justice held in Ryanair v. PR Aviation (2015) that when a database is protected by neither copyright nor the database right, EU database law does not stop its owner from restricting its use by contract.
Fake accounts are the weakest position a scraper can be in: the scraper has accepted the terms and is collecting from behind a login. Both the hiQ judgment and LinkedIn's 2026 case against ProAPIs ban them. In ProAPIs, the defendants, while denying liability, agreed to a final judgment in September 2026 that bars them from scraping LinkedIn, using fake accounts or selling LinkedIn data, and required them to certify that they had deleted it.
3. Collecting personal data (GDPR, CCPA)
In the EU and the UK, data protection law covers personal data whether or not it is public. The European Data Protection Board's 2026 draft guidelines on web scraping for generative AI training say the GDPR applies to web scraping whenever it involves personal data, and that putting information on a public page is not consent to having it scraped. The board adopted the draft on July 7, 2026, and public comments close on October 30, 2026.
Private companies often rely on legitimate interest. That needs three things: a legitimate interest, a real need for the personal data to pursue it, and a balancing test in which the rights of the people concerned do not override it. The guidelines list safeguards such as collecting only what the purpose needs, filtering out sensitive data, and leaving out sites that clearly object to scraping through robots.txt, ai.txt or CAPTCHAs.
Regulators enforce this. France's CNIL fined KASPR €240,000 in December 2024 for collecting contact details of LinkedIn users who had limited them to their first- and second-degree connections, keeping data too long and starting to inform people four years late. The Dutch regulator's 2024 guidance goes further and says scraping personal data by private organizations is "almost always" a GDPR violation. That view rests on purely commercial interests not being legitimate, which the EU Court of Justice rejected in KNLTB in October 2024.
In the UK, the ICO fined Clearview AI £7.5 million in 2022 for scraping images of people in the UK into a facial recognition database. A tribunal overturned the fine in 2023 for lack of jurisdiction. In October 2025 the Upper Tribunal reversed that and sent the appeal against the fine back, and Clearview has permission to take that ruling to the Court of Appeal.
The US has no general federal privacy law, and the state laws are narrower here. California's CCPA excludes "publicly available" information from personal information. That covers government records, information a business reasonably believes the consumer or widely distributed media lawfully made available to the general public, and information the consumer shared without limiting it to a specific audience. Biometric data collected without the person's knowledge never counts as publicly available.
4. Copying protected content (copyright)
Copyright protects expression, not facts. Prices, dates and stock levels are facts; articles, photos and reviews are usually protected. Short editorial text can count too: in Thomson Reuters v. Ross, the Third Circuit in September 2026 treated Westlaw's headnotes as original works and found that copying them to train a non-generative AI legal search tool was not fair use. Republishing copied articles or photos carries more risk than analyzing them.
In the EU, databases have a separate right of their own. In CV-Online Latvia v. Melons (2021), the Court of Justice held that a specialized search engine that copies and indexes another site's database is extracting and re-using its contents, which the database maker can prohibit where it harms the maker's investment. The UK has its own database right, which protects databases built with a substantial investment.
5. Getting around anti-bot systems (DMCA 1201)
This is where the newest cases are. Section 1201 of the DMCA makes it unlawful to get around a technological measure that controls access to a copyrighted work. In October 2025, Reddit sued SerpApi, Oxylabs, AWMProxy and Perplexity, alleging that they took Reddit content from Google's search results by defeating Google's and Reddit's technical defenses. In July 2026 the court found that SearchGuard, Google's bot detection, counts as an access control as Reddit described it, and let most of Reddit's circumvention claims go forward against SerpApi and Perplexity, the two defendants that had moved to dismiss.
Google's own case against SerpApi failed the same month, though the New York court said the two rulings do not conflict. Judge Yvonne Gonzalez Rogers dismissed all of Google's anti-circumvention claims: those over search results with no copyrighted material for good, "because the DMCA does not protect works that are not copyrighted", and the rest because Google had not alleged that the copyright owners authorized SearchGuard, which Reddit had done. Google refiled narrower claims on August 10, and SerpApi moved to dismiss them again.
Neither ruling is final. Until appeals courts weigh in, the safer reading is that getting past CAPTCHAs, paywalls and bot detection can create liability that reading public pages does not. The EU's AI code of practice and the EDPB guidelines point the same way.
Web scraping laws in the US, Europe and the UK
| United States | European Union | United Kingdom | |
|---|---|---|---|
| Access | CFAA: public pages likely not "without authorization" (hiQ, 2022) | National computer crime laws | Computer Misuse Act 1990 |
| Terms of service | Bind users who accepted them; logged-off scraping contested | Can restrict use of unprotected databases (Ryanair v. PR Aviation) | Contract law applies |
| Personal data | No general federal law; the CCPA excludes publicly available information | GDPR applies to public personal data | UK GDPR, enforced by the ICO |
| Copyright and databases | Fair use, decided case by case | Copyright, database right, mining exception with opt-out | Copyright, database right, mining exception for non-commercial research only |
| AI training | Fair use rulings split | AI Act: general-purpose AI model providers need a policy to respect opt-outs | No exception for commercial training |
| Bottom line | Public pages are low risk; logins, accepted terms and circumvention are not | Personal data needs a lawful basis even when it is public | No mining exception for commercial AI training |
Is AI scraping illegal?
Not by itself. AI training adds copyright questions on top of everything above, and the answers differ by country.
United States. Courts decide fair use case by case, and the early results point in different directions.
- Bartz v. Anthropic (June 2025). The court found training on books "quintessentially transformative" and fair use. The source of the copies mattered: scanning books Anthropic had bought into a digital library was fair use, while keeping pirated copies in that library was not. The case then settled for $1.5 billion, approved in July 2026.
- Kadrey v. Meta. Meta won on fair use because the authors made the wrong arguments, and the judge said the ruling does not mean Meta's use was lawful.
- Thomson Reuters v. Ross. The Third Circuit ruled against the AI developer, as described above.
- The New York Times v. OpenAI and Microsoft. Summary judgment motions on fair use were filed in September 2026, and there is no ruling yet.
European Union. The Copyright Directive allows text and data mining of lawfully accessible content unless the rights holder has reserved those rights in an appropriate way. For content published online, the directive expects machine-readable means. The AI Act requires providers of general-purpose AI models to have a policy for complying with EU copyright law, including those reservations, and the Commission's AI Office has been able to enforce these obligations since August 2, 2026. The code of practice published in July 2025 is voluntary. Signatories commit to follow robots.txt, honor other machine-readable opt-outs, not get around paywalls, and skip sites that EU courts or authorities have found to infringe copyright persistently on a commercial scale.
The courts are still filling in the details:
- Kneschke v. LAION. The Hamburg Higher Regional Court held in December 2025 that downloading a photo to build an AI training dataset was allowed under the mining exceptions, because the photo agency's reservation was not machine-readable. Germany's Federal Court of Justice heard the appeal in September 2026, and reports suggest a referral to the EU Court of Justice is likely.
- GEMA v. OpenAI. A Munich court found in November 2025 that memorizing song lyrics in a model and reproducing them in answers infringed copyright, and that the mining exception did not cover it. OpenAI has appealed.
Personal data in training sets falls under the EDPB guidelines described above.
United Kingdom. The UK's mining exception covers only research for a non-commercial purpose, and the government's March 18, 2026 report said a broad exception for AI training with an opt-out is no longer its preferred way forward, so the current law stays. In Getty Images v. Stability AI, the High Court held in November 2025 that the trained model was not an "infringing copy" of Getty's photos. Getty had dropped its training claim because the training did not take place in the UK. Getty has permission to appeal.
How to scrape legally: a 10-point checklist
- List the sources and the fields. Collect only what the purpose needs. Under the GDPR that is data minimization, and it also cuts every other risk on this list.
- Stay on public pages. Do not scrape behind a login with an account that accepted the site's terms, and never create fake accounts.
- Read the terms of the sites you depend on. Note what they say about automated access, especially where your company has accounts.
- Check for personal data. If the data identifies people and the EU or UK GDPR applies, you need a lawful basis, filters for sensitive data and a way to tell people what you collect.
- Separate facts from content. Prices are facts; articles and photos are not. Analyze, don't republish.
- Respect robots.txt and AI opt-outs. For AI training data in the EU, they are part of the legal test, not etiquette.
- Treat CAPTCHAs, paywalls and bot detection as a legal line, not only a technical one. That is where the 2026 circumvention cases are, and the risk is highest for copyrighted content and EU personal data. Never get around a paywall; decide the rest with counsel.
- Keep the load low. In a footnote, the hiQ court said, without deciding, that scraping beyond the site owner's consent may support a trespass to chattels claim, at least when it causes demonstrable harm. Rate limits and retries with backoff are good engineering and good defense.
- Keep records. Save the source list, the rules you applied and the collection dates, including the robots.txt in force when you collected.
- Get legal advice for high-stakes projects. That means large volumes of personal data, AI training sets, or sources that have sued scrapers before.
Who collects your data matters
Getting data from a vendor does not keep you out of court. Reddit sued Perplexity alongside the scrapers, alleging that it obtained Reddit content, including through SerpApi, and in July 2026 the court let those claims go forward.
At Ivinco, we agree with the client on the list of sources and the collection rules before a scraping project starts. The client sees a free sample of the data before we talk about price. A lead engineer is the client's contact, and the team behind them does the work. Tell us the sites and fields you need.
Frequently Asked Questions
Is it illegal to scrape LinkedIn or job sites?
Scraping public LinkedIn profiles while logged out is unlikely to violate the CFAA after hiQ v. LinkedIn. But LinkedIn's User Agreement bans tools that "scrape or copy the Services, including profiles", members accept it, and LinkedIn has obtained consent judgments against hiQ and ProAPIs that ban fake accounts. In the EU, profiles are also personal data. On other job sites, titles, employers and pay are facts, but recruiter names and candidate details are personal data.
Is Amazon web scraping legal?
Product pages are public and prices are facts. Amazon's Conditions of Use exclude "any collection and use of any product listings, descriptions, or prices" and "any use of data mining, robots, or similar data gathering and extraction tools" from the license they grant, and since 2025 they have required AI agents to identify themselves and not get around CAPTCHAs. Those terms bind customers who accept them; whether they bind a logged-out scraper is less clear.
Does YouTube allow scraping?
No. YouTube's Terms of Service prohibit accessing the service "using any automated means (such as robots, botnets or scrapers)". The exceptions are public search engines that follow YouTube's robots.txt, and access with YouTube's prior written permission. For video metadata, the official YouTube Data API is the safer route.
Is it legal to scrape Reddit?
Reddit's User Agreement, effective July 1, 2026, allows crawling within its robots.txt but prohibits scraping without Reddit's prior written consent. Reddit enforces this in court. It is suing Anthropic in California state court on contract and other state-law claims, and SerpApi and Perplexity in New York over getting around technical protections.
Does Google allow web scraping?
Google's Terms of Service, effective July 30, 2026, prohibit "using automated means to access content from any of our services in violation of the machine-readable instructions" on its pages, such as robots.txt. Google's DMCA case against SerpApi was dismissed in July 2026; Google refiled narrower claims in August, and SerpApi has moved to dismiss them again.
Is web scraping for commercial use legal?
Yes, the same rules apply. A commercial purpose narrows two defenses: US fair use weighs whether a use is commercial under 17 U.S.C. § 107, and the UK's mining exception covers only non-commercial research.
Can ChatGPT do web scraping?
ChatGPT can read a page you give it or find through search, and language models are good at pulling fields out of messy pages. It does not run crawlers at scale or deliver a steady data feed; that takes crawling infrastructure. The legal rules are the same whoever does the collecting. Our decision tree for cURL, Playwright and LLM scrapers covers when each fits.
Is screen scraping legal?
Screen scraping is an older name for reading data off what a program or website displays. The same questions about access, terms, personal data and copyright decide whether a project is lawful, and publicly visible data collected without logging in is the safest case.
Can you be sued for web scraping?
Yes. Site owners sue under contract, the CFAA and state computer laws, trespass to chattels, copyright and, increasingly, the DMCA's anti-circumvention rule. hiQ ended with a $500,000 consent judgment against it and an order to delete the data. Most scraping disputes are civil, and many end in settlements or consent judgments rather than rulings.
Need web data collected within these rules? Tell us the sites and fields you need.
Related posts

cURL vs. Playwright vs. LLM Scraper: A Decision Tree for 2026
Same scraping job costs $5, $150, or $6,000/month depending on the tool you pick. A 4-step decision tree for choosing between HTTP clients, headless browsers, and LLM-native scrapers in production.

Debugging Scrapers in Production: Logs, Screenshots, Video Replay, and Failure Forensics
Status codes aren't evidence. Screenshots, DOM snapshots, Playwright traces, and structured logs are. The Evidence Bundle pattern for turning scraper failures from hours of investigation into dashboard queries.

Playwright vs. Puppeteer vs. Selenium for Production Scraping: A 2026 Comparison
Most headless browser comparisons test on localhost. Production scraping at 100+ concurrent sessions reveals different winners — here's the data on memory, detection, and Kubernetes deployment.