Google’s DMCA Stance on Scraping: A Hypocrisy Challenging the Open Web
The Hypocrisy of Google’s Data Frontier
Last week, a US court handed Google a loss against SerpApi, a company that scrapes Google’s search results. Google immediately stated its intent to appeal, reaffirming its commitment to blocking what it calls unauthorized data collection by ‘AI bots.’ This isn’t just about a search engine protecting its turf; it’s a foundational challenge to the future of data access and the very principles of the internet Google helped define, particularly as large language models clamor for more training data.
Google’s lawsuit against SerpApi, filed last December, centers on the Digital Millennium Copyright Act (DMCA), accusing SerpApi of circumventing anti-scraping technologies to sell data via an unauthorized ‘Google Search API.’ The irony is so thick one could cut it with a knife: the company that meticulously crawled the entire internet for decades, creating the most valuable advertising business in history from that aggregated data, now cites copyright and ‘disruption to relationships with rights holders’ when its own aggregations are targeted.
This isn’t merely a corporate squabble; it exposes Google’s attempt to redefine the rules of engagement for an internet increasingly powered by AI infrastructure. For years, Google benefited from the ‘implied license’ to crawl public web pages. Now, it seems that permission only flows one way. When SerpApi extracts data from Google’s publicly displayed search results – results that themselves are aggregations of other people’s content – Google suddenly shifts from being an aggregator to a gatekeeper, demanding exclusive control over the data it presents.
The true incentive here is clear: control over the data supply chain for the next generation of AI. Google sees external scrapers not just as a revenue threat, but as competitors in the race to build superior AI models. By claiming ownership over the API economy of its own search results, Google aims to throttle the data pipelines of emerging AI companies, ensuring its proprietary access remains unassailable. This isn’t about protecting content creators; it’s about protecting its nascent AI empire from a flood of new entrants armed with increasingly sophisticated web crawling tools.
DMCA as a Weapon Against Data Access
The invocation of the DMCA is particularly telling. Google’s claim is that its anti-scraping tech protects copyrighted content in search results, and that SerpApi’s circumvention threatens Google’s relationships with rights holders, especially those licensing content for ‘knowledge panels.’
This is a stretch. The DMCA was designed to combat copyright infringement and the circumvention of technologies protecting copyrighted works, like DRM on music or movies. Applying it to publicly displayed search results, which are themselves compilations, is legally tenuous and arguably a misapplication of the law. Search results are, by their nature, pointers to other content.
While Google’s presentation of those results, especially knowledge panels, might have some original elements, the underlying data often remains public or licensed for display. The ‘circumvention’ in question is not breaking into a secure database of copyrighted works, but extracting information from a public-facing interface. To argue this is equivalent to pirating a film is not just disingenuous; it’s an attempt to expand intellectual property law into realms it was never intended to govern. This move effectively tries to transform Google’s search index from a public utility — albeit a proprietary one — into a walled garden where even data sovereignty over aggregated public information is contested.
Redefining the Open Web for the AI Age
This legal skirmish has far wider implications than a single court ruling. It’s a bellwether for the coming battles over data access for AI training. If Google succeeds in establishing a precedent that scraping publicly displayed search results—even those aggregated from elsewhere—constitutes a DMCA violation, it could effectively cripple smaller AI firms and researchers who rely on open web data.
Such a precedent would grant immense power to incumbent platforms to dictate who can access and utilize the internet’s vast information repository, creating a significant barrier to entry in the AI space. Consider the historical context: Google itself built its dominance on the ability to crawl and index the web without explicit permission from every site owner. This open ethos allowed for the creation of a universal index.
Now, as the internet shifts from human consumption to machine learning, Google attempts to retroactively apply restrictive terms to its own output, which is primarily a derived work. This isn’t just about antitrust or market dominance; it’s about shaping the fundamental infrastructure of knowledge distribution for the AI age. The question isn’t whether Google has a right to protect its systems from undue load, but whether it can claim an exclusive right over the aggregation and re-use of publicly available information, simply because it was the first to aggregate it at scale.