The API Hole
A new beef over an old gripe
Sometime last year, the New York Times, the Guardian, USA Today and a few other major publishers began blocking the Internet Archive’s Wayback Machine from scraping their sites. As I wrote back in May, the publishers’ beef with the Archive had to do with its APIs, which allow third parties to access content from the Wayback Machine (including of course journalists working for the very publishers now blocking it). In the publishers’ view, the Archive’s APIs had become a back door for AI companies to scrape their content without authorization.
“A lot of these AI businesses are looking for readily available, structured databases of content,” the Guardian’s head of business affairs and licensing, Robert Hahn, told Nieman Lab in January. “The Internet Archive’s API would have been an obvious place to plug their own machines into and suck out the IP.”
The particular beef was new, but the gripe is an old one. Complaints about unguarded doors providing disapproved electronic access to information long pre-date AI scraping and even the internet itself.
In the Before Time, when playback of digital media from physical sources or via transmission was first happening, the Hollywood studios in particular worked themselves into a lather over the so-called “analog hole.” While digital-to-digital connections between devices, or between devices and transmission lines, could be encrypted to prevent the data passing through them from being intercepted and copied, analog signals passing through analog connections could not be, which meant the signals could be recorded. Thus, the “analog hole.”
The studios’ solution to that perceived problem was to lobby the government to allow them to disable any unprotected analog outputs on devices by means of an embedded signal, and to require that devices recognize and comply with the flag. Counter lobbying by device makers and some AstroTurfed “grass roots” pushback foiled the scheme.
The latest dispute over an unprotected door pits Google against SerpApi. SerpApi sells subscriptions to a service it describes as a “Google Search API.” It generates automated searches, then organizes the resulting information and packages it into structured search-results data for clients. Google complained that the automated searches imposed significant processing costs on it without producing any advertising revenue for Google to offset those costs. So, in January 2025 Google introduced SearchGuard, which blocks automated search requests it does not recognize to prevent SerpAPI and similar services from scraping it search results.
An arms race then ensued, with SerpAPI tweaking its system to evade SearchGuard and Google tweaking back. In December 2025, Google finally sued, accusing Serp of circumventing a technical protection measure and trafficking in a circumvention technology in violation of §1201 of the DMCA. The complaint echoed an earlier lawsuit brought by Reddit accusing SerpAPI of illegally scraping its archive, as well as scraping its archive content from Google, which is licensed by Reddit to use and display its content.
Setting aside the irony of Google complaining about another service diverting its ad revenue, the DMCA claims in its complaint against SerpAPI stuck many legal observers as odd. Search results, being collections of facts, generally are not copyrightable, so why was Google bringing its case under the Copyright Act?
Google argued that the the “knowledge panels” it displays alongside some search results often include copyrighted as well as non-copyrightable elements. By scraping them, therefore, SerpApi was committing copyright infringement.
Last week, however, in unusually swift fashion, U.S. District Judge Yvonne Gonzalez Rogers of the Northern District of California, tossed Google’ DMCA claims. In a 21-page order, Gonzalez Rogers ruled that since Google is not the copyright owner of any of the copyrighted material in the knowledge panels, and had not shown that it had been expressly authorized to protect that material by any rights owner, it lacked standing to bring a DMCA claim. Further, she held, a platform cannot conjure a copyright right from uncopyrightable information simply by placing an anti-bot system in front of it.
She gave Google 21 days to file an amended complaint, however, if it can show that any of its licensing agreements with rights owners authorized it to protect the licensor’s content from third parties.
SerpApi, as well as some commentators, hailed the judge’s ruling as a major win, for SerpApi, and for the open internet generally. Others have argued that Judge Gonzalez Rogers’ ruling was more nuanced than that, and that Google has a plausible path to reinstating its claims.
I don’t have the expertise to evaluate the legal arguments of either side. But the ruling, if it stands up, is likely to have practical implications for rights owners thinking of licensing their content.
Given the huge increase in the amount of automated web scraping activity triggered by AI, rights holders not wanting to see their content exfiltrated through an unguarded door may need to make sure any agreements with licensees expressly authorize them to implement active measures to prevent such scraping. They might further want to demand the right to specify or approve those measures, and to audit their effectiveness.
Rights holders may also need to weigh the value of putting more third parties under an obligation to protect their content by initiating additional licensing deals vs. the value of preserving traffic to their own sites by limiting the number of licensees.
The lawyers on both sides may also need to clarify who is liable in the event scraping happens anyway and what rights of action the parties have.
For their part, licensees may seek offsetting concessions in the price or terms of the license if the agreement imposes significant financial or operational costs on them from the need to implement bot-blocking measures.
Following the ruling, Google vowed to renew the fight. In comments to Ars Technica, Google’s spokesperson, José Castañeda, said the company plans to amend the complaint and is “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” otherwise attempting to dispute Google’s standing.
“We look forward to filing an amended complaint, as the Court invited us to do,” he added. “We remain committed to protecting our services and partners from unauthorized access.”
In yet another irony of the case, rights owners may find themselves hoping Google succeeds in its DMCA claims and hands them a few legal bricks to plugs the API hole.

