OpenAI moved for summary judgment on September 4 in the news publishers’ cases inside the copyright MDL, In re: OpenAI, Inc. Copyright Infringement Litigation, 25-md-3143, before Judge Stein in the Southern District of New York. The brief’s headline number: OpenAI’s expert searched for the asserted articles in a sample of 20 million ChatGPT conversation logs produced in the case and found 24 instances of verbatim regurgitation, in what the brief calls an alleged regurgitation rate of 0.00012%. The argument that matters more to site owners sits deeper, in the section on Browse. OpenAI argues that copies its browsing agent made before a publisher blocked it were “impliedly licensed.”
A Motion, Not a Ruling
The filing is Defendants’ Memorandum of Points and Authorities in Support of Motion for Summary Judgment, a public, partly redacted version filed under a Defendants’ caption but signed by OpenAI’s counsel alone. It covers five News cases, including The New York Times, The Intercept, and the Ziff Davis and Daily News plaintiffs. Judge Stein has not ruled on these motions.
News Plaintiffs cross-moved the same day. Their notice seeks, among other things, summary judgment on the prima facie elements of their infringement claims against OpenAI and Microsoft, summary judgment against the defendants’ fair-use affirmative defenses, and partial summary judgment on three elements of their DMCA claims against OpenAI. Both sides requested oral argument. Neither filing gives a date, and none appears in the sealing order’s recital of the schedule.
Two Technologies, One Argument Confined to Browse
The brief covers two technologies, by its own framing: the large language model, “trained on a vast and diverse array of material, including news,” and Browse, “OpenAI’s automated web browsing service.” Pretraining is defended purely as fair use, in Section IV.A. Browse gets two separate defenses in Section IV.B: fair use for extracting factual information, and, only for Browse, implied license.
That scoping matters. The brief does not argue implied license covers pretraining. It applies to Browse, and only to copies made before a publisher updated its robots.txt file or otherwise blocked the crawler.
The Implied-License Passage, Verbatim
OpenAI’s theory rests on how search engines have long operated: “This copying has long been understood to be impliedly licensed unless the website owner opts out via mechanisms like the robots.txt protocol.” The brief argues OpenAI “publicly announced the Browse user agent and instructed the world on how to block it.” It says publishers “chose not to act.” Its conclusion: “to the extent Browse copied any Asserted Works before News Plaintiffs chose to update their robots.txt files or otherwise block Browse, those copies were impliedly licensed.”
Footnote 15 sets Browse apart from the Meltwater precedent: “OpenAI has set up an array of distinct user-agents so that publishers might do exactly what the Meltwater court suggested: ‘communicat[e] which types of use the copyright holder is permitting the web crawler to make of the content.'”
The Times, Two Blocks, Two Dates
The brief says The Times “waited nearly a year before it blocked OpenAI’s user agent” through robots.txt. It separately says The Times “implemented a hard block for ChatGPT-User in April 2024.” Those are two different mechanisms. The public text does not give the date OpenAI announced the Browse agent; it cites a separate statement of undisputed facts that is not part of this filing. So “nearly a year” cannot be checked against an announcement date in the brief itself.
OpenAI’s plugins documentation for the browsing agent dates to March 2023. PPC Land reports OpenAI disclosed the ChatGPT-User agent that month and said it honored robots.txt. Inimino checked Wayback Machine snapshots of nytimes.com/robots.txt independently of the brief. A February 26, 2024 capture lists no ChatGPT-User entry. A March 1 capture adds “User-agent: ChatGPT-User” with “Disallow: /”. The robots.txt block landed in that four-day window, at least a month before the hard block the brief dates to April.
OpenAI Says Over 96% of the Disputed Browse Copies Were Bing Snippets
A separate footnote goes after the dispute’s scale. OpenAI’s brief states that “over 96%” of the instances where News Plaintiffs claim Browse made an infringing copy were cases where “OpenAI merely received search results from Bing, containing only a short snippet of the underlying website.” The brief says OpenAI did not send a request for content to the associated website in those instances.
What Is OpenAI’s Implied-License Argument in the Publishers’ Case?
OpenAI argues that Browse copies made before a publisher blocked its crawler, via robots.txt or a server-level block, were impliedly licensed. The theory rests on decades of search-engine crawling practice and OpenAI’s public announcement of the Browse user agent. It applies only to Browse, not to pretraining, and only to copies predating a block; Judge Stein has not ruled on it.
The schedule below comes from the stipulated sealing order Judge Stein signed the day before the motions, and from the two September 4 filings.
| Date | What happens | Source |
|---|---|---|
| Sept 3 | Sealing order signed by Judge Stein | Sealing order |
| Sept 4 | Cross-motions for summary judgment filed | OpenAI brief; publishers’ notice |
| Sept 14 | Sealing statements due from parties and third parties | Sealing order |
| Sept 17 | Briefs re-filed publicly, unredacted except where sealing was sought | Sealing order |
| Oct 9 | Oppositions due | Sealing order |
| Oct 15 | Oppositions re-filed unredacted | Sealing order |
| Oct 16 | Amicus brief deadline | Sealing order |
| Nov 6 | Replies due | Sealing order |
| Not set | Oral argument requested by both sides; no date in either filing or in the sealing order’s schedule | OpenAI brief; publishers’ notice; sealing order |
What Site Owners Should Take From This
Whatever the court decides, the brief’s own logic points to one practice. Footnote 15 argues that OpenAI’s separate user agents let publishers “communicat[e] which types of use” they permit. Read the other way: a publisher who never sets a per-agent rule has, on OpenAI’s own theory, said nothing to block.
Keep a dated changelog of robots.txt changes and server-level blocks, agent by agent. Treat the date a Disallow line goes live as the fact that will matter if this argument is tested. Publishers weighing whether to block a search engine’s crawler outright have faced a version of this choice before. Tools that sync bot preferences into robots.txt can update those rules without a manual edit for every new agent; the dated record of what changed, and when, is still the site owner’s to keep.
