NYT seeks sanctions against OpenAI, alleging it hid ChatGPT log evidence for two years
The New York Times and the Daily News plaintiffs filed a sanctions motion against OpenAI on July 9, 2026, accusing the company of concealing for more than two years that it already had the technical ability to search large samples of ChatGPT logs for copyrighted content — the same ability it had told the court and plaintiffs it lacked.
What's new
According to the motion, reported in detail by Ars Technica, OpenAI privacy engineer Vincent Monaco revealed during an April re-deposition — ordered after the court found his earlier testimony inadequate — that OpenAI held two large de-identified log samples, of roughly 10 million and 78 million conversations, that could have been made available to news plaintiffs early in discovery. OpenAI had already searched those samples for New York Times content while building a filter to block regurgitation of copyrighted material, according to the filing.
Instead of disclosing those samples, OpenAI directed plaintiffs to search a heavily redacted 20-million-log "sandbox" over eight months — far short of the 120 million logs originally requested — after allegedly misrepresenting its ability to search larger datasets. That sandbox sample was further degraded by roughly 19 billion AI-generated redactions, which the court itself found rendered it "unusable." NYT lead counsel Ian Crosby said in a statement: "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court." He added that OpenAI "claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy — while at the same time concealing that it had already done such searches."
OpenAI disputed the characterization. A company spokesperson said the sanctions motion was a "late litigation effort to access more logs and infringe more users' privacy," and argued that NYT dropping some claims earlier in the case was a sign the plaintiffs' case was weakening, not OpenAI's defense.
Context
The case is part of the roughly two-year-old copyright lawsuit brought by the Times and other news organizations against OpenAI and Microsoft, alleging their models were trained on and can reproduce copyrighted journalism. Discovery over ChatGPT's output logs has been one of the most contested fronts in the case, since the logs are central evidence for whether OpenAI's outputs constitute infringement or transformative fair use. NYT has previously narrowed some of its claims while adding claims against Microsoft, which the paper has characterized as streamlining rather than weakening its case.
Why it matters
Sanctions motions built on alleged discovery misconduct, rather than the underlying infringement claims, can reshape a case regardless of the merits of the original copyright dispute — adverse inference instructions or evidentiary sanctions could hand plaintiffs a significant advantage independent of whether training on news content is ultimately ruled fair use. The allegations, if credited by the court, would also feed into the broader pattern of scrutiny AI companies face over candor during discovery in the wave of copyright litigation now working through U.S. courts, with implications for how other plaintiffs suing OpenAI, Google, Meta, and others approach log-based discovery demands going forward.
Corroborating sources
- Techcrunch
https://techcrunch.com/2026/07/09/new-york-times-says-openai-hid-evidence-in-chatgpt-copyright-trial/
- Arstechnica
https://arstechnica.com/tech-policy/2026/07/openai-faked-inability-to-search-training-data-hid-billions-of-logs-nyt-says/
“It claimed searching ChatGPT outputs for copies of The Times’ and the Daily News Plaintiffs’ content was infeasible, burdensome, and invasive of users’ privacy—while at the same time concealing that it had already done such searches.”