Skip to content
News

Microsoft: Copilot Rarely Reproduces News Articles

Microsoft's Copilot rarely reproduces even full sentences from news articles and books, and almost never generates substantive passages that could stand in for the original, the company argues in new legal filings. The statements come as Microsoft defends itself against copyright claims brought by...

Microsoft: Copilot Rarely Reproduces News Articles - Microsoft Copilot
Microsoft's Copilot rarely reproduces even full sentences from news articles and books, and almost never generates substantive passages that could stand in for the original, the company argues in new legal filings. The s

Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, and almost never generates substantive passages that could stand in for the original, the company argues in new legal filings. The statements come as Microsoft defends itself against copyright claims brought by publishers, including The New York Times, along with a group of book authors.

The dispute centers on whether AI systems built with Microsoft and OpenAI technology can regurgitate protected content in ways that harm the businesses that created it. Microsoft’s latest response leans heavily on data drawn from its own product logs to counter that argument.

What the Copilot Chat Logs Show

As part of the lawsuit’s discovery process, Microsoft turned over 8.2 million Copilot chat logs to an expert hired by the news publishers. According to the company, those logs were not a random sample. They were specifically selected because they “hit on keywords implicating use of News Plaintiffs’ websites,” making them the records most likely to contain the publishers’ works.

In other words, Microsoft says the dataset was deliberately weighted toward the scenarios most favorable to the plaintiffs’ case. Even under those conditions, the company contends, the analysis revealed only a small number of instances tied to the publishers’ material. Microsoft cites a figure of 59,545 relevant results emerging from the millions of logs examined, framing that outcome as evidence that reproduction of protected articles is an exceedingly rare event rather than a routine function of the chatbot.

The Broader Copyright Fight

The case is one of several high-profile legal battles testing how copyright law applies to generative AI. News organizations and authors have argued that large language models are trained on and can output their work without permission or compensation. Microsoft, in turn, is positioning the data from its own systems as proof that Copilot does not function as a substitute for subscribing to a news outlet or purchasing a book.

By emphasizing that the chatbot seldom produces meaningful chunks of source text, Microsoft aims to undercut a core claim in the litigation: that its tools let users bypass the original works. The company’s framing suggests that even sentence-level reproduction is uncommon, and that longer passages capable of replacing the source are rarer still.

For a US tech audience watching the AI and copyright landscape, the filings offer an unusual window into how a major platform is measuring and defending the behavior of its consumer AI assistant. The outcome could shape how courts weigh technical evidence about model outputs in future disputes between AI developers and content owners.

Microsoft says the expert reviewed the 8.2 million logs and identified 59,545 results connected to the publishers’ works.

Source
Image: theverge.com

The US tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your US list is ready.

Shop Amazon Tech Deals Shop Amazon Tech Deals