You may have seen the headlines about a Microsoft employee calling AI scraping “the largest theft of labor in human history.” It made the rounds on social media, with plenty of people insisting it was a slam dunk admission of criminality. I think I saw this gif applied to it multiple times:

The "Oh my God, He Admit It!" Meme

It certainly doesn’t look great. But does it actually matter? An exec’s statement doesn’t change the underlying facts: fair use isn’t copyright infringement, and even when something is infringement, it still isn’t theft. That someone colloquially calls it theft isn’t supposed to change the legal analysis, no matter how emotionally satisfying it feels.

For basically all of Techdirt’s history, we’ve explained that copyright infringement is not “theft,” and that using that term is misleading in dangerous ways. It’s been a few years since we last mentioned it, but the Supreme Court’s ruling in Dowling vs. the US has always been the clearest legal statement on this point.

Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.

That’s not to say that infringement is necessarily okay. But it’s a distinct issue from “theft.”

For years, most of the internet seemed to agree. But the arrival of generative AI has triggered a remarkable amount of backsliding, as people rush to (incorrectly) label training as “theft.” That’s wrong on two levels. First, if it were infringement, it still wouldn’t be “theft.” Theft removes something from someone’s possession. Copyright infringement doesn’t. It makes a copy without a license. Those are not the same things. At all.

But there’s an even bigger issue when it comes to AI training, which is that there’s a fairly strong argument that training is fair use. And at least one judge in one of the (many) cases exploring this issue has agreed. As Judge William Alsup noted:

To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.

And here, too, the law is clear. Fair use is not a “defense” to infringement. Rather, a fair use “is not an infringement of copyright” at all. That doesn’t mean that there can’t be some aspects that are infringing and are punishable: indeed, in that very case where Alsup determined that Anthropic’s training was fair use, he also dinged them for widespread mass infringement for creating and storing “pirate libraries” of content without a license.

The point is that Alsup actually looked at the specifics of each use. An AI data scientist firing off an emotional internal message is not legal analysis, nor is it someone who understands the elements of either copyright infringement or fair use, let alone “theft” in the legal sense.

Of course, for decades now, the large copyright interests have polluted the discourse on this by deliberately equating infringement to theft (and simultaneously minimizing, dismissing, or demonizing fair use). Hollywood and others (including Microsoft!) spent years poisoning the language around copyright until calling infringement “theft” became the default, and “fair use” got treated as a grudging loophole or limited defense, rather than the public’s actual right.

Bill Patry, who knows more about the modern history of copyright law than maybe anyone, wrote a wonderful book about how the large copyright players used the language of “theft” and “piracy” to influence policy discussions.

Given all that as background, it’s somewhat hilarious that people are acting like it’s a huge deal that a Microsoft employee called AI training “theft.” This came out in a recently unredacted filing from the NY Times in its case against OpenAI.

This case is about, as Microsoft’s Director of Applied Science put it, “an astonishing theft of unprecedented proportions”; SF1437, perhaps the “largest theft of labor in human history.” SF1652. Defendants repeatedly copied millions of Plaintiffs’ copyrighted articles in their entirety without permission to produce substitutive commercial AI products.

The NY Times is trying to win a copyright lawsuit, so of course it has every interest in portraying these quotes as damning. The rest of the media doesn’t have to accept that framing — especially when it’s not how copyright law actually works. Just because a random employee of one company colloquially calls it “theft” doesn’t magically make it so, either legally or morally.

And, yes, that same filing goes after the fair use argument by quoting an OpenAI employee calling AI an “existential threat” to publishers, then insisting that this “undermines” any fair use claim. Notably, the filing quotes so little of the surrounding context that it’s not even clear what the person was referring to, but just because one employee makes such a statement doesn’t make fair use disappear. That’s not how fair use is determined.

There’s also a more basic problem with the “existential threat” argument. Lots of things can be “existential threats” to companies that refuse to adapt and change. That doesn’t make their competitors illegal. It’s just how competition itself works. This is also true of the “doom loop” quote from the filing:

A Microsoft document recognizes that nobody wins that contest: “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

Read the full filing, though, and the Times’ lawyers seem to be playing a neat bit of sleight of hand here, conflating statements about search results that give users the facts they were looking for (rather than driving them to a separate website) with the claim that chatbot output is a substitute for news. Those are two separate things. Take the Satya Nadella testimony the filing leans on. He’s talking about chatbots answering questions a user has, not about anyone going to ChatGPT and asking it to replace the NY Times:

Microsoft’s CEO Satya Nadella agreed under oath that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

And, in the end, what matters most is what users actually want, because that’s what they’re going to do regardless. The NY Times might not like that someone looking for a quick answer doesn’t want to read a long article, but that’s not the fault of an AI system. It’s how people work: the AI tools (or the search results) may simply be meeting that reader’s needs in that moment better than a long-form article does.

As an entity engaged in long-form reporting ourselves, that certainly represents a challenge for us, but we try to respond to that by providing something that can’t be replaced merely by a straight answer to a question. Instead, we focus on providing more value that makes it worthwhile to read our full commentary.

Why does the NY Times think it can’t do that? Does it really value what it does so little that it can’t compete with a word generator?

The courts still have a long way to go on the copyright questions around training, but the rush to wave a few cherry-picked quotes from the NY Times filing around as “evidence” that AI training broke the law is getting silly.

Fair use isn’t infringement. And if it is infringement, it’s not “theft.” And even if a use competes with you, competition doesn’t make it illegal. Indeed, the “effect of the use upon the potential market” is just one of four factors, and it isn’t supposed to let ordinary competition swallow the entire fair use analysis. That these employees (who are not, themselves, copyright experts) said random things does not change the actual underlying analysis of fair use. Or, at least, it shouldn’t.

Leave a Reply