← Back to Geekish
AI Policy • September 17, 2026

Unsealed AI scraping filings put publishers back in the spotlight

Newly unsealed filings are adding sharper language to the fight over AI training data, news content, and publisher economics.

Illustration of court filings, news pages, and AI training data flowing into a model
Image: Geekish-generated illustration based on Ars Technica and TechCrunch reporting linked below.
AI COPYRIGHTGeekish sourced quick read

Original Geekish context based on the sources linked below.

The short version

Ars Technica reports that newly unsealed court filings in a copyright fight involving news publishers, Microsoft, and OpenAI include internal Microsoft warnings about scraping news for AI training. TechCrunch separately reports on the same filings and the publisher concerns behind them.

The quote getting attention

According to Ars, plaintiffs cited Microsoft Director of Applied Science Brent Hecht warning that scraping news for AI training was "an astonishing theft of unprecedented proportions" and potentially the "largest theft of labor in human history."

Why this matters

The legal fight is not only about one dataset. It is about whether AI companies can treat the open web, paywalled reporting, and journalism archives as training material under fair use, and what happens to publishers if AI products replace the traffic those stories used to earn.

Geekish take

AI copyright arguments often sound abstract until the internal emails surface. If the filings hold up, they make the industry's "we did not know this was a problem" posture much harder to sell.

Want more tech without boring tech-site energy?

Follow Geekish for sourced quick reads, AI, gadgets, apps, creator tools, and internet culture.

Get the tech drop