Context Window 89
Happy Friday. Thanks to Robin Sloan for mentioning the newsletter this week, and welcome to everyone who subscribed as a result. In recent issues, we’ve looked at issues of content authenticity and provenance from different angles: this week, the discussion is about whether AI is a solution as well as a problem—either through marking LLM outputs, or using agents for fact-checking.* The most significant news this week was Anthropic announcing how it intends to comply with the transparency obligations in the EU AI Act. Claude models introduced after 2 August will add watermarks to text content and, for supported file types like images, provide C2PA provenance metadata (future support for watermarking in older models is promised). Forthcoming documentation promises to show how to detect those watermarks. In line with the piece from a few weeks ago about how the act is influencing regulation globally, these changes will apply around the world.
For an industry like publishing which has grappled with issues of authenticity of writing, this is a significant step. But it’s worth unpacking the limits to the announcement. Obviously this only applies to outputs generated after the feature was introduced, and to Claude (other LLMs are available). The devil is in the documentation, much of which is not available yet. And most importantly: all that a watermark proves by itself is that text was processed through Claude. It cannot distinguish between, for example, generating a complete chapter, and providing short editorial feedback on existing material. I can’t wait for the first witch hunt started by someone who doesn’t get that distinction.
That detail means that for many authors and publishers the simplest way of avoiding the stigma of a Claude watermark may just be not using it for writing. That was the general conclusion of a lovely new essay by Craig Mod on how he uses LLMs. It also has some really nice practical ideas that everyone could be inspired by.
For me, the most interesting and moving part of Craig’s essay is describing using AI agents to research a family history that was unknown to him. His description of that tireless, automated research process inspired me to write something about what this sort of AI research tool means for non-fiction publishing, particularly memoirs. It’s a nuanced and timely issue and one that I think writers, literary agents and non-fiction publishers should be getting ahead of now, before it comes for them.
Moving from trade to academic publishing: STM has opened a second stage of consultation on a global reporting standard for AI disclosure in academic research. Interested parties—many of you, I hope—have until 16 October to respond.
From big picture issues and risks to something more practical: as a summer holiday project, I just updated my research on the UK publishing industry based on public records, including mapping more than 12,000 registered publishing companies.
The wider relevance of this for readers outside UK publishing was the use of AI in the process. Last year with ChatGPT, it took three hours. This year with Claude Cowork: thirty-three minutes, of which I was actively engaged for seventeen. I’ve written up the process and findings on my website: it poses interesting questions for anyone using AI for data analysis.
Spotify announced a raft of new AI features, including playlist curation, asking natural language questions of books and better signalling of series data. The first two are pretty much table stakes for book consumption platforms in 2026. The last one is what makes me pause. If publisher data feeds had complete, consistent series data, there would be no need for an AI feature. Unfortunately, from what I have seen in recent years, the story is really about what is often poor metadata, not the AI remedy.
A rare example of candour in executive communication this week, which also speaks to consumer attitudes: the CPO of streaming platform Twitch, asked why it had made training on user data by its parent Amazon an opt-out not opt-in feature, replied bluntly, “If it was opt-in, nobody would opt-in.”
Verity Harding is one of my favourite writers on AI—it’s a couple of years old now, but if you haven’t read her book, you’re missing out. This week she published a new essay on AI in Hollywood, seen through the lens of Christopher Nolan’s views on the technology versus George Lucas’s. There are lots of applicable lessons for publishers and creators in general: in particular, cheaper production doesn’t in itself guarantee success, and, relevant to the previous link, audience acceptance is as important as tech capability.
Finally, an interesting new project called FineBook intends to test and benchmark OCR models for their accuracy in digitising old books and documents, and to reprocess existing public domain digital libraries using state-of-the-art models to improve the accuracy of digitisation. One for librarians, archivists and archive publishers to watch.