Context Window 90
Welcome back. Some of this week’s news reads like a continuation of the trends we’ve seen over the summer: how books are used for training data, and issues of authenticity. It’s good to season that with some more practical advice around ChatGPT’s advertising platform, and the first clear explanation of how Claude’s new watermarking will really work…
Starting the newsletter with something practical: after a long period of using OpenAI models, I’m currently using Claude and Gemini as my daily drivers. I went back to ChatGPT this week to check an archived prompt, and the biggest change that I noticed was adverts—displayed alongside outputs and also a prompt to sign up to advertise on the platform.
It’s an interesting development, and I’ll be experimenting with some ads for my book in the coming weeks. Recommended CPCs in the $2-5 range seem high compared to Meta and Google ads. For any publishers also minded to experiment, the best independent guide I’ve found to OpenAI’s ad platform is here. Forward that to your marketing team—and if anyone has any experience with the platform that they’d like to share, get in touch.
While you’re talking to your marketing colleagues, you can also show them this new research from Shopify, which suggests that AI discovery and traditional search are not an either/or scenario—both need to be optimised. The most interesting idea is what they call “buyer journey compression”: AI takes the entire process of research and deliberation about a product and reduces it to a single conversation thread. AI-assisted purchasers land directly on products, convert at higher rates and have a higher average order value.
Anthropic released more information on how text watermarking in Claude will work: not so much adding identifying data to the file, as using the choice and pattern of words themselves. There’s quite a good lay person’s explanation here: if you’ve ever played a game of Monopoly, you’ll be able to understand it. It confirms the impression I had last week: this will not be a definitive proof of LLM use, just another directional data point.
Congratulations to 404 Media on the story of the week: they placed an AirTag in a shipment of rare books and tracked it all the way to an Amazon scanning facility in Nevada. The Amazon premises in question have a mascot of a fearsome dinosaur about to tear through a book (optics much?)
In Semafor, Reed Albergotti suggested that some of the thousands employed by tech companies could be assigned to preventing this kind of PR disaster, or, charmingly, that Amazon could build a new library next to its book scanning facilities. I suspect the simple answer is that no-one making the decisions cares—though if any of the Amazon execs who subscribe want an idea for free, I made a suggestion about preservation of books several weeks ago.
Bucking recent anxiety about AI authorship, business publisher Harriman House announced a new book by investor Raoul Pal generated using an AI model grounded in the author’s other content. Developed with the support of parent company Macmillan, it will be labelled as “human originated, AI-facilitated”. This won’t be for everyone, but there’s some evidence from other media types such as paid newsletters and podcasts that suggests consumers will accept AI-generated premium content in the right context. I think I would feel differently if this were literature or memoir, but for financial non-fiction and Pal’s audience, it seems a better fit. It’s going to be fascinating to see how the market responds to this product, and it’s good to see public experimentation from a major publisher.
The Macmillan press release stresses that the book creation process was grounded in the author’s IP, though of course there is an LLM behind the scenes that was trained on a broad range of inputs. That introduces a broader question of how much a particular input to a model can be identified, a subject addressed in a new research paper in Nature which argues that with enough training inputs, it’s hard to link outputs to any particular input—what they call “attribution decay”. Caveat that the paper is based on image diffusion models rather than LLMs, but this feels important for anyone basing legal assumptions or business models around attribution.
Another week, another authenticity story in publishing, with PRH pulling an autumn title whose journalist author has been accused of plagiarism. Not an AI story per se—but compare it with this Nieman Lab piece on an experimental Google AI tool, Backstory, being used by media outlets to identify disinformation, or this story in Nature about AI agents being used to systematically identify errors in published papers. As I wrote last week, proactive AI fact checking presents serious issues around accuracy, bias and procedural fairness, but it would be naive to imagine it isn’t going to become a bigger thing in publishing.
A new study shows that over a third of academics are now starting their research workflows with AI tools, and a further quarter are using general search tools with AI elements such as summaries or overviews. The discoverability implications for academic publishers are clear. Worryingly, the same research saw librarians reporting a decrease in use of primary sources, suggesting that some researchers may not be getting past the summary.
Another piece of research that caught my eye this week. A new pre-print analyses over 800,000 hours of podcast episodes and suggests that LLMs don’t just influence writing, but that they shape the word choices people make in verbal communication.
Finally this week, I’m pleased to say that I am joining my friends and colleagues at BISG for a new webinar next month looking at managing AI risks in publishing: if you’re free on 24 September, I do hope you’ll join us.