Context Window 87
Provenance and authenticity are the story this week: hyped books facing questions about authorship, alongside new research and regulation on the same theme. Stick with it to the end, though, because there’s a practical note on MCP tools worth bookmarking, and a genuinely delightful, AI-assisted publishing project that’s a nice reminder of what this technology can be used for.
Questions about provenance and authenticity abounded this week. The Atlantic reported that a bestselling self-published book by H.M. Wolfe taken on by Simon & Schuster in a seven figure deal was classified by AI detection company Pangram as containing 60% AI-generated text. The New York Times and Financial Times published long reads on author and publisher insecurities. And having prepared this newsletter on Thursday, I had to rewrite quickly this morning to add a Bookseller story on a major crime novel being pulled from submission after doubts about the author’s creative process.
The Atlantic’s reporting spins out of a new academic paper (not yet peer reviewed) which studied nearly 15,000 self-published titles and found that 20% of books in the sample showed substantial amounts of AI text. That research doesn’t identify individual authors, but preview data singling out Wolfe was provided to the Atlantic.
More significantly, the paper argues that surging generative AI use is flooding and diluting the market for books. The presence of one of the leading scholars of copyright and IP, Columbia Law’s Jane Ginsburg, on the list of corresponding authors suggests that the paper is making a legal argument as much as a technical or economic one—it reads to me like one of the first detailed academic attempts to develop Judge Chhabria’s comments on market dilution in Kadrey v. Meta. That is important in the context of new class actions running the same argument.
Returning to the authenticity point, it’s important to stress that the authors involved—H.M. Wolfe in the Atlantic story and an unnamed author in The Bookseller piece—deny using AI. Wolfe’s editor commented: “We stand behind our author, and we do not recognise third-party AI detection tools as a fair or reliable measure of authenticity when used on their own.”
The reference to third-party detection tools points to the biggest winner from all of this, Pangram, which is getting extraordinary publicity. It was ever thus: the Salem Witch Trials were probably quite good for Cotton Mather’s book sales.
UTA’s Christy Fletcher makes an interesting point in the FT coverage, identifying a schism “between fiction, where there is ‘more dogma’ over AI use, and non-fiction, where its value for research is often prized.” On that question of research, OpenAI announced this week that it will provide free access to frontier models to 100,000 researchers in science, math and engineering. This strategy of integrating AI tools into research workflows can also be seen in products like OpenAI’s LaTeX editor Prism, and Claude Science. As I’ve previously suggested, this gives the AI companies insight and training data on the research process, rather than just outputs. And for publishers, if this has the intended effect of accelerating research, the second and third-order implications are clear: increased researcher output and greater challenges in determining what should be published.
With all of this in mind, it was an interesting week for LinkedIn to release a new feature allowing readers to report posts that read like AI slop, including privately flagging to authors that readers reported their content as potentially inauthentic. This feedback will also be used to train detector systems. How this works with LinkedIn’s own AI drafting features isn’t completely clear to me.
New research from the Thomson Reuters Foundation shows that, like other regulation before it such as GDPR, the EU’s AI Act is already having a global influence. Alongside the research, the Foundation published a short checklist that is worth reviewing to determine your exposure to the regulation.
Parallel research by BISG and the IPG last year showed that fewer than a third of publishers had a clear AI policy. Congratulations to The O’Brien Press for publishing and publicising theirs this week. It probably won’t surprise you that it’s not the line I would personally take, but I respect the clarity of the position. It would provide a useful thought starter for others with a critical view of AI to balance more AI-forward policies from, for example, John Wiley.
An opportunity to shape the landscape for UK-based readers: a petition calling on government to “introduce a statutory duty requiring AI developers to disclose the copyrighted works used to train generative AI in enough detail for creators to identify use of their work, enforce their rights, negotiate licences and be fairly paid” is closing on the 10,000 signatures required for government to provide a response.
If you are looking use AI practically, Simon Willison has a useful resource on using MCP servers with Claude and ChatGPT. Integrating third party data and services through a chat interface is one of the most useful AI applications for me: for example, next Wednesday morning, Claude will run a scheduled task with MCP and Zapier integrations to pull together all of the data on this newsletter from Kit, Google Analytics, socials and email responses to generate a marketing dashboard for me. If you’re interested in this kind of marketing and data automation, do get in touch.
Finally this week, a reminder that the story of AI and publishing is not only about risk and authenticity. I’m a couple of weeks late to this but Plotlines is an absolutely beautiful web publishing project mapping the course of classic novels. It was developed by the Guardian’s Head of Editorial Innovation and AI as a personal project, and besides the public facing project, there’s a helpful essay on how it was built using Claude Code.