After the Flood: AI and Fact-Checking
Craig Mod posted a new essay this week on how he uses LLMs: emphatically not for writing, but for research, analysis and administration. There are lots of good, practical ideas, but the one that stuck with me is Craig using Claude agents and an Ancestry.com account to research his family history: from a starting point of little knowledge, this “swarm of blood robots” (that phrase!) constructed a family history and folk memory. Craig concludes: “[C]ompanies should anticipate a coming deluge… LLMs are tireless and will happily exhume millions of ‘data graves’ for you given enough time and energy. Cheap, continuous exhumation has more profound implications around family, crime, and history than we give it credit for.”
I think he’s right and there are immediate implications for authors and publishers, particularly of memoirs. In recent years there has been a series of controversies over the authenticity of non-fiction books, from Raynor Winn to Jason Arday. This post is not about any of those specific cases: my interest is in the underlying dynamics. In every case, an initial allegation or investigation was followed by a flood of further allegations when the matter became interesting enough for the media to dedicate resources to follow it up—generally at the point where the story tips from one newspaper making the weather to all of the papers covering it and looking for a unique angle. Sooner or later, an op-ed writer covering the latest literary scandal will ask, why don’t publishers fact-check non-fiction? In fact, several months ago, New York Magazine ran a piece arguing that trade publishing is unprepared for the AI wave that is about to hit it, claiming “nonfiction publishing is uniquely vulnerable to AI because the industry has long neglected to do anything to ensure the books it publishes are factually accurate”.
The basic answer is that publishing has a different culture to US journalism, really the only example of institutionalised fact-checking and itself not error-free despite those best efforts. Publishers have traditionally relied on contractual warranties from the author. Higher risk books may be given a legal read, though that is to minimise the risk of a libel or defamation claim, not to ensure accuracy. Of course, time and money—neither of which is in abundant supply in many publishers—also play a part. Extensive fact-checking and editorial due diligence are lengthy and expensive processes.
External critics were similarly constrained: historically, it required a particular level of newsworthiness, outrage or animus towards an author to comb forensically through their life and work. But subjecting a manuscript to the equivalent of Craig’s digital bloodhounds costs almost nothing. It’s not necessarily 100% accurate, though in my experience agentic research with the right tools is generally very good, and getting better. It wouldn’t substitute for the sort of careful, on-the-ground reporting that Chloe Hadjimatheou did on The Salt Path, though it might guide the early stages of it.
When time and cost barriers drop away, anyone putting themselves into the public domain may be subject to the sort of rigorous background checking that was previously reserved for candidates for high political office.1 In that case, I could imagine a scenario where in a year or two, non-fiction submissions are routinely—potentially pre-acquisition—subjected to the equivalent of Craig’s digital bloodhounds. Instinctively that’s a deeply unlovely scenario, but I could see the logic of an argument that pre-emptive due diligence using AI might be preferable to the same sorts of systematic research being deployed post-acquisition by a journalist or third party. The best defence is based on a clear understanding of attack vectors. Literary agents or publishers could red-team the author’s backstory and understand the level of reputational and financial exposure. A swarm of fact-checking agents would cost pennies on each acquisition dollar.2
The implications of this are ugly but I think the industry should be talking about them. The tools will be available equally to investigative journalists pursuing matters of public interest and any individual with a grudge. There is no practical way to control who uses them. Might this have a chilling effect on first-person non-fiction? Or are we in a sufficiently post-truth world that reader doubt is just priced into the book? Critically, it isn’t necessary for the author or book to be at fault—a good book caught in the blast radius of a fallible process takes as much damage. There are therefore fundamental questions about procedural fairness and recourse if a book or author is investigated. At least if this is done in-house, publishers and literary agents could develop standards around it.
-
I’m reminded of Warren Ellis’s satirical comic Transmetropolitan, which featured a vice presidential candidate grown in a lab vat so as to be free of any incriminating past. ↩
-
It would be naive to imagine this isn’t already happening on some level. I know of at least one editor who used a Deep Research agent to assess a book on a particular specialist subject and conclude it was not as ground-breaking as a confident submission suggested. ↩