Pangram’s AI Detector Is Nearly Perfect in the Lab: Publishing Is Treating Its Score Like a Verdict

0
1
Pangram’s AI Detector Is Nearly Perfect in the Lab: Publishing Is Treating Its Score Like a Verdict


A tool with a false-positive rate near zero in independent testing has already helped cancel a book deal, void a literary prize, and end at least one journalist’s byline. Pangram Labs built an AI-text detector that multiple universities say is the most accurate on the market. Publishing has started treating its single percentage score as a verdict.

What Pangram Is

Pangram Labs is a Brooklyn-based startup founded by Max Spero and Bradley Emi, two Stanford computer science graduates who previously worked at Google and Tesla, respectively. The company raised a seed round of roughly $4 million in mid-2025, led by Haystack VC and ScOp, then closed a $9 million round in mid-2026 to expand from text detection into image detection, bringing its total funding to about $13 million, according to SiliconANGLE and BusinessWire. Pangram offers a free consumer checker alongside paid enterprise access, and counts universities, publishers, businesses screening for fake product reviews, and the journalist-sourcing platform Qwoted among its customers.

The company claims a false-positive rate around 1 in 10,000, meaning it wrongly flags human writing as AI-generated in roughly one out of every ten thousand cases, and says the detector generalizes to new AI models without retraining. The claims are strong for a category that has struggled with reliability since AI detectors first appeared.

What Independent Testing Actually Shows

Three academic evaluations back up at least part of the claim. A University of Chicago Becker Friedman Institute study from August 2025 found Pangram’s false-positive rate at 0.001 and its false-negative rate at 0.01, ahead of competitors including GPTZero and Originality.ai. A Vrije Universiteit Brussel study covering 160 academic papers in June 2026 recorded zero false positives and 97.5 percent detection on fully AI-generated text, with 95 percent detection even on AI text run through humanizing tools. A University of Maryland evaluation put Pangram’s detection rate between 98 and 99.3 percent on paraphrased AI text, with a false-positive rate of 2 to 2.7 percent in that harder scenario.

Pangram highlights all three studies on its website, which does not make them wrong, but does mean the company selected which comparisons to publish. None of the figures above have been independently reproduced outside the original research teams, and the studies test detection in controlled academic settings rather than the messier conditions of a publishing house or an award committee.

What Happens When the Score Leaves the Lab

WIRED’s Lexi Pandell reported that Pangram’s scores have already shaped real careers. Hachette canceled the release of Mia Ballard’s novel Shy Girl after Pangram’s CEO posted that it scored 78 percent AI-generated. A New York Times Modern Love column reportedly registered 100 percent. A thriller called Call Me, I’ll Hide the Body, sold for $2.4 million, scored 97 percent. The winning entry in the 2026 Commonwealth Short Story Prize, Jamir Nazir’s The Serpent in the Grove, scored 100 percent on Pangram, according to Slate and The Week, and the author’s explanation, that he used AI only for research and drew influence from Derek Walcott’s poetry, has not resolved the dispute.

Publishing consultant Jane Friedman told WIRED that writers have grown to resent detection tools nearly as much as the AI models themselves, saying some see them as “just as evil, if not more evil, than the AI companies.” None of the cases above came with a public appeals process, a published confidence interval, or a second independent test before the consequences landed.

A Good Statistic Is Not a Good Process

The tension in Pangram’s story is not about whether the tool works. On the evidence available, it works better than its named competitors. The tension is about what a benchmark accuracy rate is allowed to mean once it leaves a controlled study and becomes the sole input into decisions about book deals, literary prizes, and bylines. A false-positive rate of 1 in 10,000 sounds reassuring until it is applied across millions of manuscripts, student essays, and submissions, at which point even a vanishingly small error rate produces real, specific, named people wrongly accused.

The Digital Staffroom, an education-focused critic of AI detection tools, has made a sharper version of this argument: the writers most likely to get flagged incorrectly tend to be the ones who write with unusual polish or precision, the exact population publishers and prize committees are trying to reward. A detector tuned for aggregate accuracy can still be a poor fit for a process that needs to protect individuals from a single wrong call.

What Should Change

Pangram’s accuracy is not the problem publishing needs to solve next. The absence of due process around how that accuracy gets used is. Any institution deploying a single AI-detection score as grounds to cancel a contract, void a prize, or end a byline should pair it with a transparent appeals mechanism and a policy that treats the score as evidence to investigate, not a verdict to act on. Pangram cannot fix that gap by itself. It sells a signal. Publishers, universities, and award committees are the ones turning that signal into a verdict, and they have mostly done it without publishing the rules.

The technology has gotten good enough that the accuracy debate is largely settled. The harder argument, over what a percentage score should be allowed to end, is only getting started.