Pangram Becomes the Benchmark for AI Detection: Is It Reliable?

As our discussion unfolds, Spero’s replies come in fits and starts, not solely because he’s eating. I inquire about his hiring philosophy. Spero murmurs “hmm” before turning away, without apology, to heat his food in the microwave. Fifteen long seconds tick by in silence. He finally turns back to me and states, “The average person is at Pangram because they care about the mission.”
To identify AI, Pangram employs a technique called “synthetic mirroring,” where it compares human writing to closely generated matches from LLMs. This process educates its model on AI writing styles. Pangram also utilizes “hard negative mining,” scanning datasets for false positives that can be synthetically mirrored to enhance its training set—essentially using errors for retraining. “All of our datasets are properly licensed, which I think is somewhat uncommon in the AI sphere today,” Spero mentions, though this is partly due to Pangram’s product being significantly less data-intensive than that of ChatGPT or Claude. “I don’t want to completely disparage the AI companies,” he adds, “but I believe they’ve lost considerable trust, particularly among creatives.” Currently, Pangram serves multiple sectors, including education, law, and recruitment, yet creative writing represents the largest portion of their training text.
The Shy Girl narrative put Pangram in the spotlight. Spero had already been highlighting suspected AI-generated content on his social media, so when he was tagged in a Reddit thread in January and subsequently received a PDF of Ballard’s manuscript, it came as no surprise. “I put it into [Pangram], posted it, and later The New York Times approached me for a comment,” he recounts. “I think people exaggerate the significance of Pangram in the Shy Girl saga.” However, that’s not exactly the full story. In a game of AI-scandal telephone, a Pangram account executive discussed the issue with a publishing industry analyst, who then brought it to the Times’ attention.
Another twist: Critics of Pangram, including the investigative project The Drey Dossier, pointed out that Spero’s manuscript came from a pirating website. When I bring this up, Spero admits, “Yeah. And, like, yeah … it is what it is. I hadn’t examined it closely. I didn’t read the entire PDF. I just input it directly into Pangram.”
Since then, Pangram hasn’t shied away from controversy. After a Commonwealth Short Story Prize winner received a high Pangram score, the company reviewed all winners since 2012, identifying three additional instances of potential AI use. Still, Spero minimizes Pangram’s influence, especially regarding canceled book deals. “Essentially, to my knowledge, every book deal hasn’t really revolved around the Pangram score,” he explains. “It factors into it, but if you speak with anyone involved, it’s just a small, small part of the larger picture.”
We shift our focus to Pangram’s Substack integration. “People shouldn’t fear disclosing this because your work should be valued for its quality and its appeal to readers, regardless of … ” Spero trails off. “Well, how do I want to phrase that?” He pauses, then continues with renewed confidence and quicker speech. “If the worth of your work hinges on misleading the end user into believing it wasn’t created by AI, then”—he hesitates again—“that’s going to be an issue.” (Later, I realize his tone changed because he was quoting his own post on X.)
