Pangram's own tests show 4% and 7% false positives; site showed only 0.01%-order
Pangram 4's self-tests misclassified AI-assisted text at up to 7%, while the site previously showcased only 0.01%-order results; wording changed after an inquiry.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
In Pangram 4's own tests, AI-assisted text was misclassified as AI-Generated at rates of 4% and 7%, while the website previously showcased only the best 0.01%-order result — a gap LessWrong author Sonia Albrecht named in a September 23 post.
She says that before she contacted the company on September 17, the site displayed "99.9%+ Accuracy" directly beside "Detects AI Assistance"; the wording has since changed, though she does not know if her message caused it. In her own test, about 30% of passages in roughly 5,000 words of her writing were labeled AI-Generated.
All of this is the critic's account and personal testing; the 4% and 7% figures are not independently verified, and Pangram's paper abstract claims an overall false positive rate of 0.0041%. False positives rise on short passages, which matters for anyone using AI editing.