Blogger's own test: Jev monitors chain-of-thought 6x faster but misses more harm
Jev's output format cuts classification cost by orders of magnitude, but the self-run data shows it detects harmful traces worse than mainstream models.
Original event 2026-09-23
A LessWrong author's self-run benchmark reports that a Jev-format model monitoring chain-of-thought for harm averaged 542 ms per classification, about 6x faster than Claude Sonnet 5 at 3,348 ms, and cost $0.056 per 1,000 classifications versus $3.17 for Sonnet, roughly 566x cheaper.
This is a first-party benchmark the author ran on 2,200 traces from the ReasoningShield dataset; it has no independent reproduction. On exact-match accuracy Jev scored 71.5%, essentially tied with Sonnet's 71.4% and below GPT-5.6 Luna's 79.4%.
The key weakness is on the harmful side: the author himself notes Jev was worse than both Sonnet and Luna at identifying genuinely harmful traces, and that catching harm matters more than fewer false positives. The headline multipliers hold only for the harmless side.