Real citations still fail to prove novelty
NovGauge tested 18 models: even with real, correctly-directed citations, over 70% of positive evidence failed to support novelty.
Even when citations are real and their conclusions point the right way, more than 70% of positive evidence still fails to logically support a paper's claim of novelty. NovGauge tested novelty judgments across 18 models using 619 paper pairs and 50 multi-paper sets; the authors report hallucination rates between 0% and 39%, with best Verified F1 around 43% to 72%.
Previously, novelty judgments relied on whether citations were real and whether their conclusions pointed the right way, but that standard misses a large share of positive evidence that does not hold up logically.
The measurement is author-reported: 619 paper pairs, 50 multi-paper sets, 18 models, hallucination rates 0% to 39%, best Verified F1 around 43% to 72%.
This is an author-built preprint benchmark that has not yet been independently reproduced; it is suited to testing whether citations support conclusions, and cannot be extrapolated to the accuracy of all AI literature reviews.