Near-domain negatives drop fraud-detection classifiers from perfect Macro-F1 to 0.65–0.68
TeleAntiFraud 2.0 freezes a monthly benchmark of 900 Chinese phone calls; authors self-report Macro-F1 falling from perfect to 0.65–0.68 on near-domain negatives.
중요도국소적증거E2 미복제
TeleAntiFraud 2.0 freezes a monthly benchmark of 900 Chinese phone-call audios (600 fraud, 300 near-domain non-fraud), letting readers test how fraud-detection classifiers really handle near-domain negatives. Previously, on unrelated or ordinary negatives, classifiers reached perfect Macro-F1, masking their failure on same-context samples.
The authors self-report that in controlled text experiments, three classifiers dropped to Macro-F1 of 0.65–0.68 once near-domain negatives were swapped in; full-audio and ASR+LLM evaluations also revealed class-prior shortcuts, prediction collapse, and snapshot sensitivity. Results are author self-reported.
Boundary: results are self-reported with no third-party replication; submitted to arXiv as a preprint by Huiyuan Liu, Zhiming Ma, and twelve other authors on September 16 (v2 updated September 17, id 2609.18748).
Sources: arXiv abstract page (2609.18748) ↗, authors' data and code repository ↗