FRAUDSkill freezes weights, tunes external skills only, self-reported 73.50% Macro-F1 on TeleAntiFraud
FRAUDSkill keeps the audio language model's weights frozen and optimizes only external skill procedures, routing policies and decision rules, with authors self-reporting 73.50% Macro-F1 on TeleAntiFraud.
BedeutungLokalBeweisE2 nicht repliziert
FRAUDSkill keeps the underlying audio language model's weights frozen and optimizes only external skill procedures, routing policies and decision rules; the authors self-report 73.50% Macro-F1 on the TeleAntiFraud benchmark, 31.96% higher than the shared frozen-model baseline, with invalid outputs down to 1.94%.
Previously, fine-tuning the model itself was required, which is costly, and the shared frozen-model baseline performed markedly worse with more invalid outputs.
The result was self-reported by Chengxian Hu and 11 co-authors — first-party data with no third-party replication yet.
Boundary: results are limited to the TeleAntiFraud benchmark; the preprint was submitted to arXiv on September 16 (updated to v2 on the 17th, id 2609.18766), with an anonymous code link provided by the authors.