One round of skill-graph self-evolution lifts retrieval reward from 52.4% to 59.4%
SE-GoS authors' own tests show one evolution round lifting SkillsBench retrieval reward from 52.4% to 59.4%, pending independent replication.
ImportanciaLocalEvidenciaE2 no replicadoAnálisisRápido
Maintaining a skill retrieval graph from execution traces alone, one round of self-evolution lifts agent skill-library retrieval reward from 52.4% to 59.4% — with no model training, no retrieval-algorithm changes, and no separate model judging which skills relate.
Previous approaches relied on fixed indexing such as full-library loading or vector retrieval, which cannot automatically adjust the graph's connections and node descriptions from execution records.
The SE-GoS authors' own tests show one evolution round lifting average reward on SkillsBench from 52.4% to 59.4%, above full-library loading and vector retrieval baselines; on a held-out split it never saw, from 52.9% to 58.3%, at about two-thirds the input tokens of loading the full library. Repeating the round adds nothing; the abstract does not explain why.
The results are self-reported by the authors with no independent replication yet; the method is by Dawei Fu and four coauthors, updated on arXiv on October 1.