Context segmentation lets gemma-4 E4B solve 18.52% of picoCTF tasks standard agents miss
Small-model agents bloat their context through accumulated tool outputs on long-horizon CTF tasks; a two-level context segmentation framework lets gemma-4 E4B solve 18.52% of tasks standard execution could not.
ImportanceLocalEvidenceE2 unreplicated
On picoCTF tasks, the memory-constrained gemma-4 E4B, using a two-level "context segmentation" framework, solved 18.52% of tasks that standard agent execution failed to complete, while acting as an "intelligent search": rewards on par with brute-force retries and better token efficiency.
Previously, small-model agents on long-horizon CTF tasks suffered context bloat from accumulated tool outputs, making long-horizon tasks hard to complete.
The result was self-reported by Sebastiano Nordio and Michele Lotto, measured on the picoCTF dataset with the gemma-4 model (authors' claim, first-party testing).
It has no independent verification yet; the preprint was submitted September 11, updated to v3 on September 15 (arXiv:2609.12839), and accepted to the non-archival ESORICS 2026 RAISE workshop.