Fathom adapts K-cache bits per query, 1.67x faster decode step
Fathom lets each query adaptively choose how many bits of the 4-bit K cache to read, making Qwen3-8B million-token decode-step GPU time 1.67x faster than 136-bit sparse scanning, self-reported by the author
BedeutungLokalBeweisE2 nicht repliziert
In the million-token setting, each query can now adaptively decide how many bits of the 4-bit K cache to read: the author self-reports that Qwen3-8B decode-step GPU time is 1.67x faster than 136-bit sparse scanning, with the KV cache and index residing in host memory.
Previously, sparse scanning read the K cache at a fixed bit width and could not adjust the read amount per query.
The comparison is the author's own measurement: timing used synthetic KV, and in a real coding-agent session 92 bits already matched the step consistency of the most accurate scan; there is no speedup when the index is in GPU memory.
Single-author arXiv preprint (2609.17652), submitted September 15, updated to v2 on the 17th, not independently verified.
Sources: arXiv preprint ↗, author's code repository ↗