News1 min read·
TaSQ Boosts 1-Bit Cache Throughput
New algorithm increases 1-bit KV cache batch size by 14x on single GPU.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.
The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.
Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.