ニュース1 分で読める·
TaSQ Boosts 1-Bit Cache Throughput
New algorithm increases 1-bit KV cache batch size by 14x on single GPU.
重要度局所的証拠E2 未複製執筆簡易
The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.
The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.
Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.