뉴스1 분 소요·
TaSQ Boosts 1-Bit Cache Throughput
New algorithm increases 1-bit KV cache batch size by 14x on single GPU.
중요도국소적증거E2 미복제작성 방식간략
The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.
The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.
Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.