Noticias1 min de lectura·
TaSQ Boosts 1-Bit Cache Throughput
New algorithm increases 1-bit KV cache batch size by 14x on single GPU.
ImportanciaLocalEvidenciaE2 no replicadoAnálisisRápido
The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.
The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.
Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.