LLoCoT Preprint: Parallel Latent Reasoning Cuts Code Generation First-Token Latency by ~36x
New framework iteratively refines latent states to drastically reduce reasoning latency while maintaining accuracy.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
The LLoCoT framework reduces the time to the first answer token by approximately 36x compared to the explicit Chain-of-Thought baseline (Reasoning SFT) on HumanEval and MBPP benchmarks.
Traditional CoT relies on autoregressive generation of intermediate steps, causing high latency; existing latent methods retain left-to-right dependencies. LLoCoT introduces a looped transformer that jointly updates a compact latent workspace through few iterations, sampling latent tokens in parallel to condition the decoder.
Authors report that the method matches Reasoning SFT in accuracy while increasing end-to-end throughput by 9.2%. These results are from a preprint and have not yet been independently reproduced.