Kernel fusion emerges as the shared answer to running trillion-parameter models
An open-source team and a vendor both bet on kernel fusion to run Kimi K3 within one week; all figures are self-reported and await third-party replication.
Two teams, one in China and one abroad, sped up the 2.8-trillion-parameter Kimi K3 within the same week using kernel fusion.
Per Zhidx reporting, on September 21 Inspur released the SD200 Ultra supernode, claiming it hosts K3 on a single machine at 5.85 ms per token; on September 23 the team Inferact open-sourced tpu-megakernels, claiming 709 tokens/s decode throughput on 16 Google TPU v7 chips. All figures are self-reported.
Both merge operators to cut data movement. K3 officially recommends 64 accelerators and runs at roughly 10 tokens/s unoptimized; the claimed speedups still await third-party replication.
Sources:https://zhidx.com/p/598610.html