Inspur launches domestic-chip supernode, claims single node runs 2.8T-parameter model
Inspur announced the SD200 Ultra supernode at AICC2026, claiming single-node hosting of Kimi K3; all performance figures are vendor claims pending third-party testing.
Original event 2026-09-21
Inspur announced the Yuannao SD200 Ultra supernode AI server at AICC2026 on September 21, built on domestic AI chips. The company says a single node can host the 2.8-trillion-parameter Kimi K3 model and supports frontier models up to 10 trillion parameters.
The system tightly couples 128 domestic AI chips with 8TB of unified-addressable memory and 64TB of system memory. Inspur claims token generation latency below 5.85ms, equivalent to 170 tokens/s per user and five times the industry average, plus a 3.5x reduction in AllReduce communication time.
Inspur also launched the HC2000 compute unit the same day, claiming 10x token throughput per unit of investment. Note that all performance figures are Inspur's own claims; the 'industry average' baseline is unspecified, and real-world results await third-party testing.