Loongson ships first full software stack for its GPU, enabling direct ONNX deployment
Loongson completes the AI inference software stack for its in-house GPU with CUDA compatibility, but all performance claims are self-reported pending real-world validation.
Original event 2026-09-22
On September 22, Loongson Technology released the first software version of its Loongson Accelerated Computing Platform, targeting the LG200 GPU cores integrated in the 2K3000 and 9A1000 chips, covering drivers, compilers, operator libraries and an inference engine.
The platform supports both OpenCL 3.0 and CUDA programming interfaces, and uses its in-house LacInfer engine as an ONNX Runtime execution backend, so ONNX models exported from PyTorch or TensorFlow can be deployed without code rewrites. Operator libraries include assembly-level optimizations for FP32 and INT8 GEMM workloads.
The company says the software already serves early customers and underpins agent development for embodied devices on the 2K3000 and 9A1000. Note that inference speed and accuracy-loss claims are Loongson's own figures with no third-party benchmarks, and the 9A1000 graphics card is not expected on sale until the first half of next year per the company's earlier statements.