iFlytek launches a speech recognition model that turns spoken language into usable text
iFlytek released Spark-ASR-2.0, trained entirely on domestic chips, targeting context-aware correction and fluent transcription; performance figures are vendor-reported.
ImportanceLocalEvidenceE3 inspectableWrite-upQuick
iFlytek released its Spark-ASR-2.0 speech recognition model on September 23, trained entirely on domestically made chips. It is already live in the iFlytek input method and available via API.
The model runs a fast non-autoregressive pass first, then uses speculative decoding with LLM-based understanding to fix errors. The company says it corrects homophones, removes filler words and self-corrections, and turns speech into text ready to use as-is, with overall inference cost up only 10% over Spark-ASR-1.0.
Hands-on testing by tech outlet Zhidx found stable transcription across dialects, subway noise, whispered speech and mixed Chinese-English terminology, though the samples were self-selected. The vendor claims word error rates below the industry's best in most scenarios, but that comparison lacks public benchmark detail, and real-world gains remain to be verified.