ONNX SmolVLA: Latency Halved but Spatial Success Falls to 41%
SmolVLA on ONNX cuts p99 latency from 1181 ms to 601/532 ms, but LIBERO Spatial success drops from 70% to 41%, per a single-author self-report.
중요도국소적증거E2 미복제
Deploying SmolVLA with ONNX Runtime on an RTX 2060 can cut p99 latency from 1181 ms to 601/532 ms, but in LIBERO simulation Spatial success drops from 70% to 41%, while Object stays at 89%.
Previously, anyone trying to cut this model's inference latency lacked latency-versus-success data for an ONNX deployment path, making it easy to misjudge capability retention after compression; an audit also found the INT8 export was actually an FP32 graph.
Single-author preprint self-report: evaluated in LIBERO simulation, Spatial success 70%→41%, Object stays at 89%, p99 latency 1181 ms→601/532 ms; the author says a language width of 24 restores it to 75% at roughly half the baseline latency, but width does not explain all of the difference.
The result is limited to an RTX 2060 and LIBERO simulation and has not yet been reproduced by others; the preprint was submitted on September 12 and updated on the 16th.