EdiTikZ self-reported: 9B beats GPT-5.6-Sol in human eval
A Qwen3.5-based 9B model for editing scientific TikZ figures, self-reported by its authors to beat GPT-5.6-Sol and match Gemini-3.1-Pro in a human eval with 9 raters and 4,320 ratings.
ImportanceLocalEvidenceE2 unreplicated
EdiTikZ, 4B/9B models built on Qwen3.5, can edit scientific TikZ figures; the authors self-report that in a human evaluation with 9 raters and 4,320 ratings, the 9B model scored above GPT-5.6-Sol and on par with Gemini-3.1-Pro.
No dedicated model for scientific TikZ figure editing existed before; this work mined 391,000 pairs of TikZ revisions from arXiv, GitHub and TeX SE and inferred 781,000 editing instructions for training.
The measurement is a self-reported human evaluation: 9 raters, 4,320 ratings in total, with the 9B model above GPT-5.6-Sol and comparable to Gemini-3.1-Pro; the results are not peer-reviewed.
Boundary: covers only scientific TikZ figure editing, with no third-party replication; the preprint was first posted September 1 and updated September 14 (arXiv:2609.01409), models and data are released, and the full code is said to be coming soon.