Multimodal Models Learn to Refuse Impossible Edits
New research introduces visual abstention; the VisTA method enables models to refuse 93% of infeasible drawing requests while maintaining high editing accuracy.
重要度重大証拠E2 未複製執筆簡易
Unified multimodal models (UMMs) can now identify impossible drawing instructions and actively refuse them, rather than generating erroneous images.
Previously, even the strongest editing models refused only 0.4% of logically infeasible modification requests, often hallucinating non-existent objects or silently altering the prompt. Explicitly prompting models to report infeasibility increased refusals but reduced editing accuracy.
The VisTA training method proposed by Chufan Shi et al. pairs feasible and infeasible examples, forcing the model to judge feasibility before generating output. The resulting VisTA-BAGEL model refuses 93.0% of infeasible requests without reminders, while completing 74.3% of feasible edits—outperforming all eight evaluated mainstream models.
These results rely on the authors' self-created Draw-or-Decline benchmark and have not yet been independently reproduced.