Paper claims decoupled latents with coupled updates lift contact accuracy in human-object interaction
Authors' own tests top contact metrics on InterAct, but margins and reproduction conditions are undisclosed; treat as a preprint.
A preprint splits body, object and hand motion into separate latents and couples their updates, with the authors' own tests topping contact metrics among compared methods.
The paper proposes TRACE, which encodes body, object and hand motion into separate latents, predicts each stream's velocity from the complete interaction state, and applies geometric losses on decoded motion to constrain contact and object-relative movement. The same model can complete any single missing stream from the other two.
The authors report on InterAct, OMOMO and BEHAVE that joint completion training improves generation and that frozen flow features improve interaction understanding. The abstract states that on InterAct, TRACE achieves the highest contact precision, recall and F1 among the compared methods.
The abstract gives no margin of improvement, hardware or training scale, and no third-party reproduction is noted, so these are the authors' own results. The paper was submitted on 26 September and revised to v2 on 29 September.
Sources:https://arxiv.org/abs/2609.32551