AI Cross-Modal Alignment Needs No Paired Data
New research shows text and image models can achieve coarse cross-modal alignment without any paired samples by leveraging shared geometry.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Independently trained AI models can achieve coarse cross-modal alignment without seeing any paired examples.
Previously, connecting different modalities like text and images relied heavily on massive amounts of manually labeled paired data. However, the Platonic Representation Hypothesis suggests that models trained on different modalities may spontaneously converge toward a shared representation geometry.
The Wasserstein Procrustes method proposed by Schnaus et al. aligns two disjoint embedding sets by estimating a single orthogonal map. Experiments show that standard geometric metrics accurately predict when this unpaired alignment is possible; in very few-pair regimes, the method outperforms existing approaches.
This result comes from a preprint submitted on October 7. It is currently an author-run benchmark and has not yet been independently reproduced.