G2G freezes the base, trains ~32M params, self-reports SOTA on four datasets
G2G freezes the MapAnything base and trains only a resampler, cross-group bridge and multi-frame pose head (~32M params); authors self-report SOTA on four datasets for relocalization and rig odometry.
重要度局所的証拠E2 未複製
With the MapAnything base frozen, G2G trains only three modules — a resampler, a cross-group bridge and a multi-frame pose head — about 32M parameters, under 6% of the full model, to estimate relative 6-DoF pose between two groups of images, using only relative pose supervision.
Previously, the usual approach was to fine-tune the whole model, which costs more training and risks eroding the base's general capabilities.
The authors self-report that on four datasets — indoor/outdoor simulation, real cross-season, and zero-shot sim-to-real — cross-sequence relocalization and multi-camera rig odometry accuracy reach what they claim is state of the art; code and weights are open-sourced.
The results are not independently reproduced; the work is by Yufei Wei and seven others, as an arXiv preprint (2606.08284) first posted June 6 and updated to v3 on September 17.