VLA adaptation cost can be diagnosed then allocated: 0.04% of parameters matches full fine-tuning on xArm-7
Authors self-report: 10 unlabeled observations diagnose VLA adaptation cost (median Spearman 0.91), budget-allocated rank-variable LoRA matches full fine-tuning on a physical xArm-7 with 0.04% trainable parameters.
ImportanciaLocalEvidenciaE2 no replicado
The adaptation cost of VLA models is structurally distributed — appearance changes concentrate in the vision encoder, instruction changes in the language backbone — so adaptation cost can be diagnosed first and then allocated by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters.
Previously, practitioners fine-tuned VLA models in full without knowing where the cost concentrates, paying unnecessary training overhead.
Authors self-report: their pipeline first uses 10 unlabeled goal observations to diagnose the cost of each region (median Spearman 0.91), then allocates rank-variable LoRA adapters by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters. Results are self-reported by Shahram Najam Syed, Jeffrey Ichnowski and two other authors.
Not yet reproduced by third parties; the preprint was submitted to arXiv on September 16 (v2 updated on the 17th, arXiv:2609.18084).