DBM's Statistical Fusion Edge Comes From Cross-Block Conditioning
Single-author preprint self-report: DBM's advantage in statistical data fusion comes almost entirely from cross-block conditioning, not generative pretraining.
ImportanceLocalEvidenceE2 unreplicated
In statistical data fusion (two samples sharing covariates, disjoint outcome blocks), the advantage of a fine-tuned deep Boltzmann machine (DBM) comes almost none from generative pretraining and instead from cross-block conditioning when predicting the other outcome block: the author reports this contribution is positive across two datasets and 40 experimental cells, with magnitudes of +0.19 and +0.36 percentage points.
Previously, attributing the DBM's advantage to generative pretraining would misidentify its source; the author reports that against a tuned baseline with the same conditioning, the DBM is best in 37 of 40 cells.
The author self-reports that imputers capable of cross-block conditioning mostly lose accuracy because of it, while the DBM benefits in every cell.
The results have not been independently verified; single-author preprint, submitted September 14, updated to v2 on the 16th, arXiv:2609.14934.