Synthetically translated corpus suffices for French biomedical pretraining
Authors released a 36.4GB French life-sciences corpus and model; performance is self-reported, worth tracking as a low-resource method signal.
Original event 2026-09-22
The authors report that pretraining on purely machine-translated text reaches state-of-the-art results on French life-sciences downstream tasks.
The paper is accepted to EMNLP 2025 Findings, and the authors released the TransCorpus toolkit, a 36.4GB French life-sciences corpus, the pretrained TransBERT-bio-fr model, and reproducible code.
Caveat: all performance claims come from the authors' own evaluations with no independent reproduction; for most readers this is a method signal for low-resource language/domain modeling, not a verified general result.