Preprint: Hypernetwork Generates LoRA Adapters for 284B Model, Boosting Accuracy to 84.9%
New hypernetwork maps context to parameters, achieving high-accuracy adapter generation on a 284B model.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Key Finding: A hypernetwork named Internalizer generates document-specific LoRA adapters for the 284-billion-parameter DeepSeek v4 Flash, achieving 84.9% top-1 teacher-forced accuracy on unseen documents compared to 63.4% for the base model.
Background: Previous research on mapping context directly to LoRA adapters was limited to base models of up to 14 billion parameters, restricting scalability.
Conclusion: Most of the hypernetwork's parameters reside in a model-agnostic trunk, requiring only thin entry and exit layers for portability. Once trained, a single forward pass converts any document into an adapter, which can be served alone for speed or alongside the document window for higher accuracy.
Boundary: Data is from author-run tests in a preprint (arXiv:2610.11715) using teacher-forced evaluation; independent reproduction of actual generation quality is pending.