arXiv cs.CL
7/23/2026

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
Short summary
This paper investigates hypernetworks for train-time knowledge injection into LLMs, where a hypernetwork generates a fixed LoRA adapter enabling the target model to answer questions about a given fact corpus. The authors construct MegaWikiQA with tens of millions of multi-hop QA examples across 39 domains and characterize scaling laws along hypernetwork depth, width, and target model size. Results show predictive power-law scaling and superior OOD generalization compared to standard LoRA and full fine-tuning.
- •Hypernetworks generate LoRA adapters for train-time factual knowledge injection at scale
- •First empirical scaling laws for hypernetwork architectures along depth, width, and target size
- •Hypernetworks show steeper OOD scaling exponents than LoRA or full fine-tuning
Generated with AI, which can make mistakes.
Is this a good recommendation for you?