Back to feed
arXiv cs.CL
arXiv cs.CL
7/23/2026
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Short summary

This paper investigates hypernetworks for train-time knowledge injection into LLMs, where a hypernetwork generates a fixed LoRA adapter enabling the target model to answer questions about a given fact corpus. The authors construct MegaWikiQA with tens of millions of multi-hop QA examples across 39 domains and characterize scaling laws along hypernetwork depth, width, and target model size. Results show predictive power-law scaling and superior OOD generalization compared to standard LoRA and full fine-tuning.

  • Hypernetworks generate LoRA adapters for train-time factual knowledge injection at scale
  • First empirical scaling laws for hypernetwork architectures along depth, width, and target size
  • Hypernetworks show steeper OOD scaling exponents than LoRA or full fine-tuning

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more