arXiv cs.CL
7/22/2026

Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
Short summary
Search-on-Graph-R1 trains an 8B model to navigate knowledge graphs for question answering using SFT followed by reinforcement learning. The approach scaffolds a frontier teacher with gold SPARQL queries so trajectories are grounded in a live Freebase graph. The 8B model surpasses all frozen frontier-LLM systems on WebQSP, CWQ, and GrailQA, achieving top CWQ results without auxiliary inference modules or LLM judges during training.
- •8B model trained via SFT+RL surpasses frozen frontier LLMs on three KGQA benchmarks
- •Teacher scaffolding with gold SPARQL queries produces grounded training trajectories on live Freebase
- •SFT and RL contribute complementary gains; RL learns to reach answers in fewer Search calls
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
