AR
arXiv CS.AI
6/30/2026

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Short summary
COMPASS is the first unified multimodal framework grounding composition intent in both perception and generation through a shared expert token. It includes Comp-11, a large-scale dataset with 11 composition categories and reasoning annotations. Experiments demonstrate substantial improvements in composition understanding and controllable image generation.
- •COMPASS unifies composition perception and generation via shared expert token steering
- •Introduces Comp-11 dataset with 11-class composition taxonomy and reasoning annotations
- •Shows significant improvements in composition understanding and controllable image generation
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

