Back to feed
AR
arXiv CS.AI
6/30/2026
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

Short summary

COMPASS is the first unified multimodal framework grounding composition intent in both perception and generation through a shared expert token. It includes Comp-11, a large-scale dataset with 11 composition categories and reasoning annotations. Experiments demonstrate substantial improvements in composition understanding and controllable image generation.

  • COMPASS unifies composition perception and generation via shared expert token steering
  • Introduces Comp-11 dataset with 11-class composition taxonomy and reasoning annotations
  • Shows significant improvements in composition understanding and controllable image generation

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more