arXiv cs.LG
7/22/2026

Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority
Short summary
Researchers present a preference-conditioned RL controller for transit signal priority that can be tuned at runtime to balance bus-priority versus overall traffic delay without retraining. Built on IntersectionZoo, the single learned policy outperforms fixed-time and rule-based baselines across a smooth trade-off frontier. Tail-delay diagnostics show non-bus externalities stay limited at moderate preference settings but rise sharply under high bus-priority weights.
- •Preference-conditioned RL policy for transit signal priority tunable at runtime via parameter w
- •Outperforms fixed-time, rule-based TSP, and fixed-weight PPO baselines on IntersectionZoo
- •Non-bus traffic delays remain limited at moderate settings but increase under high bus-priority weights
Generated with AI, which can make mistakes.
Is this a good recommendation for you?