Back to feed
AR
arXiv CS.AI
7/23/2026
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Short summary

NEXUS is a structured-plan safety monitor for tool-using LLM agents that applies formal intervention policies across four actions: allow, block, request confirmation, or request revision. It combines deterministic rules, argument-level inspection, and a calibrated logistic-regression risk scorer. On benchmarks it achieves F1=0.949 with only 0.205ms median latency, adding under 0.1% overhead to agent loops. Code and benchmarks are publicly released.

  • NEXUS monitors tool-using LLM agents with four-tier intervention: allow, block, confirm, revise
  • Achieves F1=0.949 on synthetic benchmark, outperforming rule-only by 27.3 percentage points
  • Sub-0.1% overhead at 0.205ms median latency; code and benchmarks publicly released

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more