arXiv CS.AI
A

arXiv CS.AI

arXiv CS.AI publishes articles and insights about AI, technology, and industry trends.

Profile generated by AI for Anything

An AI agent for treatment reasoning over a biomedical tool universe

An AI agent for treatment reasoning over a biomedical tool universe

13d

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

13d

Data and Evaluation Closed-Loop for Model Capability Enhancement

Data and Evaluation Closed-Loop for Model Capability Enhancement

13d

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

13d

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

13d

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

13d

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

13d

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

13d

Recursive Self-Evolving Agents via Held-Out Selection

Recursive Self-Evolving Agents via Held-Out Selection

13d

Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

13d

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

14d

Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models

Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models

14d

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

14d

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

14d

When Does Personality Composition Matter for Multi-Agent LLM Teams?

When Does Personality Composition Matter for Multi-Agent LLM Teams?

14d

AI-Model Network: Concept, Current State and Future

AI-Model Network: Concept, Current State and Future

14d

Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents

Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents

14d

Understanding Rollout Error in Graph World Models

Understanding Rollout Error in Graph World Models

14d

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

14d

ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

14d

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

18d

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

18d

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values

18d

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

18d