arXiv CS.AI
A

arXiv CS.AI

arXiv CS.AI publishes articles and insights about AI, technology, and industry trends.

Profile generated by AI for Anything

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

21h

ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models

ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models

21h

PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs

PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs

21h

AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics

AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics

21h

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

21h

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

21h

Benchmarking the Personalization Capabilities of Large Language Models

Benchmarking the Personalization Capabilities of Large Language Models

21h

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

21h

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

21h

The original title is "JAXBench: Benchmarking Autonomous TPU Kernel Optimization"

The original title is "JAXBench: Benchmarking Autonomous TPU Kernel Optimization"

21h

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience

1d

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

1d

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

1d

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

1d

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

1d

Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems

Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems

1d

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

1d

Information Discernment in Large Language Models

Information Discernment in Large Language Models

1d

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

1d

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation

1d

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

2d

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

2d

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

2d

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance

2d