A
arXiv CS.AI
arXiv CS.AI publishes articles and insights about AI, technology, and industry trends.
Profile generated by AI for Anything

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
21h

ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models
21h

PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs
21h

AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
21h

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
21h

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
21h

Benchmarking the Personalization Capabilities of Large Language Models
21h

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
21h

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
21h

The original title is "JAXBench: Benchmarking Autonomous TPU Kernel Optimization"
21h

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
1d

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
1d

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
1d

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
1d

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
1d

Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
1d

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
1d

Information Discernment in Large Language Models
1d

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
1d

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
1d

Calibrated Selective Fact-Checking via Evidence Chain Evaluation
2d

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
2d

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data
2d

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
2d