| 2026-09-05 |
Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving |
| 2026-09-05 |
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents |
| 2026-09-05 |
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs |
| 2026-09-05 |
ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize |
| 2026-09-05 |
SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center |
| 2026-09-04 |
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems |
| 2026-09-04 |
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining |
| 2026-09-04 |
Post-Training Language Models for Gold-Medal Performance in Coding Competitions |
| 2026-09-04 |
Untangling the Mechanisms of Misleading Context in Medical Question Answering |
| 2026-09-04 |
Language Models Can Control Their Own Attention |
| 2026-09-03 |
When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models |
| 2026-09-03 |
How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling |
| 2026-09-03 |
Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis |
| 2026-09-03 |
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI |
| 2026-09-03 |
Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling |
| 2026-09-01 |
Memory-First Fact-Checking: A Knowledge-Graph-Grounded Multi-Agent System for Misinformation Detection |
| 2026-09-01 |
Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory Consistency |
| 2026-09-01 |
Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents |
| 2026-09-01 |
Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security |
| 2026-09-01 |
Cross-lingual Functional Vectors for Emotion Detection in Large Language Models |
| 2026-08-31 |
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses |
| 2026-08-31 |
COVER: Identifiable Evaluation of Coalition Routing |
| 2026-08-31 |
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge |
| 2026-08-31 |
When Verified Source Becomes Attack Input: Defending Smart Contracts Against LLM-Based Vulnerability Scanning |
| 2026-08-31 |
CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents |
| 2026-08-30 |
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution |
| 2026-08-30 |
Boosting LLM Exploration via Weak-Model Guidance in RLVR |
| 2026-08-30 |
When Context Gets Root: Privilege Escalation in LLM Harnesses |
| 2026-08-30 |
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City |
| 2026-08-30 |
RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature |
| 2026-08-29 |
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit |
| 2026-08-29 |
RCMN: Understanding Misleadingness in Influential Public Discourse |
| 2026-08-29 |
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO |
| 2026-08-29 |
Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification |
| 2026-08-29 |
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting |
| 2026-08-28 |
Puro-2B: Poor Lab’s Qwen2-1.5B Trained on RTX 5090 within $5090 |
| 2026-08-28 |
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution |
| 2026-08-28 |
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes |
| 2026-08-28 |
How Language Models Organize and Structure Moral Knowledge |
| 2026-08-28 |
CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases |
| 2026-08-27 |
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks |
| 2026-08-27 |
ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs |
| 2026-08-27 |
AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs |
| 2026-08-27 |
PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology |
| 2026-08-27 |
Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems |
| 2026-08-25 |
Prime Agent: A Self-Improving RLM Harness |
| 2026-08-25 |
MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters |
| 2026-08-25 |
On the Threat Model of Weird Generalization and Emergent Misalignment |
| 2026-08-25 |
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards |
| 2026-08-25 |
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems |
| 2026-08-22 |
Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents |
| 2026-08-22 |
BreakGuard: Towards Detecting Dependency Breaking Changes with LLM-Generated Tests |
| 2026-08-22 |
MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use |
| 2026-08-22 |
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction |
| 2026-08-22 |
RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation |
| 2026-08-21 |
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection |
| 2026-08-21 |
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization |
| 2026-08-21 |
Pandora’s AI Model Routing Box: Efficient Allocation with Costly Value Estimation |
| 2026-08-21 |
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement |
| 2026-08-21 |
MidTool: Mid-training Data Synthesis for Agentic Tool Use |
| 2026-08-20 |
SPADE: Self-Play in Adaptive Synthetic Executable Environments |
| 2026-08-20 |
What is Missing from AI Post-Training AI: An Empirical Analysis |
| 2026-08-20 |
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets |
| 2026-08-20 |
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning |
| 2026-08-20 |
Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering |
| 2026-08-19 |
Recirculation |
| 2026-08-19 |
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering |
| 2026-08-19 |
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach |
| 2026-08-19 |
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating |
| 2026-08-19 |
Chain-of-Experience for Continual LLM Improvement |
| 2026-08-18 |
Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models |
| 2026-08-18 |
HarnessEval-W: Agentifying the Evaluation of Visual Worlds |
| 2026-08-18 |
ClawGym II: Exploring Black-Box RL on Agent Harness |
| 2026-08-18 |
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments |
| 2026-08-18 |
LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing |
| 2026-08-15 |
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist |
| 2026-08-15 |
Vero: Can AI Agents Build Formally Verified Software Repositories? |
| 2026-08-15 |
TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval |
| 2026-08-15 |
CAPRI: Contract-Aware Proof Repair for Isabelle |
| 2026-08-15 |
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives |
| 2026-08-14 |
LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles |
| 2026-08-14 |
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data |
| 2026-08-14 |
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination |
| 2026-08-14 |
Intern-S2-Preview: Scientific Agentic Foundation Model |
| 2026-08-14 |
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design |
| 2026-08-13 |
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation |
| 2026-08-13 |
V-FiLLM: Verified Financial LLM Reasoning Benchmark |
| 2026-08-13 |
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment |
| 2026-08-13 |
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents |
| 2026-08-13 |
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization |
| 2026-08-12 |
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness |
| 2026-08-12 |
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents |
| 2026-08-12 |
Stealing Reasoning Traces from Proprietary LLM APIs |
| 2026-08-12 |
Towards Expert-level Medical AI for Real-time Video Consultations |
| 2026-08-12 |
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA |
| 2026-08-11 |
KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs |
| 2026-08-11 |
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models |
| 2026-08-11 |
Activation Probes Surface Code-Security Signals that the Model’s Output Misses |
| 2026-08-11 |
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models |
| 2026-08-11 |
Defining Decentralization: An Ontological Perspective |
| 2026-08-09 |
RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer |
| 2026-08-09 |
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints |
| 2026-08-09 |
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models |
| 2026-08-09 |
MASS: Multiplayer World Models with Authoritative Shared State |
| 2026-08-09 |
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation |
| 2026-08-08 |
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images |
| 2026-08-08 |
TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN |
| 2026-08-08 |
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data |
| 2026-08-08 |
NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering |
| 2026-08-08 |
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction |
| 2026-08-07 |
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization |
| 2026-08-07 |
A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance |
| 2026-08-07 |
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks |
| 2026-08-07 |
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories |
| 2026-08-07 |
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering |
| 2026-08-06 |
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition |
| 2026-08-06 |
Chained Recursive Language Models for Multi-Iteration Reasoning |
| 2026-08-06 |
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding |
| 2026-08-06 |
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking |
| 2026-08-06 |
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning |
| 2026-08-05 |
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training |
| 2026-08-05 |
ACEM: A Cost Estimation Model for Agentic Software Engineering |
| 2026-08-05 |
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents |
| 2026-08-05 |
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation |
| 2026-08-05 |
Fast and Accurate Quotation Attribution in Literary Texts |
| 2026-08-05 |
Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning |
| 2026-08-05 |
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use |
| 2026-08-05 |
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning |
| 2026-08-05 |
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents |
| 2026-08-05 |
MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents |