Radar
Which research topics are accelerating, and where earlier papers have already landed in shipped models.
Papers this week
50
Trending topic
Agents
▼ 6%papers this window vs prior window
Papers linked to models
155
linked by the desk, past 90 days
Median days paper → model
—
lower is faster
papers this window by topic · Δ vs prior window · a paper can carry more than one topic
| Topic | Papers this window | Prior window | Δ |
|---|---|---|---|
| Agents | 115 | 122 | -6% |
| Multimodal | 34 | 29 | +17% |
| Evaluation & Benchmarks | 31 | 37 | -16% |
Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.
RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
arXiv:2610.11358AI-inferred · highOn the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
arXiv:2610.10833AI-inferred · highPlan-and-Patch: Diffusion Language Models for Agentic Planning
arXiv:2610.10786AI-inferred · highInternalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
| 23 |
| 15 |
| +53% |
| Training & Optimization | 19 | 22 | -14% |
| Safety & Alignment | 19 | 23 | -17% |
| Efficiency & Inference | 18 | 21 | -14% |
| Generative Models | 14 | 9 | +56% |
| Reinforcement Learning | 13 | 9 | +44% |
| Retrieval & RAG | 10 | 16 | -37% |
| Interpretability | 9 | 8 | +13% |
| Code Generation | 8 | 9 | -11% |
| Robotics & Embodied AI | 8 | 9 | -11% |
| Speech & Audio | 7 | 6 | +17% |
| Privacy & Security | 4 | 4 | 0% |
| Long Context | 3 | 2 | +50% |
| Language Understanding | 3 | 4 | -25% |
| Mixture of Experts | 2 | 3 | -33% |
| Computer Vision | 2 | 1 | +100% |
| Other | 139 | 179 | -22% |
Reasoning
Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
arXiv:2610.11384AI-inferred · highBalancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
arXiv:2610.11128AI-inferred · highWriting for the Reviewer: Defensive Writing in GPT Models
arXiv:2610.11355AI-inferred · highMultiWorldBench: Do Independently Controlled Views Describe One Shared World?
arXiv:2610.11723AI-inferred · highHarness Evolution Hits a Ceiling: When Weight Training Should Begin
arXiv:2610.11655AI-inferred · highTypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
arXiv:2610.11392AI-inferred · mediumMine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild
arXiv:2610.11328AI-inferred · highDivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
arXiv:2610.11317AI-inferred · high| Paper | Linked model | Relation | Confidence |
|---|---|---|---|
| RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation (2610.11358) | Qwen-Audio-3.1-Realtime | evaluates | 98% |
| On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents (2610.10833) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| Plan-and-Patch: Diffusion Language Models for Agentic Planning (2610.10786) | Qwen-Audio-3.1-Realtime | evaluates | 99% |
| Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models (2610.11715) | DeepSeek-V4-Flash-0731 | technique used | 95% |
| Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation (2610.11384) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL (2610.11128) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| Writing for the Reviewer: Defensive Writing in GPT Models (2610.11355) | GPT-6.1 Sol Ultrafast | evaluates | 80% |
| MultiWorldBench: Do Independently Controlled Views Describe One Shared World? (2610.11723) | Solaris | evaluates | 100% |
| Harness Evolution Hits a Ceiling: When Weight Training Should Begin (2610.11655) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models (2610.11392) | Jev | related | 55% |
| Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild (2610.11328) | DeepSeek-V4-Flash-0731 | evaluates | 100% |
| Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild (2610.11328) | Claude Haiku 5.5 | evaluates | 100% |
| Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild (2610.11328) | GPT-6.1 Sol Ultrafast | evaluates | 100% |
| DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition (2610.11317) | Qwen-Audio-3.1-Realtime | evaluates | 98% |
| When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction (2610.11123) | Claude Haiku 5.5 | evaluates | 95% |
| OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences (2610.11118) | GPT-6.1 Sol Ultrafast | evaluates | 100% |
| AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks (2610.11050) | GPT-6.1 Sol Ultrafast | evaluates | 100% |
| How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense (2610.11005) | GLM-5.3 | evaluates | 100% |
| How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense (2610.11005) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents (2610.10942) | Claude Haiku 5.5 | evaluates | 80% |
| StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents (2610.10942) | Qwen-Audio-3.1-Realtime | technique used | 100% |
| StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents (2610.10942) | DeepSeek-V4-Flash-0731 | evaluates | 100% |
| Does the Model Use the Feature? Separating Steering from Mechanism in LLMs (2610.07270) | Gemma 4 | evaluates | 90% |
| Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs (2610.07646) | Gemma 4 | evaluates | 98% |
| Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs (2610.07646) | Qwen-Audio-3.1-Realtime | evaluates | 98% |
| Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs (2610.07646) | Gemini 4 Argon | evaluates | 98% |
| Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs (2610.07646) | Sonnet 5.5 | evaluates | 98% |
| Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs (2610.07646) | GPT-6 Astra Ultrafast | evaluates | 98% |
| Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models (2610.07620) | GPT-6 Astra Ultrafast | evaluates | 100% |
| A Systematic Investigation of Bias in Large Language Models for Advertising Relevance (2610.07544) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| A Systematic Investigation of Bias in Large Language Models for Advertising Relevance (2610.07544) | GPT-6 Astra Ultrafast | evaluates | 100% |
| Offline AI Modules: Voice-First Offline Architecture, Hardware Reference Stack, Quantization and Benchmarking (2610.07026) | Gemma 4 | evaluates | 100% |
| Cascadia: Resident 975B MoE Inference on Eleven AI PCs (2610.07219) | Inkling | evaluates | 99% |
| Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite (2610.02826) | Qwen-Audio-3.1-Realtime | evaluates | 90% |
| iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD (2610.02815) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses (2610.02267) | Jev | evaluates | 100% |
| A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns (2610.02664) | GPT-6 Astra Ultrafast | evaluates | 100% |
| Labels Override Definitions in Jev-Style Typed Decision Models (2610.02586) | Qwen-Audio-3.1-Realtime | evaluates | 95% |
| What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute (2610.02491) | Qwen-Audio-3.1-Realtime | evaluates | 85% |
| PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation (2610.02982) | GPT-6 Astra Ultrafast | evaluates | 100% |
| Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning (2610.02715) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge (2610.02405) | Claude Sonnet | technique used | 90% |
| DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents (2610.02351) | Claude Sonnet | evaluates | 100% |
| DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents (2610.02351) | Qwen-Audio-3.1-Realtime | evaluates | 100% |
| Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models (2610.01222) | GPT-6 Astra Ultrafast | evaluates | 99% |
| Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models (2610.01222) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 99% |
| Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models (2610.01222) | Claude Sonnet | evaluates | 99% |
| Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model (2610.01006) | Jev | evaluates | 100% |
| VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks (2610.00972) | Claude Sonnet | evaluates | 98% |
| VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks (2610.00972) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 98% |
| Finding the Right Fit: Model-Harness Interactions across Agent Tasks (2610.00917) | Kimi K3 2.8T-A50B | evaluates | 95% |
| Finding the Right Fit: Model-Harness Interactions across Agent Tasks (2610.00917) | Claude Sonnet | evaluates | 95% |
| Finding the Right Fit: Model-Harness Interactions across Agent Tasks (2610.00917) | GPT-6 Astra Ultrafast | evaluates | 95% |
| OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework (2610.00912) | Claude Sonnet | evaluates | 88% |
| Kepler: Auditable World Models for ARC-AGI-3 (2610.00834) | GPT-6 Astra Ultrafast | evaluates | 88% |
| Kepler: Auditable World Models for ARC-AGI-3 (2610.00834) | Claude Sonnet | evaluates | 99% |
| Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents (2610.00613) | Qwen3.8-LiveTranslate | technique used | 98% |
| Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents (2610.00609) | Claude Sonnet | evaluates | 100% |
| Benchmarking Prompt Optimization of Large Language Models With Chess (2610.00416) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | technique used | 90% |
| When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents (2610.00372) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks (2610.00084) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 100% |
| Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs (2610.00047) | Qwen3.8-LiveTranslate | evaluates | 99% |
| Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness? (2610.00025) | Qwen3.8-LiveTranslate | evaluates | 100% |
| What Do Rationales Communicate? A Message-Intervention Study in Role-Specialized QA (2610.00018) | DeepSeek V4-Flash | evaluates | 95% |
| Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts (2610.00314) | DeepSeek V4-Flash | evaluates | 90% |
| Learning Process Rewards via Reasoning State Propagation (2609.39220) | Qwen3.8-LiveTranslate | evaluates | 95% |
| BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence (2609.39139) | Qwen3.8-LiveTranslate | evaluates | 100% |
| BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence (2609.39139) | GPT-6.1 Sol | evaluates | 100% |
| Consistent Plan-Act for Long-Horizon Agentic Tasks (2609.38891) | GPT-6.1 Sol | evaluates | 90% |
| Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions (2609.38869) | Qwen3.8-LiveTranslate | evaluates | 95% |
| When Context Changes: Understanding Update Failures in LLMs (2609.38866) | Qwen3.8-LiveTranslate | evaluates | 90% |
| OpenJev-RLCD: A Working RLCD Implementation (2609.38850) | Qwen3.8-LiveTranslate | evaluates | 98% |
| Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation (2609.38282) | Qwen3.8-LiveTranslate | evaluates | 100% |
| UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement (2609.38721) | Qwen3.8-LiveTranslate | evaluates | 99% |
| When Scientific Contradictions Are Lost in Translation (2609.38621) | Sonnet 5.5 | evaluates | 100% |
| When Scientific Contradictions Are Lost in Translation (2609.38621) | GPT-6.1 Sol | evaluates | 100% |
| Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization (2609.38757) | GPT-6.1 Sol | evaluates | 99% |
| VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation (2609.38512) | GPT-6.1 Sol | evaluates | 99% |
| VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation (2609.38512) | Gemini 4 Argon | evaluates | 99% |
| VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation (2609.38512) | Grok | evaluates | 99% |
| NAQD Env: A benchmark for selective withdrawal in language agents (2609.38460) | Qwen3.8-LiveTranslate | evaluates | 100% |
| MetaPersona: Task-Grounded Synthetic Populations from Empirical Social Science (2609.38392) | GPT-6.1 Sol | technique used | 80% |
| Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits (2609.38386) | Qwen3.8-LiveTranslate | evaluates | 100% |
| AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks (2609.38288) | Qwen3.8-LiveTranslate | related | 95% |
| EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents (2609.36746) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 95% |
| EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents (2609.36746) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents (2609.36746) | GPT-6.1 Sol | evaluates | 95% |
| EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents (2609.36746) | Qwen3.8-LiveTranslate | evaluates | 95% |
| MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development (2609.36679) | Qwen3.8-LiveTranslate | evaluates | 99% |
| Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents (2609.36611) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It (2609.36585) | Qwen3.8-LiveTranslate | evaluates | 99% |
| Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning (2609.36572) | Qwen3.8-LiveTranslate | evaluates | 99% |
| Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference (2609.36437) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Persona Dosing: Calibrated Activation Steering for Graded Trait Control (2609.36388) | Gemma 4 | evaluates | 100% |
| Persona Dosing: Calibrated Activation Steering for Graded Trait Control (2609.36388) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations (2609.36278) | GPT-6.1 Sol | evaluates | 100% |
| Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations (2609.36278) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations (2609.36278) | Gemma 4 | evaluates | 100% |
| More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev (2609.36154) | Jev | evaluates | 100% |
| SAGE: A Statistical Acceptance Gate for Self-Evolving Agents (2609.36043) | DeepSeek-V4-Flash-0731 | evaluates | 100% |
| Is Human-Readable Text Necessary for Effective LLM Fine-Tuning? (2609.35868) | Qwen3.8-LiveTranslate | evaluates | 95% |
| Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT (2609.35953) | GPT-6.1 Sol | evaluates | 98% |
| Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden? (2609.35897) | Open Dreamer | evaluates | 80% |
| Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training (2609.35890) | GPT-6.1 Sol | evaluates | 98% |
| FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models (2609.30954) | Qwen3.8-LiveTranslate | evaluates | 99% |
| FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models (2609.30954) | GPT-6 Astra | evaluates | 99% |
| LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information (2609.30706) | Jev | related | 90% |
| HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases (2609.30571) | Qwen3.8-LiveTranslate | evaluates | 98% |
| Pretrained ASR Pseudo-labeling for Noisy Police Audio (2609.30469) | Qwen3.8-LiveTranslate | evaluates | 98% |
| ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure? (2609.30325) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Externalized CPDAG Summaries Improve LLM Causal Deduction (2609.31071) | GPT-6 Astra | evaluates | 99% |
| Externalized CPDAG Summaries Improve LLM Causal Deduction (2609.31071) | Qwen3.8-LiveTranslate | evaluates | 99% |
| MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory (2609.30939) | GPT-6 Astra | evaluates | 99% |
| MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory (2609.30939) | Qwen3.8-LiveTranslate | technique used | 99% |
| Self-Play Search Distillation for Large Language Model Reasoning (2609.30936) | Qwen3.8-LiveTranslate | evaluates | 95% |
| ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning (2609.30796) | GPT-6 Astra | technique used | 95% |
| Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More (2609.30768) | Qwen3.8-LiveTranslate | evaluates | 100% |
| The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading? (2609.30705) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 99% |
| The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading? (2609.30705) | GPT-6 Astra | evaluates | 99% |
| The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading? (2609.30705) | DeepSeek-V4-Flash-0731 | evaluates | 99% |
| Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders (2609.31056) | Gemma 4 | evaluates | 95% |
| Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders (2609.31056) | Qwen3.8-LiveTranslate | evaluates | 95% |
| MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens (2609.30967) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 95% |
| Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents (2609.31430) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Semantic Navigation for Issue Localization in Code Repository (2609.31176) | Gemma 4 | evaluates | 90% |
| JevSoup: System-One Routing for Training-Free LoRA Composition (2609.30922) | Qwen3.8-LiveTranslate | evaluates | 98% |
| ORCA: Evaluating LLMs on Data Science Code Translation (2609.30749) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 100% |
| Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents (2609.30725) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering (2609.30489) | Grok | evaluates | 99% |
| BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering (2609.30489) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 99% |
| BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering (2609.30489) | GPT-6 Astra | evaluates | 99% |
| Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476) | Qwen3.8-LiveTranslate | evaluates | 100% |
| Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675) | Qwen3.8-LiveTranslate | evaluates | 98% |
| Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928) | GPT-6 Astra | evaluates | 99% |
| Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754) | GPT-6 Astra | evaluates | 99% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | Gemma 4 12B | evaluates | 95% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | GPT-6 Astra | evaluates | 95% |
| A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure (2608.13626) | Qwen3.8-LiveTranslate | evaluates | 98% |
| MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883) | GPT-6 Astra | related | 80% |
| Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958) | GPT-6 Astra | evaluates | 98% |
| Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290) | Qwen3.8-LiveTranslate | related | 90% |
| iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model (2609.29626) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model (2609.29626) | GPT-6 Astra | evaluates | 95% |
| Pistis Technical Report (2609.28554) | Qwen-Image-2.1 | related | 95% |
| CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars (2609.29075) | GPT-6 Astra | evaluates | 95% |
| CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars (2609.29075) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 95% |
| Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench (2609.29251) | GPT-6 Astra | evaluates | 98% |
| CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment (2609.29109) | Qwen-Image-2.1 | evaluates | 98% |
| Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races (2609.29522) | Qwen-Image-2.1 | evaluates | 99% |
| The Last Human Gate: Forward Deployed Engineering for Governance Automation (2609.29345) | DeepSeek V4-Flash | evaluates | 99% |
| The Last Human Gate: Forward Deployed Engineering for Governance Automation (2609.29345) | GPT-6 Astra | evaluates | 99% |
| The Last Human Gate: Forward Deployed Engineering for Governance Automation (2609.29345) | Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS | evaluates | 99% |
| Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons (2609.29781) | Gemma 4 | evaluates | 98% |
| Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (2609.29429) | Jev | evaluates | 99% |
| Sequential knowledge editing breaks a model's ability to tell good evidence from bad, without costing it accuracy (2609.29587) | Qwen-Image-2.1 | evaluates | 99% |
| Reinforcement Learning with Verifiable Rewards for Small Search Agents (2609.28765) | Qwen-Image-2.1 | evaluates | 98% |
| COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference (2609.26913) | GPT-6 Astra | evaluates | 98% |
| Learning the Cost of Reliable Inference (2609.28322) | Qwen-Image-2.1 | evaluates | 98% |
| Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training (2609.27197) | Qwen-Image-2.1 | evaluates | 99% |
| Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance (2609.27333) | Mistral OCR 4 | evaluates | 98% |
| Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms (2609.27321) | Qwen-Image-2.1 | evaluates | 85% |
| Provably Complete Generalized Planning with LLMs (2609.27105) | GPT-6 Astra | technique used | 99% |
| Validation and Simulation Catch Different Errors: Four Levels of Evaluation for LLM-Generated Circuits (2609.26830) | GPT-6 Astra | evaluates | 98% |
| Not What You Meant: Can LLMs Follow a Specified Negation Semantics? (2609.27517) | GPT-6 Astra | evaluates | 98% |
| Are Stated Reasoning Steps Causally Load-Bearing? (2609.27038) | Qwen-Image-2.1 | evaluates | 100% |
| TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents (2609.26911) | GPT-6 Astra | evaluates | 98% |
| Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model (2609.22939) | Qwen-Image-2.1 | evaluates | 99% |
| When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used (2609.22910) | Qwen-Image-2.1 | evaluates | 95% |
| UniK: Universal Knowledge Perception for Digital and Physical AI (2609.23971) | GPT-Live-1 | evaluates | 95% |
| DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents (2609.24092) | Qwen-Image-2.1 | evaluates | 99% |
| EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation (2609.24115) | Gemma 4 | evaluates | 95% |
| SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models (2609.24198) | Qwen-Image-2.1 | technique used | 99% |
| LIMIT: Less Is More for Instruction Tuning in Text-to-SQL (2609.24186) | Qwen-Image-2.1 | evaluates | 95% |
| FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering (2609.24002) | GPT-Live-1 | evaluates | 98% |
| Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review (2609.22694) | Gemma 4 | evaluates | 98% |
| PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI (2609.23695) | Grok | evaluates | 99% |
| PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI (2609.23695) | GPT-Live-1 | evaluates | 99% |
| Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events (2609.23074) | TimesFM 2.5 | evaluates | 98% |
| AutoRecLab: Describe the Experiment, Get the Code! (2609.21863) | GPT-Live-1 | technique used | 95% |
| TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers (2609.21139) | Qwen-Image-2.1 | evaluates | 99% |
| DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement (2609.21423) | GPT-Live-1 | evaluates | 80% |
| Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation (2609.21208) | Qwen-Image-2.1 | evaluates | 95% |
| RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models (2609.20971) | Qwen-Image-2.1 | evaluates | 98% |
| Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models (2609.21094) | GPT-Live-1 | evaluates | 98% |
| Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design (2609.22086) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development (2609.21293) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal (2609.21996) | Mistral OCR 4 | evaluates | 98% |
| A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal (2609.21996) | Qwen3.8-LiveTranslate | evaluates | 98% |
| A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal (2609.21996) | Gemma 4 12B | evaluates | 98% |
| AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory (2609.21940) | Qwen3.8-LiveTranslate | evaluates | 98% |
| Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake (2609.21149) | GPT-Live-1 | evaluates | 95% |
| SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? (2609.21190) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | GPT-Live-1 | evaluates | 99% |
| Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552) | GPT-Live-1 | evaluates | 100% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking | technique used | 98% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | GPT-Live-1 | technique used | 98% |
| ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145) | GPT-Live-1 | evaluates | 98% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking | evaluates | 95% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | GPT-Live-1 | evaluates | 95% |
| The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588) | Qwen3.8-LiveTranslate | evaluates | 98% |