Sens.aiAI signal desk
TodayRadarBriefing
Admin sign in
Sens.aiSensif.aiAI signal desk
ModelsResearchCountriesOpen vs ClosedComputeBetaCommentary
Loading…
TodayRadarBriefingAdmin

Radar

Research Radar

Which research topics are accelerating, and where earlier papers have already landed in shipped models.

Filtered to Reinforcement Learning — 50 papersClear filter

More papers match this topic — see “Show more” below.

Papers this week

50

Trending topic

Agents

▲ 16%

papers this window vs prior window

Papers linked to models

124

linked by the desk, past 90 days

Median days paper → model

—

lower is faster

Topic velocity

papers this window by topic · Δ vs prior window · a paper can carry more than one topic

AgentsAgents
135▲16%
Evaluation & BenchmarksEvaluation & Benchmarks
37▲37%
MultimodalMultimodal
23▼30%
Safety & Alignment

Topic momentum

prior window → this window

Agents

135 this window ▲16%

Evaluation & Benchmarks

37 this window ▲37%

Multimodal

23 this window ▼30%

Where research lands: paper → model linkage

Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

arXiv:2608.06474AI-inferred · high

Divergent Response Modes in Frontier Language Models Under Steering Pressure

arXiv:2608.06578AI-inferred · high

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

arXiv:2608.07169AI-inferred · high

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Filtered to Reinforcement Learning — 50 papersClear filter

More papers match this topic — see “Show more” below.

Papers on Reinforcement Learning

Clear filter
dynamic context schedulingout-of-distribution generalizationpublished Aug 24, 2026 · arXiv:2608.20799

Dynamic Context Scheduling: Learning Beyond the Static Universe

reinforcement learningscientific discoverypublished Aug 24, 2026 · arXiv:2608.20686

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Safety & Alignment
20▲33%
Training & OptimizationTraining & Optimization
20▲43%
Efficiency & InferenceEfficiency & Inference
18▼14%
ReasoningReasoning
16▲33%
InterpretabilityInterpretability
11▼27%
Reinforcement LearningReinforcement Learning
10▼29%
Retrieval & RAGRetrieval & RAG
9▼18%
Code GenerationCode Generation
8▼33%
Privacy & SecurityPrivacy & Security
7▲133%
Language UnderstandingLanguage Understanding
5▲25%
Generative ModelsGenerative Models
5▼67%
Robotics & Embodied AIRobotics & Embodied AI
5▼38%
Long ContextLong Context
4▼20%
Speech & AudioSpeech & Audio
4▼33%
Computer VisionComputer Vision
3▼63%
Mixture of ExpertsMixture of Experts
20%
Data & Synthetic DataData & Synthetic Data
1▼75%
OtherOther
163▲9%
View as table
TopicPapers this windowPrior windowΔ
Agents135116+16%
Evaluation & Benchmarks3727+37%
Multimodal2333-30%
Safety & Alignment2015+33%
Training & Optimization2014+43%
Efficiency & Inference1821-14%
Reasoning1612+33%
Interpretability1115-27%
Reinforcement Learning1014-29%
Retrieval & RAG911-18%
Code Generation812-33%
Privacy & Security73+133%
Language Understanding54+25%
Generative Models515-67%
Robotics & Embodied AI58-37%
Long Context45-20%
Speech & Audio46-33%
Computer Vision38-62%
Mixture of Experts220%
Data & Synthetic Data14-75%
Other163150+9%

Safety & Alignment

20 this window ▲33%

Training & Optimization

20 this window ▲43%

Efficiency & Inference

18 this window ▼14%
arXiv:2608.07885AI-inferred · high

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

arXiv:2608.08055AI-inferred · high

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762AI-inferred · high

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

arXiv:2608.07651AI-inferred · high

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

arXiv:2608.07813AI-inferred · high

The Authority Expectancy Effect in Multi-User Conflict

arXiv:2608.08026AI-inferred · high
GPT-5.62 links
claude-fable-5, claude-sonnet-5, claude-opus-51 link
GPT-5.6 Luna5 links
DeepSeek-V4-Flash-07313 links
GLM-5.31 link
Gemini 3.7 Flash1 link
Grok 4.61 link
View as table
PaperLinked modelRelationConfidence
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader (2608.06474)GPT-5.6evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)GPT-5.6evaluates99%
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169)GPT-5.6 Lunatechnique used98%
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885)GPT-5.6 Lunaevaluates98%
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055)DeepSeek-V4-Flash-0731technique used98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)DeepSeek-V4-Flash-0731evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GLM-5.3evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GPT-5.6 Lunaevaluates98%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)GPT-5.6 Lunaevaluates99%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)Gemini 3.7 Flashevaluates99%
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813)DeepSeek-V4-Flash-0731evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Grok 4.6evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)GPT-5.6 Lunaevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Gemini 3.7 Flashevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089)GPT-5.6 Lunaevaluates99%
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729)Qwen 3.8 Maxevaluates98%
Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670)Qwen 3.8 Maxevaluates98%
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938)Qwen 3.8 Maxevaluates99%
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797)Qwen 3.8 Maxtechnique used98%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)GPT-5.6 Lunaevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Qwen 3.8 Maxevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Gemma 4evaluates99%
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755)GPT-5.6 Lunatechnique used95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Gemma 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Mistral OCR 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Qwen 3.8 Maxevaluates95%
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397)Qwen 3.8 Maxevaluates99%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Gemma 4evaluates98%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Qwen 3.8 Maxevaluates98%
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209)Qwen 3.8 Maxevaluates99%
TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961)Qwen 3.8 Maxevaluates98%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Gemma 4evaluates99%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Qwen 3.8 Maxevaluates99%
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574)Grok 4.6evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)GPT-5.6 Lunaevaluates99%
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775)GPT-5.6 Lunaevaluates98%
CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090)DeepSeek-V4-Flash-0731evaluates98%
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341)GPT-5.6 Lunaevaluates98%
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616)GPT-5.6 Lunatechnique used98%
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403)Qwen 3.8 27Bevaluates99%
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676)Qwen 3.8 27Bevaluates95%
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097)GPT-5.6 Lunaevaluates95%
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692)Qwen 3.8 27Bevaluates99%
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994)GPT-5.6 Lunaevaluates95%
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381)Qwen 3.8 27Btechnique used98%
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233)Qwen 3.8 27Btechnique used99%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)GPT-5.6 Lunaevaluates95%
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates86%
Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224)GPT-5.6 Lunaevaluates95%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)GPT-5.6 Lunaevaluates99%
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249)claude-fable-5, claude-sonnet-5, claude-opus-5technique used95%
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210)Qwen 3.8 27Bevaluates95%
From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219)GPT-5.6 Lunaevaluates98%
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241)claude-fable-5, claude-sonnet-5, claude-opus-5technique used90%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)GPT-5.6 Lunaevaluates95%
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420)Qwen 3.8 27Bevaluates98%
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156)Qwen 3.8 27Bevaluates95%
Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409)trace-supervised symbolic neural CPUtechnique used99%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Qwen 3.8 27Bevaluates99%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)AlphaEvolverelated94%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)GPT-5.6 Lunaevaluates98%
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476)Qwen 3.8 27Bevaluates99%
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675)Qwen 3.8 27Bevaluates95%
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928)GPT-5.6 Lunaevaluates99%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Lunaevaluates95%
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754)GPT-5.6 Lunaevaluates99%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)GPT-5.6 Lunaevaluates95%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Lunaevaluates95%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Lunaevaluates100%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxrelated95%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek-V4-Flash-0731evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Lunaevaluates95%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Lunaevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek-V4-Flash-0731evaluates98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Lunatechnique used98%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Lunaevaluates95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Lunaevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Lunaevaluates98%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Lunaevaluates95%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)Qwen 3.8 Maxevaluates98%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)GPT-5.6 Lunaevaluates98%
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230)GPT-5.6 Lunaevaluates98%
MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314)Qwen 3.8 Maxtechnique used98%
Automatic bioinformatic software named entity recognition from literature (2608.19201)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Grokevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Gemini 3.7 Flashevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)GPT-5.6 Lunaevaluates90%
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220)GPT-5.6 Lunatechnique used99%
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199)Athena-Brain-8Bevaluates98%
Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290)Qwen 3.8 Maxevaluates98%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)Gemini 3.7 Flashevaluates95%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)DeepSeek-V4-Flash-0731evaluates95%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)Gemini 3.7 Flashevaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)GPT-5.6 Lunaevaluates99%
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)DeepSeek-V4-Flash-0731evaluates98%
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535)Qwen 3.8 Maxevaluates99%
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299)GPT-5.6 Lunaevaluates80%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grokevaluates98%
Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106)Qwen 3.8 Maxevaluates99%
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103)DeepSeek V4-Flashtechnique used99%
Abliteration Mitigation via Refusal Aliases (2608.18093)Gemma 4evaluates99%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Gemma 4evaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Qwen 3.8 Maxevaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Mistral OCR 4evaluates85%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Qwen 3.8 Maxevaluates99%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Mistral OCR 4evaluates99%
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161)Qwen 3.8 Maxevaluates98%
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591)Qwen 3.8 Maxtechnique used100%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)Qwen 3.8 Maxevaluates98%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324)Qwen 3.8 Maxevaluates98%
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142)Mistral OCR 4evaluates99%
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078)DeepSeek V4-Flashevaluates95%
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100)GPT-5.6 Lunaevaluates99%
Breaking the weakest link to evade vision language models (2608.18938)Qwen 3.8 Maxevaluates99%
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884)Qwen 3.8 Maxevaluates98%
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)Qwen 3.8 Maxevaluates98%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)GPT-5.6 Lunaevaluates80%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Qwen 3.8 Maxevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Gemini 3.7 Flashevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)GPT-5.6 Lunaevaluates99%
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273)GPT-5.6 Lunaevaluates95%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)GPT-5.6 Lunaevaluates98%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)Gemini 3.7 Flashevaluates98%
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684)Qwen 3.8 Maxevaluates98%
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638)GPT-5.6 Lunaevaluates95%
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588)GPT-5.6 Lunaevaluates95%
Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528)Qwen 3.8 Maxevaluates99%
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393)Qwen 3.8 Maxevaluates99%
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336)Qwen 3.8 Maxevaluates98%
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning (2608.17301)Qwen 3.8 Maxevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grok 4.6evaluates99%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Gemini 3.7 Flashevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)GPT-5.6 Lunaevaluates98%
FedPref: Federated Preference Learning for Structured Radiology Report Extraction (2608.16971)Qwen 3.8 Maxtechnique used99%
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification (2608.17247)GPT-5.6 Lunaevaluates98%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)GPT-5.6 Lunaevaluates99%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek V4-Flashevaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Solevaluates99%
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 (2608.14992)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Small Models Scout Bottleneck Order for Large-Model Data Control (2608.14936)Qwen 3.8 Maxevaluates95%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Soltechnique used98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)Qwen 3.8 Maxevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Solevaluates99%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)Gemini 3.7 Flashtechnique used98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Soltechnique used99%
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads (2608.15117)Qwen 3.8 Maxevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek V4-Flashevaluates99%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Solevaluates95%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Soltechnique used95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)Gemini 3.7 Flashevaluates98%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Solevaluates98%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Solevaluates99%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Soltechnique used95%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Solevaluates99%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxtechnique used85%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Solevaluates99%
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754)GPT-5.6 Solevaluates98%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)Gemma 4evaluates95%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)GPT-5.6 Solevaluates98%
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure (2608.13626)Qwen 3.8 Maxevaluates98%
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking (2608.13565)Qwen 3.8 Maxevaluates98%
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality (2608.05263)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167)GPT-5.6 Solevaluates99%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)AlphaEvolvecites88%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)GPT-5.6 Solevaluates98%
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese (2608.12373)Gemini 3.7 Flashevaluates98%
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese (2608.12373)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675)Qwen 3.8 Maxevaluates98%
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156)Qwen 3.8 Maxevaluates98%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Mistral OCR 4evaluates99%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Qwen 3.8 Maxevaluates99%
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420)Qwen 3.8 Maxevaluates98%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)Gemini 3.7 Flashevaluates95%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)GPT-5.6 Solevaluates95%
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476)Qwen 3.8 Maxevaluates98%
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928)GPT-5.6 Solevaluates98%
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036)Claude Opus 5evaluates90%
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994)GPT-5.6 Solevaluates95%
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676)Mistral OCR 4evaluates95%
world modelsmodel-based reinforcement learningpublished Aug 24, 2026 · arXiv:2608.20401

World models of environment, agent and joint agent-environment systems

dual-level credit assignmentstructured reasoningpublished Aug 21, 2026 · arXiv:2608.20161

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

quality-diversity reinforcement learninghierarchical skill policiespublished Aug 21, 2026 · arXiv:2608.19684

Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

game world modelingreinforcement learningpublished Aug 20, 2026 · arXiv:2608.18079

Position: Profiling Game Worlds by Transition Complexity

model-based reinforcement learningworld modelspublished Aug 19, 2026 · arXiv:2608.17959

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

generative AIlarge language modelspublished Aug 19, 2026 · arXiv:2608.16907

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

offline reinforcement learningeducational technologypublished Aug 18, 2026 · arXiv:2608.14851

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

signal temporal logicreinforcement learningpublished Aug 17, 2026 · arXiv:2608.13625

Reward Machines for Signal Temporal Logic

reinforcement learningonline educationpublished Aug 13, 2026 · arXiv:2608.11245

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

reinforcement learningQ-learningpublished Aug 12, 2026 · arXiv:2608.10549

Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization

reinforcement learningautonomous drivingpublished Aug 12, 2026 · arXiv:2608.10403

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

reinforcement learningpublic transit operationspublished Aug 12, 2026 · arXiv:2608.10207

Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding

deep reinforcement learningexplainable AIpublished Aug 12, 2026 · arXiv:2608.09967

SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning

reinforcement learningreward shapingpublished Aug 11, 2026 · arXiv:2608.08158

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

vehicle routingdeep reinforcement learningpublished Aug 10, 2026 · arXiv:2608.06668

Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry

large language modelsweb developmentpublished Aug 10, 2026 · arXiv:2608.06474

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

LinkedGPT-5.6
interactive video world modelslong-horizon planningpublished Aug 6, 2026 · arXiv:2608.04964

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

reinforcement learninglarge language modelspublished Aug 5, 2026 · arXiv:2608.03119

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

GRPOcredit redistributionpublished Aug 5, 2026 · arXiv:2608.03467

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

reinforcement learninghierarchical reinforcement learningpublished Aug 5, 2026 · arXiv:2608.02993

Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning

hypergradient methodssample complexitypublished Aug 3, 2026 · arXiv:2607.28849

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

reinforcement learningpolicy decompositionpublished Aug 3, 2026 · arXiv:2607.29246

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

multi-objective reinforcement learningpreference-based reinforcement learningpublished Aug 3, 2026 · arXiv:2607.29559

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

deep reinforcement learningrandom CNN feature extractorspublished Jul 31, 2026 · arXiv:2607.26059

Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning

reinforcement learning with verifiable rewardslarge language model reasoningpublished Jul 30, 2026 · arXiv:2607.24833

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

maritime surveillanceheterogeneous sensor networkspublished Jul 28, 2026 · arXiv:2607.22667

Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

document classificationhierarchical reinforcement learningpublished Jul 28, 2026 · arXiv:2607.22644

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

metal-organic frameworkscrystallographic information filespublished Jul 23, 2026 · arXiv:2607.19935

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

LLM preference alignmentreinforcement learningpublished Jul 23, 2026 · arXiv:2607.19824

Rewarding Better Thinking for LLM Preference Alignment

offline reinforcement learninglarge language modelspublished Jul 23, 2026 · arXiv:2607.19450

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

RLHFpreference-based reinforcement learningpublished Jul 22, 2026 · arXiv:2607.18258

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

parallel reasoningShapley value attributionpublished Jul 22, 2026 · arXiv:2607.18979

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

deep reinforcement learningBaghchalpublished Jul 22, 2026 · arXiv:2607.18296

Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal

partially observable reinforcement learningtemporal knowledge graphspublished Jul 22, 2026 · arXiv:2607.18368

Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability

reinforcement learning from human feedbackpreference datapublished Jul 21, 2026 · arXiv:2607.16195

Rater State Bias in RLHF Preference Data: An Audit Framework

reinforcement learningpolicy verificationpublished Jul 21, 2026 · arXiv:2607.16210

A Survey on the Verification of Reinforcement Learning Policies

reinforcement learningRLVRpublished Jul 21, 2026 · arXiv:2607.16206

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization

reinforcement learning with verifiable rewardsoff-policy reinforcement learningpublished Jul 21, 2026 · arXiv:2607.16205

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

multi-agent reinforcement learninggamespublished Jul 21, 2026 · arXiv:2607.17560

Reinforcement Learning: From Algorithms To Foundation Models

reinforcement learningexplainable AIpublished Jul 20, 2026 · arXiv:2607.15459

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

world modelsmodel-based reinforcement learningpublished Jul 17, 2026 · arXiv:2607.15142

Concept-Guided Spatial Regularization for World Models in Atari Pong

reinforcement learningarchive-based explorationpublished Jul 14, 2026 · arXiv:2607.09971

TopoExplore: Topological Discrimination for Archive-Based Exploration

AlphaZeroMonte Carlo Tree Searchpublished Jul 13, 2026 · arXiv:2607.08984

AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

general-utility Markov decision processesrisk-aware objectivespublished Jul 13, 2026 · arXiv:2607.09298

Risk-Aware General-Utility Markov Decision Processes

reinforcement learningprompt optimisationpublished Jul 13, 2026 · arXiv:2607.08837

Prompt-Driven Exploration

deep reinforcement learningreinforcement learning evaluationpublished Jul 10, 2026 · arXiv:2607.07769

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

AI-generated training dataknowledge distillationpublished Jul 10, 2026 · arXiv:2607.08255

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

portfolio optimizationdeep reinforcement learningpublished Jul 9, 2026 · arXiv:2607.06610

Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

Show more