Radar
Which research topics are accelerating, and where earlier papers have already landed in shipped models.
More papers match this topic — see “Show more” below.
Papers this week
50
Trending topic
Agents
▲ 19%papers this window vs prior window
Papers linked to models
123
linked by the desk, past 90 days
Median days paper → model
—
lower is faster
papers this window by topic · Δ vs prior window · a paper can carry more than one topic
Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.
Divergent Response Modes in Frontier Language Models Under Steering Pressure
arXiv:2608.06578AI-inferred · highAgent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
arXiv:2608.07169AI-inferred · highReason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885AI-inferred · highSodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
More papers match this topic — see “Show more” below.
| Topic | Papers this window | Prior window | Δ |
|---|---|---|---|
| Agents | 136 | 114 | +19% |
| Evaluation & Benchmarks | 38 | 27 | +41% |
| Multimodal | 24 | 33 | -27% |
| Safety & Alignment | 20 | 15 | +33% |
| Training & Optimization | 20 | 14 | +43% |
| Efficiency & Inference | 18 | 21 | -14% |
| Reasoning | 16 | 12 | +33% |
| Interpretability | 10 | 15 | -33% |
| Reinforcement Learning | 10 | 14 | -29% |
| Retrieval & RAG | 9 | 11 | -18% |
| Code Generation | 8 | 12 | -33% |
| Privacy & Security | 7 | 3 | +133% |
| Language Understanding | 5 | 4 | +25% |
| Generative Models | 5 | 15 | -67% |
| Robotics & Embodied AI | 5 | 7 | -29% |
| Long Context | 4 | 5 | -20% |
| Speech & Audio | 4 | 6 | -33% |
| Computer Vision | 3 | 8 | -62% |
| Mixture of Experts | 2 | 2 | 0% |
| Data & Synthetic Data | 1 | 4 | -75% |
| Other | 164 | 150 | +9% |
Safety & Alignment
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608.07762AI-inferred · highAn Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
arXiv:2608.07651AI-inferred · highWhen the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
arXiv:2608.07813AI-inferred · highThe Authority Expectancy Effect in Multi-User Conflict
arXiv:2608.08026AI-inferred · high| Paper | Linked model | Relation | Confidence |
|---|---|---|---|
| Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578) | GPT-5.6 | evaluates | 99% |
| Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169) | GPT-5.6 Luna | technique used | 98% |
| Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885) | GPT-5.6 Luna | evaluates | 98% |
| SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055) | DeepSeek-V4-Flash-0731 | technique used | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | GLM-5.3 | evaluates | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | GPT-5.6 Luna | evaluates | 98% |
| An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651) | GPT-5.6 Luna | evaluates | 99% |
| An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651) | Gemini 3.7 Flash | evaluates | 99% |
| When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | Grok 4.6 | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | GPT-5.6 Luna | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | Gemini 3.7 Flash | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089) | GPT-5.6 Luna | evaluates | 99% |
| Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729) | Qwen 3.8 Max | evaluates | 98% |
| Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670) | Qwen 3.8 Max | evaluates | 98% |
| No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938) | Qwen 3.8 Max | evaluates | 99% |
| Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797) | Qwen 3.8 Max | technique used | 98% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | GPT-5.6 Luna | evaluates | 99% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | Qwen 3.8 Max | evaluates | 99% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | Gemma 4 | evaluates | 99% |
| Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755) | GPT-5.6 Luna | technique used | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Gemma 4 | evaluates | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Mistral OCR 4 | evaluates | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Qwen 3.8 Max | evaluates | 95% |
| Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397) | Qwen 3.8 Max | evaluates | 99% |
| Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378) | Gemma 4 | evaluates | 98% |
| Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378) | Qwen 3.8 Max | evaluates | 98% |
| PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 98% |
| Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209) | Qwen 3.8 Max | evaluates | 99% |
| TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961) | Qwen 3.8 Max | evaluates | 98% |
| Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631) | Gemma 4 | evaluates | 99% |
| Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631) | Qwen 3.8 Max | evaluates | 99% |
| FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574) | Grok 4.6 | evaluates | 99% |
| StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414) | GPT-5.6 Luna | evaluates | 99% |
| SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775) | GPT-5.6 Luna | evaluates | 98% |
| CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341) | GPT-5.6 Luna | evaluates | 98% |
| AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616) | GPT-5.6 Luna | technique used | 98% |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403) | Qwen 3.8 27B | evaluates | 99% |
| XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676) | Qwen 3.8 27B | evaluates | 95% |
| Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097) | GPT-5.6 Luna | evaluates | 95% |
| HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692) | Qwen 3.8 27B | evaluates | 99% |
| Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994) | GPT-5.6 Luna | evaluates | 95% |
| From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381) | Qwen 3.8 27B | technique used | 98% |
| Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233) | Qwen 3.8 27B | technique used | 99% |
| Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232) | GPT-5.6 Luna | evaluates | 95% |
| FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 86% |
| Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224) | GPT-5.6 Luna | evaluates | 95% |
| Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343) | GPT-5.6 Luna | evaluates | 99% |
| An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 95% |
| Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210) | Qwen 3.8 27B | evaluates | 95% |
| From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219) | GPT-5.6 Luna | evaluates | 98% |
| RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 90% |
| Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585) | GPT-5.6 Luna | evaluates | 95% |
| Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420) | Qwen 3.8 27B | evaluates | 98% |
| Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156) | Qwen 3.8 27B | evaluates | 95% |
| Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409) | trace-supervised symbolic neural CPU | technique used | 99% |
| Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063) | Qwen 3.8 27B | evaluates | 99% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | AlphaEvolve | related | 94% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | GPT-5.6 Luna | evaluates | 98% |
| Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476) | Qwen 3.8 27B | evaluates | 99% |
| Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675) | Qwen 3.8 27B | evaluates | 95% |
| Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928) | GPT-5.6 Luna | evaluates | 99% |
| From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787) | GPT-5.6 Luna | evaluates | 95% |
| Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754) | GPT-5.6 Luna | evaluates | 99% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | GPT-5.6 Luna | evaluates | 95% |
| MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883) | GPT-5.6 Luna | evaluates | 95% |
| Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958) | GPT-5.6 Luna | evaluates | 100% |
| Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290) | Qwen 3.8 Max | related | 95% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | DeepSeek-V4-Flash-0731 | evaluates | 99% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | GPT-5.6 Luna | evaluates | 95% |
| Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552) | GPT-5.6 Luna | evaluates | 99% |
| When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | GPT-5.6 Luna | technique used | 98% |
| ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145) | GPT-5.6 Luna | evaluates | 95% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | GPT-5.6 Luna | evaluates | 98% |
| The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588) | GPT-5.6 Luna | evaluates | 98% |
| LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927) | GPT-5.6 Luna | evaluates | 95% |
| Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338) | Qwen 3.8 Max | evaluates | 98% |
| Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338) | GPT-5.6 Luna | evaluates | 98% |
| When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230) | GPT-5.6 Luna | evaluates | 98% |
| MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314) | Qwen 3.8 Max | technique used | 98% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | Grok | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | Gemini 3.7 Flash | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | GPT-5.6 Luna | evaluates | 90% |
| Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220) | GPT-5.6 Luna | technique used | 99% |
| A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199) | Athena-Brain-8B | evaluates | 98% |
| Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290) | Qwen 3.8 Max | evaluates | 98% |
| ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974) | Gemini 3.7 Flash | evaluates | 95% |
| ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | Gemini 3.7 Flash | evaluates | 99% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | GPT-5.6 Luna | evaluates | 99% |
| SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842) | Qwen 3.8 Max | evaluates | 98% |
| Can Agent Memory Systems Track Evolving State? (2608.19652) | Qwen 3.8 Max | evaluates | 98% |
| Can Agent Memory Systems Track Evolving State? (2608.19652) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535) | Qwen 3.8 Max | evaluates | 99% |
| Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299) | GPT-5.6 Luna | evaluates | 80% |
| When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909) | Grok | evaluates | 98% |
| Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106) | Qwen 3.8 Max | evaluates | 99% |
| DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103) | DeepSeek V4-Flash | technique used | 99% |
| Abliteration Mitigation via Refusal Aliases (2608.18093) | Gemma 4 | evaluates | 99% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Gemma 4 | evaluates | 85% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Qwen 3.8 Max | evaluates | 85% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Mistral OCR 4 | evaluates | 85% |
| Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089) | Qwen 3.8 Max | evaluates | 99% |
| Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089) | Mistral OCR 4 | evaluates | 99% |
| Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161) | Qwen 3.8 Max | evaluates | 98% |
| Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591) | Qwen 3.8 Max | technique used | 100% |
| A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389) | Qwen 3.8 Max | evaluates | 98% |
| A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324) | Qwen 3.8 Max | evaluates | 98% |
| Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142) | Mistral OCR 4 | evaluates | 99% |
| Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078) | DeepSeek V4-Flash | evaluates | 95% |
| Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100) | GPT-5.6 Luna | evaluates | 99% |
| Breaking the weakest link to evade vision language models (2608.18938) | Qwen 3.8 Max | evaluates | 99% |
| Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884) | Qwen 3.8 Max | evaluates | 98% |
| CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 80% |
| FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 100% |
| Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 90% |
| Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336) | Qwen 3.8 Max | evaluates | 98% |
| Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 80% |
| Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111) | GPT-5.6 Luna | evaluates | 80% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | Qwen 3.8 Max | evaluates | 99% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | Gemini 3.7 Flash | evaluates | 99% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | GPT-5.6 Luna | evaluates | 99% |
| Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273) | GPT-5.6 Luna | evaluates | 95% |
| StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050) | GPT-5.6 Luna | evaluates | 98% |
| StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050) | Gemini 3.7 Flash | evaluates | 98% |
| Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684) | Qwen 3.8 Max | evaluates | 98% |
| The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638) | GPT-5.6 Luna | evaluates | 95% |
| TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588) | GPT-5.6 Luna | evaluates | 95% |
| Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528) | Qwen 3.8 Max | evaluates | 99% |
| LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393) | Qwen 3.8 Max | evaluates | 99% |
| TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336) | Qwen 3.8 Max | evaluates | 98% |
| SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning (2608.17301) | Qwen 3.8 Max | evaluates | 98% |
| When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909) | Grok 4.6 | evaluates | 99% |
| When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909) | Gemini 3.7 Flash | evaluates | 98% |
| When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909) | GPT-5.6 Luna | evaluates | 98% |
| FedPref: Federated Preference Learning for Structured Radiology Report Extraction (2608.16971) | Qwen 3.8 Max | technique used | 99% |
| Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification (2608.17247) | GPT-5.6 Luna | evaluates | 98% |
| GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890) | GPT-5.6 Luna | evaluates | 99% |
| GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | DeepSeek V4-Flash | evaluates | 99% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | GPT-5.6 Sol | evaluates | 99% |
| Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 (2608.14992) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 100% |
| Small Models Scout Bottleneck Order for Large-Model Data Control (2608.14936) | Qwen 3.8 Max | evaluates | 95% |
| LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927) | GPT-5.6 Sol | technique used | 98% |
| The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588) | Qwen 3.8 Max | evaluates | 98% |
| The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588) | GPT-5.6 Sol | evaluates | 99% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | Gemini 3.7 Flash | technique used | 98% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 98% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | GPT-5.6 Sol | technique used | 99% |
| Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads (2608.15117) | Qwen 3.8 Max | evaluates | 99% |
| When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680) | DeepSeek V4-Flash | evaluates | 99% |
| ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145) | GPT-5.6 Sol | evaluates | 95% |
| ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145) | GPT-5.6 Sol | technique used | 95% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | Gemini 3.7 Flash | evaluates | 98% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | GPT-5.6 Sol | evaluates | 98% |
| Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552) | GPT-5.6 Sol | evaluates | 99% |
| MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883) | GPT-5.6 Sol | technique used | 95% |
| From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787) | GPT-5.6 Sol | evaluates | 99% |
| Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290) | Qwen 3.8 Max | technique used | 85% |
| Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958) | GPT-5.6 Sol | evaluates | 99% |
| Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754) | GPT-5.6 Sol | evaluates | 98% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | Gemma 4 | evaluates | 95% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | GPT-5.6 Sol | evaluates | 98% |
| A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure (2608.13626) | Qwen 3.8 Max | evaluates | 98% |
| Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking (2608.13565) | Qwen 3.8 Max | evaluates | 98% |
| OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality (2608.05263) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 90% |
| NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167) | GPT-5.6 Sol | evaluates | 99% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | AlphaEvolve | cites | 88% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | GPT-5.6 Sol | evaluates | 98% |
| Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese (2608.12373) | Gemini 3.7 Flash | evaluates | 98% |
| Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese (2608.12373) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675) | Qwen 3.8 Max | evaluates | 98% |
| Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156) | Qwen 3.8 Max | evaluates | 98% |
| Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063) | Mistral OCR 4 | evaluates | 99% |
| Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063) | Qwen 3.8 Max | evaluates | 99% |
| Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420) | Qwen 3.8 Max | evaluates | 98% |
| Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585) | Gemini 3.7 Flash | evaluates | 95% |
| Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585) | GPT-5.6 Sol | evaluates | 95% |
| Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476) | Qwen 3.8 Max | evaluates | 98% |
| Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928) | GPT-5.6 Sol | evaluates | 98% |
| Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036) | Claude Opus 5 | evaluates | 90% |
| Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994) | GPT-5.6 Sol | evaluates | 95% |
| XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676) | Mistral OCR 4 | evaluates | 95% |
| XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676) | Qwen 3.8 Max | evaluates | 98% |