Sens.aiAI signal desk
TodayRadarBriefing
Admin sign in
Sens.aiSensif.aiAI signal desk
ModelsResearchCountriesOpen vs ClosedComputeBetaCommentary
Loading…
TodayRadarBriefingAdmin

Radar

Research Radar

Which research topics are accelerating, and where earlier papers have already landed in shipped models.

Filtered to Evaluation & Benchmarks — 146 papersClear filter

Papers this week

50

Trending topic

Agents

▲ 22%

papers this window vs prior window

Papers linked to models

137

linked by the desk, past 90 days

Median days paper → model

—

lower is faster

Topic velocity

papers this window by topic · Δ vs prior window · a paper can carry more than one topic

AgentsAgents
133▲22%
Evaluation & BenchmarksEvaluation & Benchmarks
39▲44%
MultimodalMultimodal
26▼21%
Safety & AlignmentSafety & Alignment
View as table
TopicPapers this windowPrior windowΔ
Agents133109+22%
Evaluation & Benchmarks3927+44%
Multimodal2633-21%

Topic momentum

prior window → this window

Agents

133 this window ▲22%

Evaluation & Benchmarks

39 this window ▲44%

Multimodal

26 this window ▼21%

Where research lands: paper → model linkage

Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.

Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems

arXiv:2608.22143AI-inferred · high

GenCoord: Skill-Path Commitments under Private Information

arXiv:2608.22055AI-inferred · high

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

arXiv:2608.22048AI-inferred · high

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

Filtered to Evaluation & Benchmarks — 146 papersClear filter

Papers on Evaluation & Benchmarks

Clear filter
mathematical reasoninglocal inferencepublished Aug 25, 2026 · arXiv:2608.22048

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

LinkedQwen 3.8 Max
language model evaluationstructured perturbationspublished Aug 25, 2026 · arXiv:2608.22138
23▲53%
Training & OptimizationTraining & Optimization
21▲50%
Efficiency & InferenceEfficiency & Inference
20▼5%
ReasoningReasoning
15▲36%
InterpretabilityInterpretability
12▼20%
Reinforcement LearningReinforcement Learning
10▼29%
Retrieval & RAGRetrieval & RAG
9▼18%
Privacy & SecurityPrivacy & Security
8▲300%
Code GenerationCode Generation
8▼27%
Robotics & Embodied AIRobotics & Embodied AI
6▼14%
Language UnderstandingLanguage Understanding
5▲25%
Generative ModelsGenerative Models
5▼67%
Speech & AudioSpeech & Audio
5▼17%
Mixture of ExpertsMixture of Experts
3▲50%
Long ContextLong Context
3▼40%
Computer VisionComputer Vision
3▼63%
Data & Synthetic DataData & Synthetic Data
1▼75%
OtherOther
165▲12%
Safety & Alignment
23
15
+53%
Training & Optimization2114+50%
Efficiency & Inference2021-5%
Reasoning1511+36%
Interpretability1215-20%
Reinforcement Learning1014-29%
Retrieval & RAG911-18%
Privacy & Security82+300%
Code Generation811-27%
Robotics & Embodied AI67-14%
Language Understanding54+25%
Generative Models515-67%
Speech & Audio56-17%
Mixture of Experts32+50%
Long Context35-40%
Computer Vision38-62%
Data & Synthetic Data14-75%
Other165147+12%

Safety & Alignment

23 this window ▲53%

Training & Optimization

21 this window ▲50%

Efficiency & Inference

20 this window ▼5%
arXiv:2608.22191AI-inferred · high

Redteaming Leading Arabic LLMs with ASAS

arXiv:2608.21985AI-inferred · high

Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning

arXiv:2608.22128AI-inferred · high

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

arXiv:2608.22061AI-inferred · high

HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews

arXiv:2608.21868AI-inferred · high

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

arXiv:2608.21811AI-inferred · high

HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

arXiv:2608.21792AI-inferred · high

ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling

arXiv:2608.21712AI-inferred · high
Qwen 3.8 Max7 links
Gemma 41 link
GPT-5.63 links
claude-fable-5, claude-sonnet-5, claude-opus-51 link
Gemini 3.7 Flash1 link
Athena-Brain-8B1 link
View as table
PaperLinked modelRelationConfidence
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Qwen 3.8 Maxevaluates98%
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Gemma 4evaluates98%
GenCoord: Skill-Path Commitments under Private Information (2608.22055)Qwen 3.8 Maxtechnique used98%
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning (2608.22048)Qwen 3.8 Maxevaluates100%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)Qwen 3.8 Maxevaluates98%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)GPT-5.6evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)GPT-5.6evaluates98%
Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning (2608.22128)Gemini 3.7 Flashevaluates95%
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds (2608.22061)GPT-5.6evaluates95%
HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews (2608.21868)Qwen 3.8 Maxtechnique used99%
Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning (2608.21811)Qwen 3.8 Maxevaluates98%
HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries (2608.21792)Qwen 3.8 Maxtechnique used99%
ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling (2608.21712)Athena-Brain-8Btechnique used99%
Context as an Environment: Programmatic Context Management for Long-Horizon Agents (2608.21690)Qwen 3.8 Maxevaluates95%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)GPT-5.6evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)Gemini 3.7 Flashevaluates98%
Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning (2608.21501)Qwen 3.8 Maxevaluates99%
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference (2608.21362)Qwen 3.8 Maxevaluates98%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)GPT-5.6evaluates100%
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents (2608.06735)DeepSeek-V4-Flash-0731evaluates95%
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167)GPT-5.6evaluates98%
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader (2608.06474)GPT-5.6evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)GPT-5.6evaluates99%
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169)GPT-5.6 Lunatechnique used98%
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885)GPT-5.6 Lunaevaluates98%
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055)DeepSeek-V4-Flash-0731technique used98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)DeepSeek-V4-Flash-0731evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GLM-5.3evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GPT-5.6 Lunaevaluates98%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)GPT-5.6 Lunaevaluates99%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)Gemini 3.7 Flashevaluates99%
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813)DeepSeek-V4-Flash-0731evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Grok 4.6evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)GPT-5.6 Lunaevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Gemini 3.7 Flashevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089)GPT-5.6 Lunaevaluates99%
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729)Qwen 3.8 Maxevaluates98%
Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670)Qwen 3.8 Maxevaluates98%
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938)Qwen 3.8 Maxevaluates99%
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797)Qwen 3.8 Maxtechnique used98%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)GPT-5.6 Lunaevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Qwen 3.8 Maxevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Gemma 4evaluates99%
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755)GPT-5.6 Lunatechnique used95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Gemma 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Mistral OCR 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Qwen 3.8 Maxevaluates95%
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397)Qwen 3.8 Maxevaluates99%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Gemma 4evaluates98%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Qwen 3.8 Maxevaluates98%
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209)Qwen 3.8 Maxevaluates99%
TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961)Qwen 3.8 Maxevaluates98%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Gemma 4evaluates99%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Qwen 3.8 Maxevaluates99%
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574)Grok 4.6evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)GPT-5.6 Lunaevaluates99%
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775)GPT-5.6 Lunaevaluates98%
CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090)DeepSeek-V4-Flash-0731evaluates98%
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341)GPT-5.6 Lunaevaluates98%
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616)GPT-5.6 Lunatechnique used98%
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403)Qwen 3.8 27Bevaluates99%
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676)Qwen 3.8 27Bevaluates95%
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097)GPT-5.6 Lunaevaluates95%
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692)Qwen 3.8 27Bevaluates99%
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994)GPT-5.6 Lunaevaluates95%
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381)Qwen 3.8 27Btechnique used98%
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233)Qwen 3.8 27Btechnique used99%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)GPT-5.6 Lunaevaluates95%
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates86%
Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224)GPT-5.6 Lunaevaluates95%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)GPT-5.6 Lunaevaluates99%
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249)claude-fable-5, claude-sonnet-5, claude-opus-5technique used95%
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210)Qwen 3.8 27Bevaluates95%
From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219)GPT-5.6 Lunaevaluates98%
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241)claude-fable-5, claude-sonnet-5, claude-opus-5technique used90%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)GPT-5.6 Lunaevaluates95%
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420)Qwen 3.8 27Bevaluates98%
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156)Qwen 3.8 27Bevaluates95%
Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409)trace-supervised symbolic neural CPUtechnique used99%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Qwen 3.8 27Bevaluates99%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)AlphaEvolverelated94%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)GPT-5.6 Lunaevaluates98%
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476)Qwen 3.8 27Bevaluates99%
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675)Qwen 3.8 27Bevaluates95%
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928)GPT-5.6 Lunaevaluates99%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Lunaevaluates95%
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754)GPT-5.6 Lunaevaluates99%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)GPT-5.6 Lunaevaluates95%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Lunaevaluates95%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Lunaevaluates100%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxrelated95%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek-V4-Flash-0731evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Lunaevaluates95%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Lunaevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek-V4-Flash-0731evaluates98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Lunatechnique used98%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Lunaevaluates95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Lunaevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Lunaevaluates98%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Lunaevaluates95%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)Qwen 3.8 Maxevaluates98%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)GPT-5.6 Lunaevaluates98%
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230)GPT-5.6 Lunaevaluates98%
MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314)Qwen 3.8 Maxtechnique used98%
Automatic bioinformatic software named entity recognition from literature (2608.19201)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Grokevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Gemini 3.7 Flashevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)GPT-5.6 Lunaevaluates90%
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220)GPT-5.6 Lunatechnique used99%
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199)Athena-Brain-8Bevaluates98%
Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290)Qwen 3.8 Maxevaluates98%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)Gemini 3.7 Flashevaluates95%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)DeepSeek-V4-Flash-0731evaluates95%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)Gemini 3.7 Flashevaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)GPT-5.6 Lunaevaluates99%
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)DeepSeek-V4-Flash-0731evaluates98%
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535)Qwen 3.8 Maxevaluates99%
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299)GPT-5.6 Lunaevaluates80%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grokevaluates98%
Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106)Qwen 3.8 Maxevaluates99%
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103)DeepSeek V4-Flashtechnique used99%
Abliteration Mitigation via Refusal Aliases (2608.18093)Gemma 4evaluates99%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Gemma 4evaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Qwen 3.8 Maxevaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Mistral OCR 4evaluates85%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Qwen 3.8 Maxevaluates99%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Mistral OCR 4evaluates99%
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161)Qwen 3.8 Maxevaluates98%
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591)Qwen 3.8 Maxtechnique used100%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)Qwen 3.8 Maxevaluates98%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324)Qwen 3.8 Maxevaluates98%
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142)Mistral OCR 4evaluates99%
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078)DeepSeek V4-Flashevaluates95%
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100)GPT-5.6 Lunaevaluates99%
Breaking the weakest link to evade vision language models (2608.18938)Qwen 3.8 Maxevaluates99%
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884)Qwen 3.8 Maxevaluates98%
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)Qwen 3.8 Maxevaluates98%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)GPT-5.6 Lunaevaluates80%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Qwen 3.8 Maxevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Gemini 3.7 Flashevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)GPT-5.6 Lunaevaluates99%
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273)GPT-5.6 Lunaevaluates95%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)GPT-5.6 Lunaevaluates98%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)Gemini 3.7 Flashevaluates98%
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684)Qwen 3.8 Maxevaluates98%
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638)GPT-5.6 Lunaevaluates95%
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588)GPT-5.6 Lunaevaluates95%
Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528)Qwen 3.8 Maxevaluates99%
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393)Qwen 3.8 Maxevaluates99%
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336)Qwen 3.8 Maxevaluates98%
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning (2608.17301)Qwen 3.8 Maxevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grok 4.6evaluates99%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Gemini 3.7 Flashevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)GPT-5.6 Lunaevaluates98%
FedPref: Federated Preference Learning for Structured Radiology Report Extraction (2608.16971)Qwen 3.8 Maxtechnique used99%
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification (2608.17247)GPT-5.6 Lunaevaluates98%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)GPT-5.6 Lunaevaluates99%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek V4-Flashevaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Solevaluates99%
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 (2608.14992)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Small Models Scout Bottleneck Order for Large-Model Data Control (2608.14936)Qwen 3.8 Maxevaluates95%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Soltechnique used98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)Qwen 3.8 Maxevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Solevaluates99%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)Gemini 3.7 Flashtechnique used98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Soltechnique used99%
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads (2608.15117)Qwen 3.8 Maxevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek V4-Flashevaluates99%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Solevaluates95%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Soltechnique used95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)Gemini 3.7 Flashevaluates98%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Solevaluates98%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Solevaluates99%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Soltechnique used95%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Solevaluates99%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxtechnique used85%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Solevaluates99%

Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

literature review agentsexpert evaluationpublished Aug 25, 2026 · arXiv:2608.21374

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

multiple-choice benchmarksleaderboard reliabilitypublished Aug 25, 2026 · arXiv:2608.21382

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

LLM-as-Judgetrust scoringpublished Aug 24, 2026 · arXiv:2608.21097

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

misinformation interventionscontent moderationpublished Aug 24, 2026 · arXiv:2608.20649

Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

AI evaluator auditingcounterfactual reasoningpublished Aug 24, 2026 · arXiv:2608.20938

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

LinkedQwen 3.8 Max
pathology foundation modelswhole-slide imagingpublished Aug 24, 2026 · arXiv:2608.21060

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

automated evaluationlanguage-model rankingpublished Aug 24, 2026 · arXiv:2608.20574

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

LinkedGrok 4.6
large language modelslegal advicepublished Aug 21, 2026 · arXiv:2608.20220

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

legal contractscontract scrubbingpublished Aug 21, 2026 · arXiv:2608.20204

ContractScrub: A benchmark for final review of legal contracts

LLM memorymemory benchmarkspublished Aug 21, 2026 · arXiv:2608.20202

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

large language modelssocial simulationpublished Aug 21, 2026 · arXiv:2608.19689

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

large language modelsair traffic controlpublished Aug 21, 2026 · arXiv:2608.19299

Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation

LinkedGPT-5.6 Luna
LLM evaluationself-preferencepublished Aug 20, 2026 · arXiv:2608.18091

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

text-space skill optimizationreference-free LLM judgespublished Aug 20, 2026 · arXiv:2608.18719

Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization

structured decompositionlanguage model evaluationpublished Aug 20, 2026 · arXiv:2608.18303

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

LLM-as-a-Judgerecommendation explanationspublished Aug 20, 2026 · arXiv:2608.18300

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

AI benchmarksleaderboard governancepublished Aug 20, 2026 · arXiv:2608.18117

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

language model benchmarkingpartial-credit evaluationpublished Aug 20, 2026 · arXiv:2608.18336

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

Linkedclaude-fable-5, claude-sonnet-5, claude-opus-5Qwen 3.8 Max
medical consultationlarge language modelspublished Aug 19, 2026 · arXiv:2608.17330

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

time series foundation modelszero-shot forecastingpublished Aug 19, 2026 · arXiv:2608.17299

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

AI evaluationautonomous scientific researchpublished Aug 19, 2026 · arXiv:2608.17271

ASI-Bench: At the Dawn of Artificial Superintelligence

large language modelshypothesis rankingpublished Aug 19, 2026 · arXiv:2608.17270

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

document intelligencevisual document parsingpublished Aug 18, 2026 · arXiv:2608.15064

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

large language modelsparenting advicepublished Aug 18, 2026 · arXiv:2608.14622

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

AI moral reasoningmoral value problempublished Aug 18, 2026 · arXiv:2608.14566

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

agentic systemsmodel routingpublished Aug 18, 2026 · arXiv:2608.14641

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

large language modelsmechanical engineeringpublished Aug 18, 2026 · arXiv:2608.14615

Large Language Models and their Awareness of Mechanics and Spatial Geometry

anchoring effectcognitive biaspublished Aug 17, 2026 · arXiv:2608.14320

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

large language modelsmmWave radarpublished Aug 17, 2026 · arXiv:2608.14179

Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding

enterprise data interfacessemantic planningpublished Aug 17, 2026 · arXiv:2608.13612

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

AI evaluationhuman-AI collaborationpublished Aug 17, 2026 · arXiv:2608.13577

AI Evaluation Should Work With Humans

autonomous systemsfailure discoverypublished Aug 17, 2026 · arXiv:2608.13719

Coverage Aware Active Evaluation for Failure Discovery with Paired Systems

mechanistic interpretabilityrepresentation learningpublished Aug 17, 2026 · arXiv:2608.13626

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

machine learningconstitutive modelingpublished Aug 17, 2026 · arXiv:2608.14063

Benchmarking data-driven material models on the classic Treloar dataset

Auto-Researchresearch process evaluationpublished Aug 15, 2026 · arXiv:2608.12788

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

LLM-as-a-judgereasoning trace evaluationpublished Aug 15, 2026 · arXiv:2608.12585

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

LinkedGPT-5.6 Luna
autonomous reasoningsymbolic AIpublished Aug 15, 2026 · arXiv:2608.12325

Position: Reasoning is a Learnable Rule-Based Process

research integrityLLM evaluationpublished Aug 15, 2026 · arXiv:2608.12345

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

medical visual question answeringvision-language modelspublished Aug 15, 2026 · arXiv:2608.12928

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

LinkedGPT-5.6 Luna
mathematical reasoningautomated theorem provingpublished Aug 13, 2026 · arXiv:2608.11941

OEIS Open: How many conjectures can language models turn into theorems?

LLM evaluationinference-time token budgetspublished Aug 13, 2026 · arXiv:2608.12150

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

large language modelsrecommendation evaluationpublished Aug 13, 2026 · arXiv:2608.11493

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

AI evaluationreal-world benchmarkspublished Aug 13, 2026 · arXiv:2608.11341

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

LinkedGPT-5.6 Luna
cybersecuritysecurity detectionpublished Aug 13, 2026 · arXiv:2509.16749

Evaluating LLM Generated Detection Rules in Cybersecurity

rubric-based evaluationLLM judgespublished Aug 13, 2026 · arXiv:2608.12097

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

LinkedGPT-5.6 Luna
financial reasoningstructured datapublished Aug 12, 2026 · arXiv:2608.11047

V-FiLLM: Verified Financial LLM Reasoning Benchmark

human action recognitionhierarchical intentionspublished Aug 12, 2026 · arXiv:2608.10765

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition

artificial lifeneuroevolutionpublished Aug 12, 2026 · arXiv:2608.10323

Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures

defensive LLMssocial engineeringpublished Aug 12, 2026 · arXiv:2608.10239

Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction

interactive analytic dashboardsdashboard generationpublished Aug 12, 2026 · arXiv:2608.10567

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

LLM evaluationmoral reasoningpublished Aug 11, 2026 · arXiv:2608.08061

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

benchmark contaminationAI evaluationpublished Aug 11, 2026 · arXiv:2608.07914

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

medical reasoningelectronic health recordspublished Aug 11, 2026 · arXiv:2608.07796

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

LLM benchmarksbenchmark verificationpublished Aug 11, 2026 · arXiv:2608.07762

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

LinkedDeepSeek-V4-Flash-0731GLM-5.3GPT-5.6 Luna
large language modelsretrieval-grounded systemspublished Aug 11, 2026 · arXiv:2608.07688

IntelliAudit: Using Large Language Models to Evaluate Audit Controls

large language modelsdocument repairpublished Aug 11, 2026 · arXiv:2608.07617

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

large language model evaluationgeospatial reasoningpublished Aug 10, 2026 · arXiv:2608.07411

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

large language modelscreative generationpublished Aug 10, 2026 · arXiv:2608.07243

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

LLM evaluationjudge panel dependencepublished Aug 10, 2026 · arXiv:2608.06940

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

automated item evaluationeducational assessmentpublished Aug 10, 2026 · arXiv:2608.06609

Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

human-AI coordinationcyber-physical systemspublished Aug 10, 2026 · arXiv:2608.06657

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

Bayesian decision-makingepistemic uncertaintypublished Aug 7, 2026 · arXiv:2608.05642

Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying

financial large language modelsinvestment decision-makingpublished Aug 7, 2026 · arXiv:2608.06108

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

personalized LLMsbehavioral personalizationpublished Aug 7, 2026 · arXiv:2608.05246

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

AI tutoringlarge language modelspublished Aug 7, 2026 · arXiv:2608.05411

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

LLM factuality evaluationclaim decompositionpublished Aug 7, 2026 · arXiv:2608.05228

TriQua: Reconciling Granularity and Context in Factuality Evaluation

embodied intelligencephysics simulationpublished Aug 7, 2026 · arXiv:2608.05948

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

NAVSIMre-simulation benchmarkspublished Aug 6, 2026 · arXiv:2608.04896

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

large language modelselectroencephalographypublished Aug 6, 2026 · arXiv:2608.04156

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

large language modelschess commentarypublished Aug 6, 2026 · arXiv:2608.04240

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

music quality assessmentreference-free evaluationpublished Aug 6, 2026 · arXiv:2608.04142

InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

AI evaluationsimulated userspublished Aug 6, 2026 · arXiv:2608.04205

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

chain-of-thought promptingzero-shot promptingpublished Aug 5, 2026 · arXiv:2608.03550

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

AI evaluationbenchmark infrastructurepublished Aug 5, 2026 · arXiv:2608.02996

On the missing benchmarks layer and a potential solution

electrocardiographyatrial fibrillation detectionpublished Aug 5, 2026 · arXiv:2608.03597

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

AI scientistsscientific reasoningpublished Aug 5, 2026 · arXiv:2608.03569

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

long-document question answeringbenchmark correctionpublished Aug 5, 2026 · arXiv:2608.03397

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

large language modelsforecastingpublished Aug 5, 2026 · arXiv:2608.03416

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

AI for researchautonomous experimental designpublished Aug 5, 2026 · arXiv:2608.03501

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

benchmark evaluationcapability measurementpublished Aug 5, 2026 · arXiv:2608.03219

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

scientific equation discoverysymbolic regressionpublished Aug 3, 2026 · arXiv:2607.28684

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

computational creativityinterpretive perspectivespublished Aug 3, 2026 · arXiv:2607.28644

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

clinical diagnosisemergency department encounterspublished Aug 3, 2026 · arXiv:2607.28788

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

federated learningfoundation modelspublished Aug 3, 2026 · arXiv:2607.28658

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

large language model evaluationoptimization modelspublished Aug 3, 2026 · arXiv:2607.29431

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

LLM evaluationbenchmark auditingpublished Aug 3, 2026 · arXiv:2607.28801

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

LLM-as-a-Judge evaluationhallucination detectionpublished Aug 3, 2026 · arXiv:2607.28641

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

AI Scientist systemsautomated peer reviewpublished Aug 3, 2026 · arXiv:2607.28631

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

language-model evaluationbenchmark validitypublished Jul 31, 2026 · arXiv:2607.26191

Position: Evaluation Scores Are Perishable Knowledge Claims

medical imagingradiologypublished Jul 31, 2026 · arXiv:2607.25589

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

synthetic usersLLM evaluationpublished Jul 31, 2026 · arXiv:2607.26348

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

AI evaluationbenchmark validitypublished Jul 31, 2026 · arXiv:2607.26159

When benchmark inferences do not compose: Projectibility in AI evaluation

large language modelschain-of-thought evaluationpublished Jul 31, 2026 · arXiv:2607.26102

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

large language modelseducational assessmentpublished Jul 31, 2026 · arXiv:2607.26067

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

masked diffusion language modelsevaluation protocolspublished Jul 30, 2026 · arXiv:2607.24763

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

large language modelsparaphrase robustnesspublished Jul 28, 2026 · arXiv:2607.22554

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

parallel programminghigh-performance computingpublished Jul 28, 2026 · arXiv:2607.22588

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

phone call assistantstarget-oriented dialogue systemspublished Jul 28, 2026 · arXiv:2607.22635

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

phonon stabilitycomputational materials sciencepublished Jul 28, 2026 · arXiv:2607.22573

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

large language model benchmarkingmultisensor hazard assessmentpublished Jul 24, 2026 · arXiv:2607.20476

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

TPU kernel optimizationJAXpublished Jul 24, 2026 · arXiv:2607.20466

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

AI agentskernel generationpublished Jul 24, 2026 · arXiv:2607.20518

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

large language modelspersonalizationpublished Jul 24, 2026 · arXiv:2607.20471

Benchmarking the Personalization Capabilities of Large Language Models

uncertainty estimationconfidence calibrationpublished Jul 23, 2026 · arXiv:2607.19367

Rethinking Uncertainty Evaluation in Large Language Models

out-of-distribution detectionbenchmark leakagepublished Jul 23, 2026 · arXiv:2607.19393

Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks

language model evaluationeconomic productivitypublished Jul 23, 2026 · arXiv:2607.19375

Economic Evaluations of Language Models

machine unlearningdistribution restorationpublished Jul 23, 2026 · arXiv:2607.19442

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

LLM ensembleswisdom of crowdspublished Jul 22, 2026 · arXiv:2607.18269

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

text-attributed graphsgraph neural networkspublished Jul 22, 2026 · arXiv:2607.19108

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

supply-chain logisticsshipment prioritizationpublished Jul 22, 2026 · arXiv:2607.18573

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

large language modelscode synthesispublished Jul 22, 2026 · arXiv:2607.18260

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis

persona consistencylong-term memorypublished Jul 21, 2026 · arXiv:2607.17191

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

large language modelscapability diagnosispublished Jul 20, 2026 · arXiv:2607.16122

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

ODRLpolicy evaluationpublished Jul 20, 2026 · arXiv:2607.15987

A Formally Grounded ODRL Evaluator: Implementation and Comparison

document extractioncorporate governancepublished Jul 20, 2026 · arXiv:2607.15879

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

large language modelsalgorithmic tradingpublished Jul 20, 2026 · arXiv:2607.15414

AI Trading: Evaluating Large Language Models for Technical Market Analysis

geospatial datasetsmulti-agent systemspublished Jul 20, 2026 · arXiv:2607.15781

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

AI evaluationhuman-aligned evaluationpublished Jul 17, 2026 · arXiv:2607.14673

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

code world modelsclassical planningpublished Jul 17, 2026 · arXiv:2607.14169

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

language-model evaluationhonesty evaluationpublished Jul 17, 2026 · arXiv:2607.14399

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

item response theoryAI evaluationpublished Jul 17, 2026 · arXiv:2607.15190

Can We Trust Item Response Theory for AI Evaluation?

channel foundation modelswireless communicationspublished Jul 17, 2026 · arXiv:2607.14975

CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models

argumentative essay scoringautomated writing evaluationpublished Jul 17, 2026 · arXiv:2607.14524

WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

large language modelsquestion answeringpublished Jul 14, 2026 · arXiv:2607.10562

CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

knowledge graphsinformation extractionpublished Jul 14, 2026 · arXiv:2607.10212

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

benchmarkingprompt wrapperspublished Jul 14, 2026 · arXiv:2607.09665

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

AI evaluationforecastingpublished Jul 14, 2026 · arXiv:2607.10972

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

battery agingdata curationpublished Jul 14, 2026 · arXiv:2607.09762

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

benchmark compressionprompt subset selectionpublished Jul 14, 2026 · arXiv:2607.09739

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

federated learningcontinual learningpublished Jul 13, 2026 · arXiv:2607.08784

HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning

large language modelsreverse engineeringpublished Jul 13, 2026 · arXiv:2607.07738

REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

health chatbotspatient simulationpublished Jul 10, 2026 · arXiv:2607.08625

The complexities of patient-centred conversational artificial intelligence

psychiatric clinical evaluationlarge language model benchmarkingpublished Jul 10, 2026 · arXiv:2607.08257

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

scientific lineage reasoninglineage-grounded idea generationpublished Jul 10, 2026 · arXiv:2607.08758

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

AI evaluationhuman-AI interactionpublished Jul 10, 2026 · arXiv:2607.08285

Psychological Competence as a Missing Dimension in AI Evaluation

LLM-as-judgeself-consistencypublished Jul 10, 2026 · arXiv:2607.08065

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

large language model agentspromptingpublished Jul 9, 2026 · arXiv:2607.07504

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

compact world modelsspatial relation groundingpublished Jul 9, 2026 · arXiv:2607.06925

Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix

AI evaluationpsychometricspublished Jul 9, 2026 · arXiv:2607.07040

Measuring Intelligence Beyond Human Scale

large language modelsmultifactor scoringpublished Jul 9, 2026 · arXiv:2607.06940

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

large language modelscultural competencepublished Jul 8, 2026 · arXiv:2607.05405

CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

small language modelsAI tutoringpublished Jul 8, 2026 · arXiv:2607.05571

CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

information retrievalscientific computingpublished Jul 8, 2026 · arXiv:2607.05443

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

Design Structure Matriceslarge language modelspublished Jul 8, 2026 · arXiv:2607.05985

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation