access, licence, and disclosure — never a single open/closed binary
Weights availability
unknown
Licence
unknown
Licence category
unknown
Commercial use allowed
unknown
Derivatives allowed
unknown
Redistribution allowed
unknown
Field-of-use restrictions
unknown
Training code availability
unknown
Inference code availability
unknown
Architecture disclosure
unknown
Training data disclosure
unknown
Deployability
raw values as extracted — free text, never inferred from branding
Self-hosting
unknown
On-premise
unknown
Edge deployability
unknown
API availability
unknown
Regions available
unknown
Benchmarks
as reported — vendor-supplied unless otherwise noted
No benchmark data on this release yet.
National-security pathway & evidence
from the assessed source item, when one is linked
Impact: lowConfidence: mediumNat-sec: noneprimary
Evidence limitations: The source is a very brief official announcement and provides no independent validation, technical details, usage data, or information about the model’s capabilities.
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specializationevaluates · 99% confidence
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluationevaluates · 98% confidence
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photographyevaluates · 99% confidence
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agentsevaluates · 99% confidence
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Adviceevaluates · 98% confidence
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agentsevaluates · 99% confidence
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helpingtechnique_used · 99% confidence
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretabilityevaluates · 98% confidence
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoningevaluates · 98% confidence
Large Language Models Show Metacognitive Sensitivity in Medical Reasoningevaluates · 99% confidence
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligenceevaluates · 100% confidence
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skillsevaluates · 98% confidence
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRLevaluates · 95% confidence
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolutionevaluates · 98% confidence
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrievalevaluates · 99% confidence
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judgesevaluates · 95% confidence
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Designtechnique_used · 95% confidence
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memorytechnique_used · 98% confidence
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routingevaluates · 95% confidence
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generationevaluates · 95% confidence
Automatic bioinformatic software named entity recognition from literatureevaluates · 90% confidence
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Modelsevaluates · 95% confidence
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scalingevaluates · 95% confidence
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backendsevaluates · 95% confidence
From Monolithic to Modular: Segment-level Automatic Prompt Optimizationevaluates · 98% confidence
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligenceevaluates · 98% confidence
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillationevaluates · 98% confidence
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?evaluates · 99% confidence
The Authority Expectancy Effect in Multi-User Conflictevaluates · 95% confidence
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)evaluates · 99% confidence
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometryevaluates · 80% confidence
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluationevaluates · 80% confidence
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasksevaluates · 95% confidence
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimizationtechnique_used · 98% confidence
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messagesevaluates · 95% confidence