Today
A ranked briefing from the ai signal desk.
Today's summary
Custom inference silicon is becoming a direct lever on agent economics and latency
The AI scaling bottleneck is broadening from accelerators to the system around them
↳ Linked from your briefing
What changed: The work presents multimodal large language models as a potential general-purpose alternative to specialist molecular encoders, with embeddings conditioned on both molecular information and natural-language context.
What changed: The authors report that marking untrusted spans as non-executable substantially reduces prompt-injection success while preserving utility and readability, including SEP separation rising from 24.3% to 96.5%, TensorTrust attack success falling from 34.8% to 6.6%, and zero compliance across four PIArena attack families.
What changed: The benchmark grounds evaluation tasks and scoring rubrics in real experimental protocol modifications rather than tasks elicited solely from experts; Claude Opus 5 achieved the highest reported normalized rubric score of 59.2%.
What changed: Compared with the Qwen3.5-4B base model, the authors report a 9.95% improvement in SciFact label prediction accuracy, a 3.7% increase in evidence quality, and a 19.63% reduction in evidence hallucination rate.
What changed: Under a shared evaluation harness, BRANCH improved accuracy in all 14 model-benchmark settings by an average of 5.98 percentage points and was the best-performing operator in 12 settings; the paper also identifies truncation recovery and paired scoring as important factors in evaluation outcomes.
What changed: It extends NL2SQL evaluation beyond simplified academic schemas by testing enterprise-scale complexity, cross-dialect generalization, and cases where executed queries return semantically incorrect results.
What changed: The work treats explicit revisions to target-defining rules as a structured maintenance signal rather than as ordinary statistical concept drift, reporting 92.3% accuracy and 90.2% Macro-F1 while reprocessing 14.7% of historical data and reducing average update latency relative to complete retraining.
What changed: The authors report that ExTS is competitive with or improves on task-specific tree-search baselines across several tasks, with an average relative gain of 5.5% using one fixed configuration.
What changed: KVBoost combines dual-hash cache keying, boundary repair, deviation-guided recomputation, KV quantization, adaptive chunking, and importance-weighted eviction; on Qwen2.5-3B it reports a 4.49x reduction in time-to-first-token compared with recomputation.
What changed: On the reported trade-finance and Tobacco-3482 evaluations, HIRA substantially improved Macro-F1 while reducing human corrections and LLM calls without updating model weights.
What changed: It adds a culturally grounded Arabic safety evaluation protocol and reports that most evaluated models failed to defend against roughly half of the unsafe prompts, particularly in weapons and illicit-substance categories.
What changed: The proposed evaluation reports per-level accuracy, magnitude-weighted stability, and family-specific collapse points instead of relying on aggregate accuracy, revealing recurring weaknesses involving conflicting instructions and impossible premises.
What changed: It presents a NotebookLM-enabled approach for rapidly generating quantitative geothermal benchmarks and describes an LLM-assisted case study for auto-parallelizing geothermal numerical models.
What changed: SAEM uses reasoning-stage detection, stage-aware caching, expert-aligned token repacking, and in-situ CPU execution to reduce memory pressure and data movement during inference.
What changed: The authors report consistent improvements across large language models on the STATE-ToxiCN and TBO benchmarks, including state-of-the-art results for bilingual multi-tuple extraction.
What changed: The authors report that capability feedback increased paired local-information success from 50% to 100%, while multi-step commitments improved held-out-template success by 6.9 points and reduced model decisions by 32%; a short DSL also reduced peer traffic and commitment time relative to free-form communication.
What changed: RO-PnR models user health literacy and belief commitment and reports the highest cost-adjusted utility across three datasets and three base models, with 30% fewer turns than always-probe baselines.
What changed: The authors report accuracy comparable to FP16 and up to 1.4x speedup in W4A8 and W8A8 configurations, with deployment validation on an Orin NX device.
What changed: It proposes linking persistent agent identities, cryptographically verified capabilities, execution performance, creator identities, and reputation within a single infrastructure.
What changed: The reported system improves Gemini 2.5 Flash-Lite material-selection accuracy from 37.5% with pure LLM reasoning to 90.0%, and achieves 75.0% printability across 96 physical validation trials with 88.9% task suitability among successfully printed samples.
What changed: Compared with outcome-level contrastive RLVR methods, LACL-GUI adds preferences within successful and failed trajectories to favour concise successful executions and distinguish failures by their divergence from successful trajectories.
What changed: It provides a large-scale benchmark and supporting historical box-office data for evaluating multimodal models on film narrative, reporting near-chance performance for many vision-language models and a maximum accuracy of 61.1% for audio-visual models.
What changed: On SWE-bench Verified, the authors report that routing arbitration increased the macro-average resolved rate from 44.9% to 48.2% for the gpt-oss family and improved over uniform selection on Qwen3.6.
What changed: The reported evaluation moves kernel optimization beyond isolated benchmarks toward deployment-aware end-to-end testing, with geometric-mean latency speedups of 3.91x on A100 and 6.98x on H100 across ten language-model inference workloads.
What changed: It reports evaluations across 27 model configurations, identifies six recurring failure modes, and presents prompting and supervised fine-tuning methods that improve performance by up to 20 percent.
What changed: The reported results indicate that residual representation can match dense retrieval quality with 30% fewer context tokens and lower scaling of per-query scoring work, while top-down progressive descent performs poorly for document selection.
What changed: It finds that schema validation is the dominant quality filter, ontology grounding provides an additional filter, staging-logic validation rejected no records in the tested fully gated corpus, and retrieval augmentation has strongly model-dependent effects.
What changed: It consolidates previously fragmented studies on the risks of recursively using AI-generated data to train subsequent models and identifies open research challenges.
What changed: The reported materials-science benchmark results show competitive answer accuracy with substantially fewer retrieved-context tokens, 2.7x lower end-to-end latency than prompt-all, and stronger tool selection, parameter validity, provenance, and licence grounding.
What changed: It reports validation across five benchmark cases spanning four PDE classes, claiming lower long-time extrapolation error than the numerical prior and better performance than the strongest competing baseline in each case.
What changed: The authors report a formal method intended to distinguish responsibility among agents and human stakeholders, detect blame-shifting attempts and remain invariant under forged evidence.
What changed: The approach replaces fixed memory compression with programmatic retrieval and transformation of persistent session state, while retaining lossless historical records and reporting strong results on LongMemEval_S, BEAM_10M, and LOCA_256K.
What changed: Across 57 repositories and 105 release transitions, the study finds that all evaluated transitions invalidate part of the prior skill set, while six frontier agents achieve only 29.9%–69.7% average-at-3 macro F1 on stale-content removal.
What changed: It defines a benchmark requiring models to extract requested textual attributes together with corresponding image bounding boxes, and reports that Q2IT improves on standalone vision-language models while still leaving a substantial performance gap.
What changed: The proposed SSKG approach produced a wider accuracy range of 44.1% to 85.2% and a monotone mastery gradient, compared with 96.8% to 100% accuracy from the tested prompt-based simulations.
What changed: It reports that distilled students can match teacher performance on adversarial text, reduce false alarms on harmless prompts, and achieve approximately 24 ms CPU classification latency for the encoder class.
What changed: It adds reusable environment- and rollout-orchestration layers that provision isolated tool environments and overlap long tool-call trajectories to improve GPU utilization, with integrations for veRL and slime.
What changed: The reported experiment found that just-in-time warnings without cognitive forcing reduced compliance with AI-steered recommendations containing dark patterns from 71.7% to 53.7%, while awareness and flagging generally showed limited improvement.
What changed: It proposes signed, offline-verifiable decision records containing hashed references to inputs, outputs, and evidence, linked through a SHA-256 chain and supported by a reference implementation and two-language conformance kit.
What changed: The authors report that directly optimised hybrid architectures, particularly MobileViTv2 models, outperformed heavier models and knowledge-distillation approaches for edge deployment, achieving up to 87.4% sensitivity, 86.5% specificity, and 97.2% negative predictive value against specialist labels.
What changed: The authors report that role assignment, particularly the final integrator, drives much of the observed performance and can produce a high-recall operating point without threshold tuning.
What changed: The authors report that the approach can synthesize selection heuristics for LU factorization that replicate expert-designed strategies.
What changed: Encoder adaptation increased regional accuracy from approximately 39.03% for zero-shot CLIP to 75.94–82.10%, while full fine-tuning reduced mean distance to the predicted region centre from 12.30 km to 3.86 km.
What changed: It provides a comparative framework-level analysis and reports that only one reviewed method supports frequency-domain explanations, only two evaluation metrics are time-series-specific, and identical XAI methods can produce substantially different explanations across frameworks.
What changed: The authors report that a design generated by the framework passed China Classification Society Approval in Principle and reduced steel mass and unit capital cost by 8.1% versus the human-optimized TuQiang baseline.
What changed: On EnterpriseRAG-Bench, the system reportedly improves the Overall score from 68.22 to 82.26 and reaches 86.50 Correctness across more than 500,000 documents and approximately 650 million tokens.
What changed: It replaces query-time SQL synthesis with selection among domain-aligned tools containing parameterized queries and server-side business logic, reporting a pooled mean score of 0.939 for the verticalized pack versus 0.666 for raw SQL and 0.605 for the generic pack.
What changed: Evaluation is expanded from final game artifacts or isolated tasks to three lifecycle stages assessed through interaction, behavioral tests, product criteria, and regression checks.
What changed: It adapts hack-verifiable environments to Terminal Bench, evaluates reward-hacking rates across frontier models, and releases the environments and agent traces.
What changed: It reports that harness choices can change individual model scores from bands rather than points, account for 95.7% of adjacent-model score gaps on average, and select different leaderboard winners.
What changed: It compares accuracy alongside runtime, energy use, output length, extraction failures, and statistical significance, finding that no model dominates and that Gemma3:4b produces more correct answers per watt-hour despite Qwen3:4b leading accuracy on two datasets.
What changed: It introduces generative, missing-signal, missing-type, and temporal prompts to model missingness-aware intra- and cross-modal relationships in EHR data.
What changed: Verifier output is reused during memory retrieval, conflict resolution, summarization, and archival rather than being applied only as a one-time admission filter.
What changed: ATHENA combines a weight-sharing supernet with cross-hospital architecture priors and multi-agent LLM search, and reportedly matches or outperforms four NAS baselines in 9 of 12 hospital-task evaluations at a search budget of 30.
What changed: It combines documentation-grounded code generation with executable feedback from ABB RobotStudio, exposing failures that static and semantic checks may miss.
What changed: On PlanBench, the authors report that the method increased average success from 35.5% for LLM+P to 70.8%, achieved 66.3% faithful success, and reduced semantic drift to 6.4%.
What changed: It reports that state-of-the-art vision-language models achieve only 24.1% average pair accuracy and presents Verdict Log-Odds Supervision as a post-training method that substantially improves open-weight backbones.
What changed: It reports that the attack achieved an average adversary-aligned response rate of 91.2% across four downstream tasks, including 86.6% on GPT-5.5, while a proposed memory-boundary defence reduced the rate to 80.6%.
What changed: It reports that most evaluated models show some activation-control ability and that this control can imperfectly evade several activation-based monitoring techniques in simple tasks.
What changed: It presents an integrated on-system alternative to cloud-GPU and conventional on-premises inference, claiming that sensitive data can remain within the LinuxONE hardware perimeter while achieving sub-two-second end-to-end RAG latency.
What changed: It reports that priors improve learning when they effectively reach the policy, that HL-Gauss critic training adds 14.4 in-domain points over MSE, and that one commonly used in-domain slice negatively correlates with cross-domain transfer while the hardest slice predicts transfer strongly.
What changed: It replaces static prompts reused across inference instances with item-specific prompts that adapt to context, constraints, evidence, and unresolved fields.
What changed: On the reported experiments, CompPO improved held-out accuracy over tuned GRPO on Qwen3-4B and improved greedy pass@1 on Qwen3-4B and Llama-3.1-8B-Instruct; the authors also report greater stability across PPO-grid runs.
What changed: The reported system improves balanced accuracy over matched baselines on temporal, medication-adverse-event, and recorded-order verification, while reducing false-support predictions under evidence masking and recovering source-traceable event chains when intermediate events are hidden.
What changed: It adds explicit multi-level orchestration, bounded evidence routing, specialist item scoring, profile auditing, targeted revision, and a persistent hierarchical evidence trace.
What changed: Across 500 environments in six domains, the reported feasible-task rate increased from 48.6% to 94.8%, while trained PPO policies showed stronger performance and improved transfer to WebArena, WebShop, and MiniWoB++ without LLM calls during evaluation.
What changed: The reported evaluation indicates that an optimised Whisper large-v3 pipeline reduced word error rate by 20.7% and 24.4% on recorded discussions, while MedGemma-RAG identified 2.3 times more MDT-concordant interventions than a proprietary cloud comparator without a significant difference in overall accuracy.
What changed: It provides an expert-calibrated evaluator, LitJudge, that reportedly raises agreement with human experts from Spearman's rho 0.467 for existing LLM-as-a-judge methods to 0.78, while finding that leading systems win only 23.0% of decisive matches against human drafts.
What changed: The benchmark shifts evaluation toward underspecified real scientific requests with attachments and no ground-truth solutions, finding that no model cleared the acceptance threshold under all three judges and that 47.6% of scored judgments fell below the threshold.
What changed: It provides a methodological framework linking geographic differences between source and target regions to transfer performance, finding that information shift and spatial shift have statistically significant and complementary explanatory power.
What changed: It extends authorship obfuscation from independently protecting individual documents to addressing correlations across multiple texts, including cross-genre attribution and verification attacks.
What changed: The authors report improved EnterpriseRAG-Bench performance over the strongest published baseline at scales from 10 million to 618 million tokens, with gains ranging from 12.23% to 4.66%.
What changed: It replaces direct averaging or heuristic aggregation of confidence estimates with soft evidence over candidate answers and a null state, and reports improved calibration on GSM8K, SVAMP, and GSM-Hard using Qwen, Mistral, and Gemma models.
What changed: The reported system improves offline action-prediction metrics and real-robot grasp success over pi0.5, including 11/30 versus 2/30 successful grasps at fast belt speed, while adding a 2.46-2.93% latency cost.
What changed: It adds a reproducible, executable-ground-truth evaluation framework for open-ended model outputs and reports Grok 4.6 with the highest point estimate at 65.1, although only 101 of 351 pairwise model comparisons were statistically resolved.
What changed: It proposes replacing fixed Top-K retrieval and unfiltered reinforcement-learning trajectories with dynamically sized evidence sets and penalties for low-confidence trajectories.
What changed: It reports a common execution interface with confidence gating and adaptive strategy selection for AI-enabled scalar, aggregate, and join operations, including a reported 358-fold cost reduction on a representative factorable join.
What changed: On a 10-target held-out split, five sampled GPT-4o policies achieved 0.589 Recall@10 versus 0.571 for the strongest single-feature Protenix binder ipTM baseline; target-conditioned GPT-5.4 policies reached 0.519 Recall@10 and 0.583 NDCG@10 on a three-target subset.