Models

New models, labs, evals, open weights — what changed in capability, cost, or license.

ModelsDispatch

Anthropic ships Claude Opus 4.8 at $5/$25 with cheaper fast mode

28 May 2026. Same list price as Opus 4.7. Fast mode $10/$50. System card: not past Mythos Preview on RSP. Prompt-injection tests vs 4.7 are mixed.

TLDR

Anthropic's 28 May 2026 post says Claude Opus 4.8 is available everywhere at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.7. Fast mode is $10/$50 and described as 2.5 times the speed and three times cheaper than prior fast mode. AWS said the same day it is on Amazon Bedrock and Claude Platform on AWS. The Opus 4.8 system card says the model does not advance the capability frontier beyond Claude Mythos Preview under Anthropic's Responsible Scaling Policy, and that in some agentic tests it was "somewhat less robust than Opus 4.7" on prompt injection until safeguards were applied.

ModelsThe Record

NIST AI 800-5 synthesis finds agent threats novel but cybersecurity principles adaptable

Published May 18, 2026; summarizes CAISI RFI NIST-2025-0035 (comments due March 9, 2026).

TLDR

NIST Trustworthy AI 800-5 summarizes RFI responses on AI agent security. Commenters widely agreed agents pose novel threats that hinder adoption and that core cybersecurity practice must adapt—not merely be reused. Government roles cited: implementation guidance, information-sharing, and standards. COSAiS SP 800-53 overlays remain forthcoming; prescriptive control IDs are UNKNOWN until NISTIR 8605D drafts.

ModelsGround Truth

One million tokens fixes long-document recall

DeepSeek-V4 reports MRCR 1M MMR 83.5 and CorpusQA 1M ACC 62.0 at Pro Max. Strong on paper. Not a buyer corpus guarantee.

TLDR

Claim: stuffing full documents into a 1M-token window reliably surfaces buried constraints without retrieval. DeepSeek-V4 (arXiv:2606.19348, April 2026) scores MRCR 1M MMR 83.5 and CorpusQA 1M ACC 62.0 at Pro Max on HuggingFace, but Non-Think mode is 44.7/35.6, Claude Opus 4.6 leads both benchmarks, and the paper notes degradation beyond 128K. Verdict: overstated for buyer corpora without placement tests on the shipped checkpoint.

ModelsThe Stack

SAM 3.1 multiplexes video objects; license is Meta SAM, not Apache

27 March 2026 checkpoints on Hugging Face `facebook/sam3.1`. Meta reports ~7x at 128 objects on one H100 versus SAM 3. Weights gated. GitHub stars UNKNOWN.

TLDR

Meta's `facebookresearch/sam3` repo dated 27 March 2026 released SAM 3.1 Object Multiplex and new checkpoints. `RELEASE_SAM3p1.md` reports about 7x speedup at 128 objects on a single H100 versus the SAM 3 November 2025 release, and VOS gains on 6 of 7 named benchmarks. `LICENSE` is the SAM License dated 19 November 2025, a limited non-exclusive grant with trade-control and no-reverse-engineering clauses, not Apache-2.0. Hugging Face `facebook/sam3.1` is the checkpoint host. Independent speed reproduction is UNKNOWN.

ModelsDispatch

Google ships Gemini 3.1 Pro in preview across app, Vertex, and Copilot

19 February 2026. Google cites a 77.1 percent ARC-AGI-2 score. General availability date is UNKNOWN. GitHub Copilot Business and Enterprise admins must flip a policy.

TLDR

Google's 19 February 2026 blog says Gemini 3.1 Pro is rolling out in preview for developers (Gemini API in AI Studio, Gemini CLI, Google Antigravity, Android Studio), enterprises (Vertex AI and Gemini Enterprise), and consumers (Gemini app and NotebookLM). Google reports a verified ARC-AGI-2 score of 77.1 percent, more than double Gemini 3 Pro. GitHub's changelog the same day puts 3.1 Pro in Copilot public preview for Pro, Pro+, Business, and Enterprise, with an admin policy gate for Business and Enterprise. Independent ARC-AGI-2 confirmation and a GA date are UNKNOWN.

ModelsDispatch

NIST CAISI opens an AI Agent Standards Initiative on three pillars

The 17 February 2026 launch ties industry-led protocols to agent security research. An RFI on agent security closed 9 March.

TLDR

NIST's Center for AI Standards and Innovation announced the AI Agent Standards Initiative on 17 February 2026. CAISI and ITL will work on industry-led standards, open-source agent protocols, and agent security and identity research. Stakeholder input was open through CAISI's Request for Information on AI Agent Security, due 9 March 2026, and ITL's agent identity concept paper, due 2 April 2026.

ModelsThe Stack

MedGemma 1.5 4B adds CT, MRI, and slide patches under HAI-DEF terms

Google Research, 13 January 2026. Weights are not Apache-2.0. Clinical diagnosis is out of intended use. Internal accuracy deltas are Google's.

TLDR

Google Research posted MedGemma 1.5 4B on 13 January 2026 as an update to MedGemma 1 4B, expanding CT/MRI volumes, whole-slide histopathology patches, longitudinal chest X-rays, anatomical boxes, and lab-report extraction. The model card says use is governed by the Health AI Developer Foundations terms, last modified 15 November 2024: recipe and inference code Apache-2.0, weights and derivatives under HAI-DEF, with a ban on uses that could make Google a medical-device manufacturer. Outputs are not intended to inform diagnosis. Benchmark gains cited in the research blog are Google's internal figures.