The Frontier

Reporting from the edge of applied AI.

An independent journal of artificial intelligence in practice.

See what's new

New releases

7 recent pieces

On the desk

Secondary picks

ModelsGround Truth

One million tokens fixes long-document recall

DeepSeek-V4 reports MRCR 1M MMR 83.5 and CorpusQA 1M ACC 62.0 at Pro Max. Strong on paper. Not a buyer corpus guarantee.

TLDR

Claim: stuffing full documents into a 1M-token window reliably surfaces buried constraints without retrieval. DeepSeek-V4 (arXiv:2606.19348, April 2026) scores MRCR 1M MMR 83.5 and CorpusQA 1M ACC 62.0 at Pro Max on HuggingFace, but Non-Think mode is 44.7/35.6, Claude Opus 4.6 leads both benchmarks, and the paper notes degradation beyond 128K. Verdict: overstated for buyer corpora without placement tests on the shipped checkpoint.

Policy & lawDispatch

DeepSeek-V4-Pro ships on Hugging Face under MIT, not a community carve-out

arXiv 2606.19348 and the model card both cite MIT on weights and repository code. Export and deployment diligence still sit with the buyer.

TLDR

DeepSeek posted DeepSeek-V4-Pro and DeepSeek-V4-Flash weights on Hugging Face under the MIT License, the model card states. Technical report arXiv:2606.19348 was published 26 April 2026. The license permits commercial use, modification, and redistribution without MAU caps or naming prefixes. Export controls on advanced computing hardware and destination rules still apply to whoever hosts the weights.

Work & moneyGround Truth

Enterprise AI is paying off at scale

McKinsey's 2026 AI Trust survey finds average RAI maturity 2.3 and only ~30% at level 3+ on governance. Not an EBIT proof point.

TLDR

Claim: enterprise AI is delivering material EBIT impact at scale across the market. McKinsey's State of AI Trust in 2026 (March 25, 2026) surveyed ~500 organizations: average RAI maturity rose to 2.3 from 2.0, but only about 30% reach level 3+ on strategy, governance, and agentic AI controls. Security and risk, not regulation, top barriers to scaling agents. Verdict: overstated as a market-wide EBIT claim without firm-level evidence.

ModelsThe Stack

SAM 3.1 multiplexes video objects; license is Meta SAM, not Apache

27 March 2026 checkpoints on Hugging Face `facebook/sam3.1`. Meta reports ~7x at 128 objects on one H100 versus SAM 3. Weights gated. GitHub stars UNKNOWN.

TLDR

Meta's `facebookresearch/sam3` repo dated 27 March 2026 released SAM 3.1 Object Multiplex and new checkpoints. `RELEASE_SAM3p1.md` reports about 7x speedup at 128 objects on a single H100 versus the SAM 3 November 2025 release, and VOS gains on 6 of 7 named benchmarks. `LICENSE` is the SAM License dated 19 November 2025, a limited non-exclusive grant with trade-control and no-reverse-engineering clauses, not Apache-2.0. Hugging Face `facebook/sam3.1` is the checkpoint host. Independent speed reproduction is UNKNOWN.

Compute & powerDispatch

H200 and MI325X China licenses go case-by-case, still not a general license

BIS final rule 2026-00789, effective 15 January 2026, reviews U.S. exports of those SKUs case by case if exporters certify U.S. supply, a 50 percent TPP cap, KYC, and U.S. third-party testing. Reexports stay denied.

TLDR

BIS published final rule 2026-00789 on 15 January 2026, effective the same day, changing license review for certain advanced computing commodities, named examples NVIDIA H200 and AMD MI325X, from presumption of denial to case-by-case for exports from the United States to end-users in China or Macau. Applicants must certify U.S. commercial availability, no delay to U.S. orders, no foundry diversion, aggregate TPP to China and Macau no more than 50 percent of U.S. end-use shipments of the same product, KYC and remote-access screening, and U.S. third-party testing. Reexports, in-country transfers, and shipments to D:5-headquartered entities remain presumption of denial. How many licenses BIS will grant is UNKNOWN.

ModelsThe Stack

MedGemma 1.5 4B adds CT, MRI, and slide patches under HAI-DEF terms

Google Research, 13 January 2026. Weights are not Apache-2.0. Clinical diagnosis is out of intended use. Internal accuracy deltas are Google's.

TLDR

Google Research posted MedGemma 1.5 4B on 13 January 2026 as an update to MedGemma 1 4B, expanding CT/MRI volumes, whole-slide histopathology patches, longitudinal chest X-rays, anatomical boxes, and lab-report extraction. The model card says use is governed by the Health AI Developer Foundations terms, last modified 15 November 2024: recipe and inference code Apache-2.0, weights and derivatives under HAI-DEF, with a ban on uses that could make Google a medical-device manufacturer. Outputs are not intended to inform diagnosis. Benchmark gains cited in the research blog are Google's internal figures.

Latest

14 stories

Agents & softwareGround Truth

JSON Schema makes agent tools production-ready

NIST AI 800-5 finds commenters agree agent threats are novel and existing controls need adaptation. Schema syntax is not the bar.

TLDR

Claim: valid JSON against a schema is enough to ship side-effecting agent tools. NIST Trustworthy and Responsible AI 800-5 (May 18, 2026) summarizes CAISI's agent-security RFI: commenters widely agreed agents present novel threats and fundamental cybersecurity practices require adaptation. The January 2026 RFI foregrounded indirect prompt injection and misaligned objectives. Verdict: overstated for workflows with side effects.

ModelsThe Record

NIST AI 800-5 synthesis finds agent threats novel but cybersecurity principles adaptable

Published May 18, 2026; summarizes CAISI RFI NIST-2025-0035 (comments due March 9, 2026).

TLDR

NIST Trustworthy AI 800-5 summarizes RFI responses on AI agent security. Commenters widely agreed agents pose novel threats that hinder adoption and that core cybersecurity practice must adapt—not merely be reused. Government roles cited: implementation guidance, information-sharing, and standards. COSAiS SP 800-53 overlays remain forthcoming; prescriptive control IDs are UNKNOWN until NISTIR 8605D drafts.

Policy & lawDispatch

Colorado enacts SB26-189, replacing its AI Act with ADMT notice rules

Official bill page lists the measure as enacted. Duties start 1 January 2027 if the attorney general finishes rules. June 2026 high-risk AI deadlines from SB24-205 are not the live calendar.

TLDR

Colorado SB26-189 repeals and reenacts the SB24-205 AI consumer-protection provisions with automated decision-making technology (ADMT) rules for consequential decisions, the General Assembly bill page states. Developer documentation, deployer notices, a 30-day adverse-outcome description, three-year records, and a consumer path to human review are in the enacted summary. The attorney general must adopt rules on post-adverse-outcome disclosures by 1 January 2027. Skadden's 9 June 2026 client alert states Governor Polis signed the act on 14 May 2026. Independent confirmation of the signature date on a governor's press page is UNKNOWN.

Compute & powerThe Record

DeepSeek-V4 paper reports 27% FLOPs and 10% KV cache at 1M tokens versus V3.2

arXiv:2606.19348 (April 26, 2026); MIT checkpoints on HuggingFace; CSA/HCA hybrid attention.

TLDR

DeepSeek-V4-Pro and V4-Flash use interleaved Compressed Sparse Attention and Heavily Compressed Attention for 1M-token contexts. At 1M tokens the paper reports V4-Pro at 27% single-token FLOPs and 10% KV cache versus DeepSeek-V3.2; V4-Flash at ~10% FLOPs and ~7% KV. MIT weights on HuggingFace. Dollar cost per million agent tokens on your stack is UNKNOWN until benchmarked.

Work & moneyField Notes

McKinsey skill-partnership math reframes headcount fights as workflow redesign

MGI report: ~57% of US work hours theoretically automatable today — a capability ceiling, not a job-loss forecast; April 2026 podcast walks the partnership frame.

TLDR

McKinsey Global Institute estimates currently demonstrated technologies could in theory automate activities accounting for about 57 percent of US work hours — roughly 44 percent via agents and 13 percent via robots — while stressing this is technical potential, not a job-loss forecast. More than 70 percent of in-demand skills appear in both automatable and non-automatable work. The McKinsey Podcast episode published 30 April 2026 discusses human–agent hybrid teams. Named customer reorg outcomes are UNKNOWN.

Policy & lawDispatch

BIS pushes Approved IC Designer applications to December 31, 2026

Federal Register 2026-06851 replaces the April 13, 2026 cutoff in Note 1 to ECCN 3A090.a. The rule took effect April 7.

TLDR

BIS published final rule 2026-06851 on 9 April 2026, effective 7 April 2026, extending the deadline to apply for Approved Integrated Circuit Designer status to 31 December 2026. The rule amends Note 1 to ECCN 3A090.a under the January 2025 advanced-computing due-diligence framework. Applicants received by that date may be treated as authorized IC designers for 180 days while BIS processes ERC review.

Work & moneyDispatch

OpenAI closes $122 billion round at an $852 billion post-money valuation

31 March 2026 company post: committed capital, not a cash-in-bank figure. Amazon, NVIDIA, and SoftBank anchored after a 27 February $110 billion step.

TLDR

OpenAI stated on 31 March 2026 that it closed a funding round with $122 billion in committed capital at an $852 billion post-money valuation. The 27 February 2026 post listed $110 billion in new investment at a $730 billion pre-money valuation, including $50 billion from Amazon, $30 billion from NVIDIA, and $30 billion from SoftBank. Enterprise share, token throughput, and revenue run-rate figures in those posts are company-reported. Independent audit of the round and of monthly revenue is UNKNOWN.

Agents & softwareField Notes

Vertex Gen AI eval pipeline replaces RAG demo scorecards

Google’s Vertex evaluation service and EvalTask score retrieval and generation separately; model-based rubrics replace hallway comparisons.

TLDR

Google Cloud documents a Gen AI evaluation service on Vertex AI: generate answers for a prompt set, score them with named rubrics, and log runs in Vertex AI Experiments. For RAG, operators use EvalTask datasets with prompt and response columns, then batch jobs for large golden sets. Groundedness and answer-quality metrics replace anecdotal pilots. Named customer before/after scores are UNKNOWN.

DeploymentField Notes

Copilot agents split on declarative vs custom engine before Control System gates

Microsoft docs: declarative agents inherit M365 compliance; custom engine agents bring your own orchestrator and hosting bill.

TLDR

Microsoft documents two Copilot agent paths — declarative agents using Copilot's orchestrator and models, and custom engine agents with external hosting — plus a Copilot Control System with security, management, and measurement pillars. Operators report choosing declarative for Graph-scoped Q&A and routing complex workflows to custom engines only after governance review. Named customer agent counts are UNKNOWN.

ModelsDispatch

Google ships Gemini 3.1 Pro in preview across app, Vertex, and Copilot

19 February 2026. Google cites a 77.1 percent ARC-AGI-2 score. General availability date is UNKNOWN. GitHub Copilot Business and Enterprise admins must flip a policy.

TLDR

Google's 19 February 2026 blog says Gemini 3.1 Pro is rolling out in preview for developers (Gemini API in AI Studio, Gemini CLI, Google Antigravity, Android Studio), enterprises (Vertex AI and Gemini Enterprise), and consumers (Gemini app and NotebookLM). Google reports a verified ARC-AGI-2 score of 77.1 percent, more than double Gemini 3 Pro. GitHub's changelog the same day puts 3.1 Pro in Copilot public preview for Pro, Pro+, Business, and Enterprise, with an admin policy gate for Business and Enterprise. Independent ARC-AGI-2 confirmation and a GA date are UNKNOWN.

ModelsDispatch

NIST CAISI opens an AI Agent Standards Initiative on three pillars

The 17 February 2026 launch ties industry-led protocols to agent security research. An RFI on agent security closed 9 March.

TLDR

NIST's Center for AI Standards and Innovation announced the AI Agent Standards Initiative on 17 February 2026. CAISI and ITL will work on industry-led standards, open-source agent protocols, and agent security and identity research. Stakeholder input was open through CAISI's Request for Information on AI Agent Security, due 9 March 2026, and ITL's agent identity concept paper, due 2 April 2026.

SecurityDispatch

NIST CAISI agent-security RFI closed March 9; summary analysis landed in May

Docket NIST-2025-0035 opened 12 January 2026. NIST Trustworthy and Responsible AI 800-5 summarizes responses published 18 May 2026.

TLDR

CAISI published a Request for Information on securing AI agent systems on 12 January 2026, docket NIST-2025-0035, with comments due 9 March 2026. The RFI focuses on threats distinct from conventional software: indirect prompt injection, poisoned models, and misaligned autonomous actions. NIST published summary analysis NIST AI 800-5 on 18 May 2026, reporting broad agreement that existing cybersecurity practice needs adaptation for agents.

Policy & lawDispatch

Texas TRAIGA took effect 1 January 2026; the Attorney General may investigate and sue

House Bill 149, 89th Regular Session. Prohibited uses, government and health disclosures, a sandbox, and a state AI council. Attorney General website mechanism due by 1 September 2026.

TLDR

The Texas Legislature's enrolled summary for House Bill 149, the Texas Responsible Artificial Intelligence Governance Act, lists an effective date of 1 January 2026. The act amends the Business and Commerce Code and Government Code. It sets disclosure duties for certain governmental and health-care users of AI systems, prohibits specified harmful developments and deployments, creates civil penalties with Attorney General investigative authority, and establishes a regulatory sandbox and a Texas Artificial Intelligence Council. Section 8 of the enrolled analysis requires the Attorney General to post a specified website mechanism by 1 September 2026.

Policy & lawDispatch

California SB 53 is in force: frontier labs must publish an AI framework

Chapter 138, effective 1 January 2026. Large frontier developers owe a public frontier AI framework and a transparency report at model deploy. The statutory compute cutoff is not restated from secondary trackers.

TLDR

California's Transparency in Frontier Artificial Intelligence Act (SB 53, Chapter 138) was approved 29 September 2025 and took effect 1 January 2026, per Business and Professions Code section 22757.12 as added by that chapter. A large frontier developer must write, implement, comply with, and clearly publish a frontier AI framework covering its frontier models. Before or concurrently with deploying a new or substantially modified frontier model, a frontier developer must publish a transparency report. The statute bars materially false or misleading statements about catastrophic risk. Whether a given lab's website posting satisfies the statute is UNKNOWN without comparing that posting to the code.