← Back to all posts
News

The NSA Now Benchmarks Frontier Models. The Rubric Is Classified.

August 4, 2026 · 01:10 UTC · News
The NSA Now Benchmarks Frontier Models. The Rubric Is Classified.

TL;DR

The White House confirmed on Monday that it met the 60-day deadline in the June 2 executive order Promoting Advanced Artificial Intelligence Innovation and Security, and that the voluntary framework for reviewing frontier AI models is finished. It then declined to publish it. The contents are unreleased, the list of who has read it is unreleased, the start date is unreleased, and the benchmark that determines whether a model is a "covered frontier model" in the first place is classified and maintained by the NSA. Staff met with OpenAI, Anthropic, and Google DeepMind on Tuesday to walk through it. Those three had already reviewed a draft.


What the order actually says

The public part is not in dispute, because the executive order text is published. Section 3 gives the Secretary of the Treasury, the Secretary of War (acting through the Director of NSA), and the Secretary of Homeland Security 60 days to do two things.

First, Section 3(a): "develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a 'covered frontier model'."

Second, Section 3(b): "design a voluntary framework with AI developers" through which a developer can ask the government whether a model under development meets that designation, and then "provide the Federal Government with access to covered frontier models ... for a period of up to 30 days before they plan to release such models to other trusted partners."

June 2 plus 60 days is August 1. That is the deadline the administration says it hit.

order to finished framework: 60 days Jun 2 Aug 1 Aug 3 Aug 4 order signed 60-day deadline framework done labs briefed contents withheld
Sixty days from signature to a finished framework nobody outside the room can read.

The part you cannot read

Per Axios, which broke the completion story, the administration would not describe the framework's contents, would not say who has seen it, and would not say when companies would begin using it. A White House official's explanation: "Just because things are unclassified that doesn't mean we are going to broadcast them to everyone."

CBS News got the same non-answer. An official confirmed completion and declined to share details.

The framework covers real operational obligations. Reporting describes it as spelling out how a developer determines scope, plus the confidentiality, cybersecurity, insider-risk, and intellectual-property terms that apply during the review window, and how the government and the developer jointly pick which trusted partners get early access. Those are exactly the terms an engineering and legal team would need weeks to work through. They are also exactly the terms nobody outside three companies can currently see.

disclosure ledger, as of aug 4 PUBLIC the executive order text the 30-day access window no-licensing disclaimer which agencies lead that it is finished WITHHELD the framework text the coverage threshold the cyber benchmarks who has read it when it takes effect
Everything load-bearing for a developer sits in the right-hand column.

Why the threshold is a cyber test, not a capability test

This is the detail most coverage skips, and it changes what the framework is. The trigger is not general capability, model size, or training compute. Section 3(a) scopes the classified benchmark specifically to "advanced cyber capabilities." The designation turns on whether a model is good at finding and exploiting software weaknesses.

That is a narrower and more defensible line than a FLOPs threshold, and it is drawn by the one agency with the most operational reason to care. It also means the criteria are never going to be published, because publishing them would tell every adversary precisely which offensive-security evals the US government uses and where the bar sits.

The practical effect for a developer is odd. You have a speed limit, the sign has been turned to face away from the road, and you find out you were over it when someone taps you on the shoulder. Legal analyses from Skadden and Mayer Brown both note that the threshold is set by classified process and shared with developers only as officials see fit.

section 3(b): a covered model on its way out the door developer asksam I covered? nsa benchmarkclassified up to 30 daysfederal access trusted partnersthen release
Voluntary at every step, and the second box decides whether the rest of it applies to you.

"Voluntary" is doing a lot of work

The order is emphatic on this point. Section 3(c) states that nothing in the section "shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models."

That is a real limit, and it is the reason this is not the licensing regime the industry spent 2023 arguing about. Nobody can be compelled to hand over a checkpoint, and nobody can be stopped from shipping.

But voluntary frameworks work by becoming the thing a reasonable actor is expected to have done. Once three of the largest labs have signed up, "did you go through the review?" becomes a procurement question, an insurance question, and eventually a liability question, without a single line of statute. A voluntary standard only functions if everyone can see what they are volunteering for, and right now most of the industry cannot.

What this means if you ship models

  • If you are not OpenAI, Anthropic, or Google, you have nothing to read. The three labs that helped edit the draft are the three that know the terms. Everyone else is waiting for a document that may never be published.
  • Open-weight releases are the awkward case. The window is defined as up to 30 days before release "to other trusted partners." Weights that go public go to everyone at once, and the trusted-partner staging model does not obviously map onto a Hugging Face upload.
  • The trigger is offensive-security capability. If you are fine-tuning for autonomous vulnerability discovery or exploit generation, you are closer to the designation than a team shipping a general assistant of the same size.
  • No enforcement mechanism exists. There is no penalty, no filing, and no registry. The pressure here is reputational and commercial, not legal.

The caveats, straight

Several load-bearing details rest on reporting rather than published text, and should be read that way. That OpenAI, Anthropic, and Google jointly reviewed a draft and submitted edits in late July comes from press reporting, not from a government document. The specific obligations attributed to the framework (confidentiality, insider risk, IP terms) likewise come from reporting on a document nobody has published.

What is directly verifiable is the executive order itself, the August 1 deadline arithmetic, the classified benchmarking mandate, the up-to-30-day window, the no-licensing disclaimer, and the administration's own confirmation that the work is done and will not be described. That is a strange combination of facts to be certain about.

Key Takeaways

  • The voluntary US frontier-model review framework was completed by its August 1 deadline, 60 days after the June 2 executive order, and the White House confirmed completion without publishing it.
  • Whether a model is a "covered frontier model" is decided by a classified NSA benchmarking process scoped to advanced cyber capabilities, not to size, compute, or general capability.
  • Participating developers give the federal government access to covered models for up to 30 days before release to other trusted partners.
  • Section 3(c) explicitly forbids reading any of this as a licensing, preclearance, or permitting requirement. Participation is voluntary and unenforced.
  • OpenAI, Anthropic, and Google reviewed a draft and met with staff on August 4. Everyone else in the industry is being asked to opt into a document they cannot read.
  • Open-weight releases fit poorly into a staged trusted-partner window, and no public guidance addresses that gap.

Sources: Executive Order, Promoting Advanced Artificial Intelligence Innovation and Security (June 2, 2026), Axios: White House finalizes AI framework behind closed doors, CBS News: Trump administration finalizes AI framework, Quartz: White House to review AI cybersecurity framework with top labs, The Next Web: The White House says its AI framework is done, Skadden: New AI Executive Order, Mayer Brown: President Trump Signs Executive Order on Advanced AI Innovation and Security

AIPolicyRegulationFrontier ModelsNSACybersecurityAI SafetyOpen Weights
CONSOLE
$