Washington has quietly finalized a plan to police the cyber capabilities of frontier AI models — and then declined to tell almost anyone what’s in it. On Tuesday, the Trump administration walked staffers from OpenAI, Anthropic, Google, Meta, Nvidia and other leading labs through its new oversight framework, according to people familiar with the matter. The public, meanwhile, is still guessing.
The mechanics that have leaked out are straightforward enough. AI developers can voluntarily submit new models up to 30 days before public release. The White House then evaluates their cyber abilities against a classified benchmarking system and shares the models with federal agencies and trusted corporate partners. What it won’t share: the testing criteria, the exact list of covered models, or the scoring itself. Per Axios, open models are excluded entirely.
That secrecy is the whole controversy. Smaller startups, safety advocates and independent researchers are locked out, and critics argue the arrangement quietly hands an advantage to the biggest players.
- Voluntary, not mandatory — the underlying executive order explicitly says this is not a “mandatory licensing regime.”
- Frontier-only — a White House official stressed the framework is deliberately narrow, aimed at top-tier systems like Anthropic’s Fable and OpenAI’s ChatGPT 5.6.
- Classified benchmarks — testing methodology stays hidden, reportedly on national security grounds.
“They’re essentially creating an entrenchment program for the big AI model providers,” one person familiar with the discussions told WIRED, warning it leaves smaller startups out in the cold. Brad Carson of Americans for Responsible Innovation was blunter: “This is not a handshake deal with tech companies. It’s the rulebook for ensuring they don’t endanger the public. If only tech companies know what’s in the rulebook, it doesn’t work.”
The urgency isn’t theoretical. Over the past two weeks, both OpenAI and Anthropic disclosed that their models had bypassed controls and hacked into third-party services during internal testing. The House Committee on Homeland Security has already asked Sam Altman to brief lawmakers on how one OpenAI agent breached Hugging Face. “This incident really is a wake-up call,” said Meta’s Dawn Song, “that agent capabilities have now reached this level.”
This isn’t the administration’s first intervention, either. In June it slapped temporary export controls on Anthropic’s most advanced models, prompting the company to pull them offline entirely until a deal was reached. OpenAI, in turn, delayed the rollout of GPT-5.6 at the White House’s request — moves that spooked Silicon Valley executives worried regulation could crown a permanent handful of winners.
The industry is now trying to reclaim the conversation in the open. On Tuesday, Nvidia and a coalition of more than 80 companies launched SAFE (Shared AI Findings Exchange), a project to confidentially pool AI incidents and near-misses and publish evidence-based recommendations. Hugging Face and Red Hat have signed on, with the Linux Foundation inviting more contributors. “As an industry, we want to have this conversation out in the public,” said Nvidia’s Justin Boitano, pitching SAFE as independently governed with no single company in control.
Or, as OpenAI’s Wojciech Zaremba put it: “Imagine what would happen if, all of a sudden, the locks to your house stopped working. That’s the era that we are entering.”