The White House announced on Monday that it had completed a voluntary framework for evaluating advanced artificial intelligence models, meeting the deadline set by President Trump's June executive order. But officials refused to disclose what the framework contains, which companies have seen it, or when it will be put into practice. The lack of transparency has raised questions about how the administration intends to govern frontier AI systems while keeping critical details hidden from the public and even from most policymakers.
“The voluntary framework outlined in the June 2nd executive order was complete by the deadline,” a White House official said, adding that “discussions with industry about next steps are underway.” The official spoke on condition of anonymity because the details are not public. The framework is meant to give the government a structure for determining whether AI models under development would be covered by the executive order, which created a 30-day pre-release review window for frontier models. That review window is designed to allow federal agencies to assess risks before a powerful model is released to the public.
What the framework contains — and what it doesn't
The framework itself is not designated as classified, according to the White House official. But the benchmarks used to assess cyber capabilities are classified, as is the threshold for determining which models fall under the review requirement. The threshold has been shared only with developers “as appropriate,” the official said. “Just because things are unclassified that doesn’t mean we are going to broadcast them to everyone,” the official explained, suggesting that even non-classified material could remain internal to the government and selected industry partners.
This opacity marks a significant departure from earlier AI policy efforts, which typically published evaluation criteria, model specifications, and safety guidance in government documents. The June executive order, titled “Preventing Threats to National Security from Artificial Intelligence,” established a 30-day pre-release review for what it called “frontier AI models” — systems with capabilities that could pose serious national security risks, particularly in cyber operations, biological weapons development, or other dual-use domains. The order tasked the National Security Agency and other agencies with developing benchmarks to measure whether a model reaches the threshold for review.
Those benchmarks remain classified. That means the public, researchers, and even members of Congress cannot independently verify whether the government’s threshold is too strict, too lenient, or appropriate. It also means that companies developing AI models have no public yardstick to know in advance whether their systems will trigger the review requirement. The White House has said it shares the threshold with individual developers “as appropriate,” but that discretionary process has drawn criticism.
Industry consultation and the Tuesday meeting
OpenAI, Anthropic, and Google provided feedback on a draft of the framework, the White House confirmed. The administration says it is engaging with “many more” industry partners beyond those three, though it declined to name them. A staff-level meeting with companies is scheduled for Tuesday to review the completed framework. The meeting is expected to include technical staff and policy leads from several AI developers, but the White House has not said which companies will attend.
The involvement of the three largest frontier AI labs suggests that the framework will directly affect the most prominent models in development. But the closed-door nature of the consultation process has fueled concerns that smaller companies and open-source developers are being left out of the conversation. Industry observers have noted that a voluntary framework with classified benchmarks could create an uneven playing field, where only companies with direct government relationships understand the rules.
The executive order frames the review as voluntary. However, the combination of classified benchmarks, undisclosed thresholds, and a 30-day government preview window creates what many legal scholars describe as a de facto gating mechanism. Because companies cannot publicly evaluate the criteria or challenge the government’s determination, they face a stark choice: comply with an opaque process or risk being cut off from federal contracts, procurement opportunities, or even face security clearance implications. This dynamic blurs the line between voluntary cooperation and coercion.
Gold Eagle and the cybersecurity connection
The framework is the companion piece to a new initiative called Gold Eagle, which the White House launched this month. Gold Eagle is designed to coordinate AI-powered cyber defense across federal agencies. It will use AI systems to hunt for vulnerabilities in government networks, detect intrusions, and automate some defensive responses. The model evaluation framework complements Gold Eagle by determining which AI models are powerful enough to require government review before release. In essence, Gold Eagle finds vulnerabilities, while the framework decides which models are powerful enough to be subjected to pre-release scrutiny.
This pairing highlights the administration’s focus on AI’s offensive and defensive capabilities. The classified benchmarks reportedly emphasize cyber capabilities, which suggests that the government is most concerned about models that can autonomously identify or exploit software vulnerabilities. Such abilities could be used for both cyber defense and cyber offense, making them a double-edged sword. The National Security Agency, which is responsible for signals intelligence and cybersecurity, has played a central role in developing the benchmarks, according to sources familiar with the matter.
The emphasis on cyber capabilities is not surprising given the growing use of AI in both offensive and defensive military operations. The Pentagon has invested heavily in AI-powered electronic warfare, autonomous surveillance, and decision-support systems. In the commercial sector, AI models are increasingly capable of writing complex code, finding bugs, and even suggesting exploits. Frontier labs like OpenAI, Anthropic, and Google DeepMind have all warned that future models could lower the barrier to cyberattacks, making advanced hacking tools available to non-experts.
A history of voluntary AI safekeeping
The new framework follows a series of voluntary AI commitments made by leading companies. In July 2023, the Biden administration secured voluntary safety commitments from seven AI companies, including OpenAI, Google, Microsoft, and Anthropic. Those commitments included internal and external security testing, information sharing, and watermarking of AI-generated content. The Biden administration also issued an executive order in October 2023 that required developers of certain high-impact models to share safety test results with the government. However, that order was later rescinded by President Trump, who replaced it with his own approach emphasizing innovation and reduced regulation.
The Trump administration’s June executive order is narrower in scope than its predecessor. It focuses specifically on national security threats, rather than broad AI safety and equity issues. The order gives the Department of Homeland Security and the Department of Defense a role in reviewing models that could affect critical infrastructure or military operations. It also establishes a streamlined process for companies to report concerns about potential misuse. The voluntary framework is the operational mechanism for that review process.
However, the voluntary nature of the framework has been questioned. While it is not a formal regulation, it carries significant weight because of the government’s purchasing power and its influence over the AI industry. Many companies are willing to submit to the review process voluntarily to maintain good relationships with federal regulators and to gain security clearances for government contracts. The White House has said the framework is voluntary because the administration does not want to stifle innovation. But critics argue that an opaque process is not genuinely voluntary — it is simply an unofficial mandate.
Transparency vs. national security
Policymakers and AI safety advocates expected to see details of the framework once it was completed. They have not. Several Democratic senators have sent letters to the White House requesting a briefing on the framework and asking for the benchmarks to be declassified to the extent possible. The responses have been terse, citing national security concerns. The White House official said that some details might be shared in private briefings to Congress, but no such briefings have been publicly announced.
The tension between transparency and national security is a long-standing issue in American governance. Classified programs are subject to oversight by the congressional intelligence committees, but that oversight is limited to a small number of lawmakers who must hold security clearances. The AI framework, however, affects a much broader swath of the economy. Tech companies, investors, and civil society groups all have legitimate interests in understanding how the government plans to evaluate cutting-edge AI systems. Keeping the criteria secret makes it impossible for these stakeholders to assess the fairness of the process or to plan for compliance.
The situation is especially troubling for open-source AI developers. Many open-source models are distributed freely, making it impossible to enforce a pre-release review requirement without restricting public access to the code. If a developer creates a model that meets the secret threshold and releases it as open source, the government might not be able to intervene before the model is publicly available. The White House has not explained how the framework would apply to open-source models, which remain a major source of AI capability outside the major labs.
International implications also loom large. The European Union is pressing ahead with the AI Act, which imposes binding obligations on high-risk AI systems. The United Kingdom has announced its own AI safety institute, which will evaluate models for public benefit. China has implemented export controls on advanced AI chips and is investing heavily in AI military applications. The US framework, by contrast, is largely invisible to the outside world. Foreign governments cannot know whether the US threshold aligns with their own risk assessments, which could complicate future international coordination on AI safety.
The de facto gate
What makes the framework particularly powerful is not the classified thresholds themselves, but the process that surrounds them. Under the executive order, companies developing a frontier model must notify the government 30 days before release. The government then evaluates the model against the secret benchmarks. If the model is deemed too risky, the administration can urge the company to delay release, add safety mitigations, or take other measures. While this request is not legally binding, the administration can apply pressure through federal contracts, procurement rules, and even export controls.
In practice, this gives the government a de facto approval role over the most advanced AI models. Companies are unlikely to risk releasing a model against the government’s wishes, particularly when they rely on federal agencies as customers or partners. The result is a gating mechanism that exists outside public scrutiny. No company can publicly litigate the threshold, because doing so would require revealing classified information. No independent researcher can verify whether the benchmark evaluations are sound. And no watchdog can determine whether the government is using its power fairly.
The White House official insisted that the framework is intended to enhance security, not to block innovation. “This is about protecting the American people from emerging threats while ensuring that the United States remains ahead in AI development,” the official said. But the lack of transparency has created a climate of uncertainty among developers. Several AI safety researchers, speaking on condition of anonymity, expressed frustration that they were being asked to trust the government without any way to evaluate the evaluation.
The question that lingers is whether a framework that nobody outside the government can read, built on benchmarks that nobody outside the NSA can see, qualifies as the transparency the executive order promised. The order itself said it would “promote innovation and transparency” in AI development. So far, the process has delivered innovation in secrecy but little in the way of transparency.