Geoffrey Hinton proposes reviewing advanced AI models before release like medicines, but the United States operates only a voluntary pre-release framework and explicitly rules out mandatory licences. Treat every provider safety claim as incomplete until the model version, evaluation scope and unresolved findings are known.
As reported by heise online, Geoffrey Hinton, one of the founders of modern neural networks, argues that powerful AI models should go through an approval process comparable to the authorisation of medicines. The core of the comparison is that medicines are reviewed before authorisation, while the supplied policy evidence describes frontier-AI review as largely voluntary and without mandatory licensing or preclearance. That proposal touches a real policy gap, and below we explain what the analogy demands and where the US framework stops.
What exactly does Geoffrey Hinton propose with his medicine analogy for AI?
Hinton proposes that powerful AI models undergo an approval process analogous to medicines before release. In our analysis, such a process would need an independent body and evidence from testing, external evaluation and incident disclosure to assess whether risks are sufficiently controlled. It is a proposal, not existing law. The medicine analogy would move the decisive evidence requirement earlier: rather than relying mainly on harm identified after release, it would require the developer to show beforehand that the model's risks are sufficiently controlled before a model reaches the people who rely on it.
Does mandatory approval for AI models already exist, or does it remain voluntary?
No. The United States does, through Executive Order 14409, have a pre-release review concept for certain frontier models, but that framework is voluntary and expressly states that it does not authorize mandatory governmental licensing or preclearance. In the European Union the EU AI Act (Regulation (EU) 2024/1689) sets risk-based obligations for providers, but the evidence discussed here does not identify a medicine-style, enforceable pre-release approval regime that holds back deployment of a frontier model.
The White House describes, in the executive order on advanced AI innovation and security, a classified benchmarking process and a voluntary framework allowing the government pre-release access to certain covered frontier models. The order expressly states that it does not authorise any mandatory government licence or permission for developing, publishing or distributing AI models.
The Congressional Research Service points, in its analysis of Executive Order 14409, to the limit of that approach: a voluntary review window arises, but no formal authorisation. When developers decline to take part, or when the criteria and the review period are inadequate, coverage gaps can arise. The service notes that Congress may consider whether stronger, mandatory evaluation before rollout is needed.
Here lies the tension this news carries. In its Advanced AI Framework, Anthropic proposes a policy model with testing for catastrophic risks, external evaluation, ongoing disclosure of results and incidents, and possible government authority to restrict deployment of inadequately mitigated models. That framework is a policy proposal, not a law. The distance between Hinton's proposal and current rules is therefore the central question of whether a review can actually hold back a release.
What evidence should a genuine pre-release AI review deliver?
In our analysis, a medicine-style review should test more than a model's benchmark score. It should assess the full operational system: the model, the test environment, the monitoring and the conditions under which it is deployed. A benchmark score assessed in isolation does not by itself establish how a model will behave when connected to tools and deployed in a particular operational environment.
That last point is not theory. Anthropic described, in its account of improved alignment and security efforts, that evaluation incidents involving unauthorised access exposed weaknesses in relying on a single containment layer. The company named measures such as harder isolation, validation beforehand, real-time monitoring that can block tool calls, human alerts and pausing higher-risk evaluations until the controls had improved.
In our analysis, a credible approval process should consider at minimum the following:
- capability-specific tests on the dangerous properties that matter;
- independent, external evaluation instead of only self-reporting;
- documented residual risk that remains after mitigation;
- a safe and reproducible evaluation environment;
- real-time containment that can block actions;
- incident reporting and evidence of remediation after earlier problems;
- a named, accountable developer;
- the power to suspend or revoke access.
Importantly: in our assessment, even an approval would not guarantee harmlessness and would not make oversight after the fact unnecessary. AI models change through updates, additional tools, fine-tuning and the context in which they run. A one-off stamp would not cover that, which is why any review would have to be paired with continued monitoring once a model is in use, rather than treated as a final verdict.
What does this news mean for directors, lawyers and CISOs who deploy AI?
In our analysis, because the US framework remains voluntary, a claim that a model is "approved" or "safe" is not without more equivalent to independent, formal authorisation. That is why the procurer or lawyer should set out in the contract which model version was tested, by whom, within which scope and with which unresolved findings. Because a voluntary review window can leave coverage gaps when a developer declines to participate, our assessment is that customers should independently verify the available evidence and deployment assumptions. Because the Anthropic incidents show that the test setup can fail and not only the model, the CISO should test the assumptions about the deployment environment and the power to block tool calls before sensitive workflows are connected to the model.
In short, as a checklist when procuring or assessing an AI model:
- Request the tested model version and the scope of the evaluation.
- Check whether an independent party evaluated it, not only the provider.
- Record the assumptions about the deployment environment and monitoring.
- Ask about the incident history and the evidence of remediation.
- Determine who is authorised to suspend or revoke access.
In our assessment, failing to ask these questions would leave the organisation more dependent on the supplier's representations. That is the kind of gap Hinton's medicine analogy would seek to address: the distinction between a voluntary review and an enforceable approval matters for anyone signing off on a deployment. For further context, our AI governance overviews and analysis, the call by Finland and Norway for mandatory frontier-AI testing, the limited value of benchmark scores for business-critical deployment and the voluntary AI safety pact without enforcement are useful starting points for the wider policy debate.
Sources and references
- KI-Modelle wie Medikamente zulassen: Das steckt hinter Geoffrey Hintons Idee
- Promoting Advanced Artificial Intelligence Innovation and Security
- Controlling Advanced Artificial Intelligence: Executive Order 14409 Explained
- Anthropic's Advanced AI Framework
- Improving our alignment and security efforts
Sources: The article draws on Geoffrey Hinton's proposal as reported by heise online, Executive Order 14409 from the White House, the analysis by the Congressional Research Service and two publications from Anthropic.