According to Tweakers, major AI firms are exploring a joint standards body for safety testing of frontier models. For your organisation this means you must be able to demonstrate per workflow which models you use, which tests they passed, the results, and how you translate those into local thresholds, usage limits and human oversight.
Tweakers reported on 14 September 2026 that major AI firms, including Anthropic, OpenAI and Google, are considering setting up their own standards body for safety testing. That ties in with broader European and Dutch discussions about how AI use can be made verifiable. In our assessment, this shifts the core of the concept of AI safety: from internal, hard-to-verify tests at a single lab to shared, multi-party evaluations that you, as a user, must take into account.
What exactly are Anthropic, OpenAI and Google discussing about a standards body?
According to an analysis by Pondero, based on reporting by The Information, working groups below CEO level at Anthropic, OpenAI and Google have been meeting regularly since July 2026 to design an industry-led standards body for AI safety. The focus is on pre-release testing: independent evaluations, safety reviews and standardised risk assessments for their frontier models.
What stands out is that this no longer stops at general statements about safety. The named sources describe concrete process discussions about how models are tested and by whom. That makes the subject practically relevant for anyone deploying these models in sensitive processes.
How does a FINRA-style testing regime for frontier models work under Hassabis's proposal?
In a research note, the Cloud Security Alliance describes the proposal that Demis Hassabis made on 14 July 2026 for a US Frontier AI Standards Body, modelled on FINRA, the industry-funded and federally supervised securities watchdog. According to the note, the core of that model is:
- Labs voluntarily submit models up to 30 days before public release for testing.
- A test battery focuses on dangerous capabilities in the areas of cyber, biology and deception.
- Independent technical experts carry out the evaluations.
- Approval could become mandatory for US deployment once the protocol proves robust enough.
This differs from the current, ad-hoc self-testing by labs: it establishes a shared procedure, external assessors and a path towards obligation, rather than relying on loose claims. For your own decisions about slowing down or admitting models, governance mechanisms per workflow for frontier AI are a logical addition.
Which government evaluations already exist for AI models?
The impression that only discussions are taking place is inaccurate. According to The Guardian, the US Center for AI Standards and Innovation (CAISI) at the Department of Commerce struck arrangements with Google DeepMind, Microsoft and xAI to assess early versions of new AI models before public release, with an emphasis on cyber and biosecurity risks. CAISI is the successor to the AI Safety Institute that was housed at the National Institute of Standards and Technology (NIST). NIST is the US standardisation institute within the Department of Commerce that develops standards and measurement methods, while CAISI now operates as the dedicated programme for predeployment evaluations. The Cloud Security Alliance research note links this programme to Hassabis's proposal and reports that CAISI has already carried out more than 40 pre-deployment evaluations and expanded voluntary testing arrangements to five labs: Google DeepMind, Microsoft, xAI, OpenAI and Anthropic.
In our assessment, this means that a future industry standards body would sit on top of an already growing layer of public testing capacity, with CAISI as an existing government evaluator embedded in the US standards system. You may therefore have to deal with multiple assessors per model.
What changes practically for organisations that use frontier models?
The shift that these sources together reveal is one from promises to verifiable testing governance. For organisations deploying frontier models in workflows involving confidential or high-risk information, a concrete question arises: can you demonstrate against which standards the models you use have been tested?
This also affects your procurement. When assessing suppliers, it helps to factor in these testing arrangements; see our explanation on vetting AI suppliers during procurement. After all, what a standards body or CAISI tests says nothing about how your own organisation then deploys the model. That translation remains your responsibility.
Which governance and verification tasks must you now record per workflow?
Based on the developments described, this is, in our assessment, a workable checklist to record per workflow:
- Which frontier models are in use in this workflow.
- Whether and how those models have been evaluated by CAISI or a future standards body.
- Which test results and residual risks are known, and where you obtain them.
- How those external evaluations have been translated into local controls: deployment limits, usage restrictions and points of human oversight.
- How you track incidents and model updates once tests or versions change.
This record-keeping ties in with the broader practice of making organisational controls demonstrable. You will find more depth in our topic hub on AI governance and accountability.
Where a verification layer can help is in making precisely that cross-section visible. Vera is not a chatbot and not its own language model, but a privacy-focused verification layer that can help map, per workflow, which model passed which tests, which evidence is available and how updates are tracked. The Semantic Privacy Shield is designed to replace sensitive document values with synthetic, session-only equivalents on EU infrastructure before the AI chain analyses the content; when a privacy check fails, nothing is sent onward. That supports control, but the professional final judgement always remains with you. The factual news value lies in the emergence of shared standards bodies and government evaluations such as those of NIST, not in the tool.
Sources and references
- Grote AI-bedrijven overwegen eigen standaardorganisatie voor veiligheidstests
- Anthropic, OpenAI, and Google have been meeting since July to design an AI safety standards body
- Hassabis's FINRA-for-AI Plan and Who Regulates Frontier Models
- US and tech firms strike deal to review AI models for national security before public release
Sources: The article draws on reporting by Tweakers and Pondero about the standards body, a research note by the Cloud Security Alliance on Hassabis's proposal and reporting by The Guardian on the CAISI arrangements.