Blog

Open-weight AI models vulnerable to jailbreaks: why choosing open or closed is a governance question

A University of Waterloo and FAR.AI study found serious jailbreak weaknesses in open-weight LLMs. What this means for choosing between open and closed AI.

· By

Two identical grey steel cabinets side by side, the left open exposing cables, the right closed with a small lock, and a tablet resting between them.
Choosing between open-weight and closed AI models is a per-workflow governance question, not a simple safety question.Image: IamVera.ai — original editorial illustration

The choice between open-source and closed AI models is no longer a simple safety question but a governance question per workflow. A study by the University of Waterloo and FAR.AI (August 2026) shows that built-in safety mechanisms in 21 popular open-weight LLMs are relatively easy to switch off, which means open models pose extra misuse risk without their own runtime controls. At the same time, a security benchmark by TELUS Digital shows that the open-source model GLM 4.7 beats several closed models on safety.

Open-weight models give more grip on data paths, data residency and audit access, but place all compliance responsibility on the organisation. Closed models offer stronger vendor guardrails and documentation, but less transparency. The real question is which combination of model type, hosting, licences and verification layers makes the risk per workflow demonstrably manageable.

What did the University of Waterloo and FAR.AI study find about open-weight models?

Researchers at the University of Waterloo and FAR.AI tested, according to reporting by TechXplore on the open-weight LLM security study, a corpus of 21 popular open-weight language models for jailbreak and security vulnerabilities. The central finding: the built-in safety mechanisms of these models are relatively easy to switch off.

The underlying point is structural, not incidental. With open-weight models, anyone can download the weights, fine-tune them and redeploy them without central mitigations. Where a provider of a closed model can adjust guardrails centrally, that control layer disappears once the weights are public. In our assessment this does not mean that open models are unsuitable, but that the attack surface grows as soon as an organisation adds no additional runtime controls of its own.

Are open-source AI models safer or less safe than closed models?

The answer is more nuanced than the label suggests. The TELUS Digital GenAI Safety Model Benchmark shows that the open-source model GLM 4.7 performs better than several closed models in a large security benchmark. The measured vulnerability range ran from 1.3 to 93 per cent, which according to the benchmark relates mainly to model design, size and reasoning capabilities, not to the open or closed label.

Two observations from the same source are relevant in practice:

  • No tested model proved immune to attacks.
  • Smaller models were significantly more vulnerable, which is directly relevant for organisations wanting to run smaller open models on-premise.

Our editorial reading: safety does not follow from openness or closedness, but from the model design plus the controls an organisation puts around it itself. That makes it a question of LLM security and verification in 2026 rather than a brand choice.

Which governance and compliance choices go with open versus closed models?

A governance analysis by Areebi describes three adoption patterns that organisations use in 2026, each with its own division of responsibility:

  • Closed SaaS via API: strong, provider-managed guardrails and documentation, but less visibility into logging and model changes.
  • Self-hosted open weights in your own VPC or on-premise: maximum control over data path, data residency, audit access and fine-tuning, but full responsibility for security and compliance lies with the deployer.
  • Hybrid deployment: closed models for generic tasks, open on-premise models for strictly confidential data.

The Areebi analysis emphasises, regarding the EU AI Act, that open-source carve-outs in Article 53 do not mean that model openness exempts an organisation from governance obligations. Compliance is not a property of weights, but of the combination of model, hosting, contracts and control plane. That aligns with the broader shift we described earlier, in which AI governance shifts from principles to concrete control duties. For those who want to steer more deeply on these trade-offs, the topic hub on AI governance is the starting point.

Practical research by Futuriom, summarised by MarketScale on the basis of more than two hundred enterprise case studies, shows that combining proprietary data with open-source models on your own infrastructure can often be safer than generic frontier SaaS, precisely because data paths, logs and retention remain internally controllable.

How do I choose between open-source and closed AI models per workflow?

Taken together, the sources point to a decision per workflow rather than one organisation-wide choice. An academic comparison of on-premise open-source SLMs with commercial LLMs in a SOC and CSIRT context (arXiv, November 2025) found that closed models achieved higher classification accuracy for security incidents, while locally deployed open-source models offered advantages in privacy, cost and data sovereignty, because incident data does not leave the infrastructure. The advice from that study is explicitly hybrid.

On the basis of these sources, a workable trade-off per workflow is:

  1. Determine the sensitivity of the data path: where does the data go and which logs and audits are possible?
  2. Assess the security posture: how easy is misuse and which runtime controls do you have yourself, such as sandboxing, tool binding and logging?
  3. Match the model type: closed for more complex generic tasks where the risk allows it, open on-premise for strictly confidential or heavily regulated data.
  4. Set out licence and AI Act obligations: know what the licence permits and which carve-outs do or do not apply.
  5. Add verification where the risk is high: multi-model checks, human review and additional logging.

The choice for verification through multiple models that check each other is therefore directly related to the data path and the sensitivity of the task.

What role does a verification layer play in these model choices?

A verification layer does not solve the governance question, but it can help to make visible which choice is running in which workflow. For professionals in law, healthcare, finance and government that means: insight per workflow into which model type is active, which data path goes with it and which security and logging layers surround it.

In that light we position IamVera.ai modestly. Vera is not a chatbot and not its own language model, but a verification layer that can route a task through selected independent models and expose verification steps, corrections, disagreements and sources for inspection. That supports control and gives more visibility into possible errors, but does not replace verification of the output itself. The Semantic Privacy Shield can replace sensitive document values before AI processing on EU infrastructure with synthetic, session-only equivalents; the workflow is fail-closed, so that if the privacy check fails nothing is sent onward. More about this is on the page about the Semantic Privacy Shield and anonymisation. The professional final judgement always remains with the user.

Sources and references

  1. Major security weaknesses found in leading open-weight LLMsTechXplore (University of Waterloo & FAR.AI security study) · 2026-08-25
  2. GLM 4.7 Beats Several Proprietary AI Models In TELUS Digital Security StudyOpen Source For You (TELUS Digital GenAI Safety Model Benchmark) · 2026-05-27
  3. Open source LLMs vs proprietary models: the 2026 enterprise governance guideAreebi · 2026-05-20
  4. Enterprises are ditching frontier AI models for open-source alternatives to protect proprietary dataMarketScale (Futuriom enterprise case study summary) · 2026-08-18
  5. On-Premise SLMs vs. Commercial LLMs: Prompt Engineering and Incident Classification in SOCs and CSIRTsarXiv · 2025-11-18

Sources: The article relies on the open-weight LLM security study by the University of Waterloo and FAR.AI (via TechXplore), the TELUS Digital GenAI Safety Model Benchmark (via Open Source For You), the governance analysis by Areebi, Futuriom case studies via MarketScale and an arXiv study on on-premise SLMs versus commercial LLMs.

← All articles in this topic ← All articles