Blog

Ivo releases contract model Sage as open source: verify the local system, not just the score

Ivo released Sage as an open-source contract model with self-reported benchmark gains. Local adaptability means legal teams now own the validation and evidence

· By

Two printouts of the same contract lie side by side on a desk, the right one with pen markings, next to a magnifying glass and a closed confidential folder.
A reported benchmark score is only a starting point; legal teams must validate their own local system against their own contracts.Image: IamVera.ai — original editorial illustration

On 30 September 2026 Ivo released Sage as an open-source contract model with self-reported gains on its own LAB tasks. Because teams can now adapt the model locally, the burden of proof shifts: anyone deploying Sage must set up a reproducible validation and retention process before outcomes in high-risk contract work rely on that score.

What exactly did Ivo release with Sage, and what do the reported figures say and not say?

Ivo describes Sage as an open-source model for long-horizon, agentic contract work, built on DeepSeek V4 Flash and post-trained with reinforcement learning using 1,040 LAB tasks. On the public model card and evaluation record, Ivo reports results on 180 held-out LAB tasks, including a rise in the pass rate on the Contracts criterion from 0.701 to 0.913. The model is a LoRA adapter, a thin layer on top of the base model, that teams can download and run locally. On the public model card and evaluation record, Ivo reports results on 180 held-out LAB tasks, including a rise in the pass rate on the Contracts criterion from 0.701 to 0.913. The model is a LoRA adapter, a thin layer on top of the base model, that teams can download and run locally.

Importantly, these are Ivo's own reported results under Ivo's own test set-up, task split and scoring procedure. There is no independent replication. The figures show that the model scores better within that framework, not that it reasons correctly in legal terms on the contracts of a specific firm.

Why do self-reported benchmark gains not prove reliable contract analysis in my practice?

A higher score on an in-house benchmark does not by itself establish whether the model is reliable in your workflow. Independent research shows how large that gap can be, especially on the narrow, realistic review tasks that contract work consists of.

The ContractScrub benchmark by Yejin Bang and colleagues, presented at the ICML AI4Law workshop, tests the final review of contracts for error types such as misuse of defined terms, incorrect references and inconsistent wording. Leading models performed poorly there, despite strong scores on broader benchmarks; only one model achieved a macro-average recall of 0.75. The CLAUSE benchmark, published in the Findings of EACL 2026, generates over 7,500 perturbed contracts across ten categories and tests whether models can detect and explain subtle legal discrepancies. The authors argue that reliability in practice calls for adversarial, fine-grained tests, not broad claims about legal reasoning.

The Stanford Institute for Human-Centered Artificial Intelligence further warns that benchmark labels do not always measure what they claim to measure, and that benchmarks claiming to measure the same property can contradict one another. For that reason, a higher Sage score under Ivo's protocol says nothing in itself about accuracy, robustness or safety in a concrete legal workflow. See also our earlier analysis on independent AI verification.

What does the open-source release of Sage mean for directors, lawyers and CISOs?

Our analysis: because Sage is open and locally adaptable, the object of verification shifts from a fixed vendor service to the whole local system — base model, adapter, prompts, tools, data, context handling and the review process; therefore your team should not assess Ivo's service but its own assembled configuration, before first deployment. Because a LoRA adapter, prompt or context limit can be changed locally without the score changing with it, performance can quietly slip after an adjustment; therefore you fix a versioned configuration and retest after every change to the model, adapter, prompt or tools before that version reaches production. Our analysis is that, because responsibility for review and accountability may shift toward the deploying team rather than remain solely with an external vendor, the ultimately accountable director or partner could face responsibility for an error that nobody validated; therefore you should appoint a single owner who manages the validation, model versions and unresolved errors, and have qualified lawyers assess the errors rather than the model itself. Because ContractScrub and CLAUSE show that models stumble precisely on subtle, rare drafting errors, the greatest danger in our assessment is not the visible miss but the plausible redline that a lawyer wrongly trusts; therefore you build adversarial and long-context cases into your own test set and place an independent checkpoint between AI analysis and formal advice, as we described in a checkpoint against cascading hallucinations.

How do I set up a reproducible validation and retention process before I deploy Sage?

Approve Sage only through a versioned, local evaluation protocol that you repeat after every adjustment and whose evidence you retain. That makes an error traceable later and prevents a vendor score from being presented as your own accountability. This is our recommendation as an editorial team, derived from the sources named above.

  • Freeze the base model, the LoRA adapter, the inference settings, the prompt and the tool harness as one named version with recorded hashes.
  • Create a confidential, representative holdout set of your own agreements that is never used for tuning.
  • Test clause extraction, consistency between documents, negotiation restraint, execution of your playbook, source and citation fidelity, and escalation judgement.
  • Include adversarial and long-context cases, such as hidden discrepancies and unusual drafting.
  • Let qualified lawyers assess the errors and compare against your current human or vendor workflow.
  • Repeat the test after every local adjustment and retain model hashes, data provenance, configuration, outputs, reviewer decisions and unresolved errors.

This approach aligns, in our analysis, with broader requirements around the GDPR accuracy principle in generative AI and with work that records agentic AI execution traceably, such as the IETF draft standard for AI audit trails. Our core editorial point is that Ivo's figures are a starting point for your own, reproducible validation, not proof that Sage is reliable for your practice.

Sources and references

  1. Ivo-Sage model card and evaluation recordIvo AI · 2026-09-30
  2. Here's Everything We Announced at InscribeIvo AI · 2026-09-30
  3. ContractScrub: A benchmark for final review of legal contractsYejin Bang et al., arXiv / ICML AI4Law Workshop · 2026-08-20
  4. Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning CapabilitiesAssociation for Computational Linguistics, Findings of EACL 2026 · 2026-03-24
  5. The Tests That Grade AI May Be Getting It WrongStanford Institute for Human-Centered Artificial Intelligence · 2026-09-25

Sources: The article relies on Ivo's own model card and announcement of Sage, the academic benchmarks ContractScrub and CLAUSE, and a measurement analysis by Stanford HAI.

← All articles in this topic ← All articles