According to tests by Harvey and Legora, GPT-6 Astra raises the quality of multi-step legal and financial checking work, but in trust-sensitive practice this is only usable once you can demonstrate per matter which model ran, which documents the agent read, which checks it performed and where a lawyer signed off. So treat Astra workflows as governed, auditable processes.
The occasion is an article by Artificial Lawyer of 7 September 2026, in which Harvey and Legora describe their first tests with OpenAI's new flagship model GPT-6 Astra. In our assessment, the most important thing here is not so much that a better chat interface appears, but that legal AI is shifting towards agents that carry out complete checking tasks under human oversight. That mainly changes what you have to record.
What do Harvey and Legora report concretely about Astra?
According to Artificial Lawyer's account, Niko Grupen, Head of Applied Research at Harvey, describes Astra as a clear quality improvement over GPT-5.6 Sol on complex legal tasks. Astra reportedly distinguishes source documents from drafts better, separates supported from unsupported assumptions and turns gaps into concrete drafting suggestions.
The case study by OpenAI on Legora puts a number on that. An agent powered by Astra carried out a financial tie-out across 41 documents in a single run, checked balances against trial balances and consolidation schedules, flagged planted discrepancies — including a discrepancy of £500,000 — and recorded each check as structured evidence for human review. OpenAI reports a nearly 40% improvement on this workflow over the previous model and roughly 3% on average on Legora's internal benchmark BAR.
Why does Astra shift legal AI towards agentic, end-to-end work?
Legora describes itself on its own product page as an agentic operating system for legal work, in which agents plan, execute, check and deliver. WebProNews places Astra in a broader context and speaks of OpenAI's direct entry into the legal workflow layer, with a model that processes long contexts, maintains state and carries out computer actions.
The practical difference for you: a lawyer no longer merely asks a question of an assistant, but delegates a multi-step task to an agent that works across documents, data sources and tools. That calls for thinking in terms of governed workflows and agents rather than isolated prompts. As the growing gap between firms operationalising AI shows, the distinction lies not in the model, but in how the work is set up.
Which safety and governance context accompanies Astra?
Astra does not arrive in a vacuum. The weekly legal AI briefing from Flank reports that an early Astra model exceeded the internal "Critical" threshold for cybersecurity capabilities in OpenAI's Preparedness Framework, after which OpenAI paused reinforcement-learning training for at least two weeks to expand red-teaming and monitoring. The same briefing notes that Harvey became the first legal AI company to be certified against AIUC-1, a standard for the security, safety and reliability of AI agents, and that Harvey's Memory feature is designed so that draft data does not train global models.
Our analysis: these signals mean that model upgrades and demonstrable safeguards must move in step. That aligns with broader obligations; see how audit rights and evidence obligations under the AI Act become contractual and technical requirements.
What must you be able to reconstruct and demonstrate per Astra task?
The performance gain is only usable in trust-sensitive environments once every Astra-driven task is as reconstructable and auditable as possible. We advise recording at least the following per matter:
- Which model version ran (Astra or a previous generation) and the reason for that choice.
- Which benchmark or evaluation evidence justified that model choice, for example internal suites such as BAR.
- Which documents the agent actually read in the run.
- Which checks were performed and which discrepancies or gaps came to light.
- Where a lawyer accepted, amended or overruled the suggestion, with a watertight sign-off.
These points align with existing expectations around working papers and audit trails. Astra's agentic use turns those expectations into technical design requirements. For outcomes that rest on text, claim-by-claim checking of AI outputs remains a sensible addition to the agent's logging.
How does a verification layer help make this visible?
A privacy-focused verification layer such as Vera can help operationalise this accountability, without itself being a language model or chatbot. Vera can route a task through selected independent AI models and expose verification steps, corrections, mutual disagreements and sources for inspection. That supports review and gives more insight into how an outcome came about; it does not promise correctness and does not remove the need to check for hallucinations.
For sensitive documents, the Semantic Privacy Shield can replace sensitive document values with synthetic, session-only equivalents on EU infrastructure before AI processing. The architecture is designed to send only anonymised content onward; if the privacy check fails, the document is not sent onward. The professional final judgement always remains with the lawyer. In our assessment, that is exactly the role that suits Astra: the model raises the ceiling, but the auditability of every step determines whether you can use it in high-trust work.
Sources and references
Sources: The article relies on Artificial Lawyer, OpenAI's case study on Legora, Flank's legal AI briefing, Legora's product page and analysis from WebProNews.