OpenAI announced Astra for Law on September 17, 2026: its newest frontier model, GPT‑6 Astra, configured specifically for legal work and wired to a purpose-built legal search index.
The announcement contains a number that most product launches would have buried. It's worth understanding, because how you read it is a small lesson in reading AI claims generally.

What Was Built
Astra for Law can search U.S. case law, statutes, regulations, court rules and administrative decisions across, in OpenAI's words, "a corpus of more than 230 million URLs, with sources added daily."
Some of that comes from an interesting place. OpenAI's work with the Free Law Project — the nonprofit behind CourtListener — brings in a case-law collection covering "more than 99.9% of published U.S. precedential case law." That is a public-interest project, built over years, now underpinning a commercial frontier product.
On top of the index sit custom instructions for legal analysis and writing: distinguishing a holding from a court's other observations, addressing cases that cut against the argument, explaining how a contract exception shifts risk.
The Number
OpenAI tested the complete setup on 200 U.S. legal research questions drawn from the private validation set of Vals AI's Legal Research Bench. The result:
"Astra for Law passed the evaluation's overall correctness check on 54.0% of questions, compared with 38.7% for GPT‑6 Astra using web search alone – a 40% relative improvement."
Take the good part first, because it's real. A 40% relative improvement from configuration alone — same underlying model, better index and instructions — is a large gain. The retrieval numbers underneath it are concrete: 24% more reference cases found on case-law questions, and up to 54% more relevant passages pulled from the correct court opinions.
Now the other half of the same sentence. A system built specifically for law, searching a dedicated legal corpus, at the highest reasoning effort, answers just over half of research questions correctly.
Both readings are true simultaneously. That's not a contradiction; it's what an honest benchmark looks like in a hard domain.
Three Qualifications
Whose benchmark, whose test. The Legal Research Bench belongs to Vals AI, a third party — and the questions came from its private validation set, which is methodologically the right choice, since a private set can't be trained on. But OpenAI ran the evaluation and published the results. The accurate phrase is "third-party benchmark, self-reported result." It is not "an independent evaluation found."
Highest reasoning effort. Both numbers are at the top setting for both systems. In ordinary use, at ordinary settings, expect less.
Research questions, not outcomes. The benchmark measures whether the model finds the right authority and whether its research answer meets evaluation criteria. That is upstream of advising a client, and a long way upstream of a filing.
Why 54% Lands Differently in Law
In most domains, a wrong answer is an inconvenience. In legal work it has a specific, documented failure mode: a confident citation to a case that doesn't say what it's cited for, or doesn't exist.
This isn't hypothetical. Courts have already sanctioned lawyers for filings containing AI-fabricated citations. The danger isn't that the model sounds unsure — it's that it sounds exactly as authoritative when it's wrong as when it's right.
Which is why a grounded legal index is the meaningful engineering here, arguably more than the model. If the system retrieves real opinions and shows them, a lawyer can check. OpenAI's own framing gestures at this: "reliable authorities the lawyer can examine for herself."
That sentence describes the correct workflow. The AI narrows; the lawyer verifies.
Who Actually Gets It
Not the public. Astra for Law "will be initially offered to selected law firms through Trusted Access in ChatGPT and Codex, and will be coming soon to the API." It appears in the model picker as "GPT‑6 Astra Law."
Eligible firms get Zero Data Retention on the API, and ChatGPT Enterprise usage excluded from human review by default — which matters when the input is privileged client material. Harvey and Legora are named as API partners, and the launch includes 26 partner-built plugins connecting ChatGPT to tools firms already use, such as Relativity, Clio, iManage and Thomson Reuters' HighQ.
Several firms have built their own tools on it: Sullivan & Cromwell an agreement analyzer, Ropes & Gray a deal diligence system, Cooley a tool for IPO preparation.
For Parents and Teachers
(This section is context, not the announcement.)
There's a transferable skill in this story, and it isn't about law.
When a company publishes a benchmark number, ask three questions. Who built the test? Who ran it? What does the score actually measure? Here the answers are: a third party, the company itself, and research questions rather than real outcomes. None of those answers is damning. All three change what the number means.
Notice too how easily 54% could be dressed either way. "A 40% improvement" and "wrong nearly half the time" describe the same result. A student who can hold both framings at once — and say which one the source itself emphasized — is reading well.
And there's a career note for any teenager eyeing law. The work being automated first is research retrieval: finding the right authority fast. The work that remains is the judgment about whether the authority actually supports the argument, and the professional responsibility for having checked. The value moves toward verification.
That's a broader pattern than law. When a tool gets good at producing plausible answers, the scarce skill becomes knowing how to check them.
Source: "Introducing Astra for Law," OpenAI, September 17, 2026: https://openai.com/index/astra-for-law/
All figures above are as reported by OpenAI in its own announcement. The Legal Research Bench is built by Vals AI; the evaluation run and reported results are OpenAI's. No independent reproduction has been published.
The sections marked as context — the three qualifications framing and the notes for parents and teachers — are ours.
Want practical AI guidance for parents and educators every week? Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter