Justee 3 posts the #3 score on PRBench Legal, ahead of GPT 6 Astra and Claude Opus 5.5
By Max Zaykov, Founder · October 2, 2026
Key Takeaways
- Justee 3 scored 53.3 on all 500 PRBench Legal tasks, the mean of three full runs (52.9, 53.8 and 53.3), graded with Scale AI's unmodified grader.
- Against Scale's public leaderboard of 30 September 2026, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) score higher; GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) score lower.
- On the 272 tasks never used during development Justee 3 scored 51.3, the more conservative estimate, and 42.7 on the 250-task hard subset.
- PRBench does not verify citations. In a separate check of 150 references from 50 answers, a live source confirmed 57% and contradicted 3%.
- Justee 3 is a product, not a model, and Scale has not verified this result. It is released in batches in the first half of October 2026 on all paid plans.
Justee 3 posts the #3 score on PRBench Legal, Scale AI's public benchmark of 500 real legal questions, in our evaluation with Scale's published grader. Across three full runs it scored 53.3 on all 500 legal tasks (52.9, 53.8 and 53.3). On Scale's public leaderboard of AI models, published on 30 September 2026, only two Muse Spark models score higher, and GPT 6 Astra and Claude Opus 5.5 score lower.
This is an experiment in what a purpose-built legal AI product adds to the models it uses. Below we set out the benchmark, exactly what we tested, the full results with their run-to-run spread, how the number sits against Scale's leaderboard, and one thing PRBench does not measure that matters a great deal to anyone relying on a legal answer: whether the legal references are right. Justee 3 is released in batches in the first half of October 2026, on all paid plans.
PRBench Legal is the legal set of Scale AI's Professional Reasoning Benchmark: 500 tasks written with lawyers who hold a JD, each graded by an o4-mini judge against an expert rubric of weighted criteria. In Justee's evaluation, run with Scale's unmodified grader (commit 7b2df0a), Justee 3 scored 53.3 on all 500 tasks, the mean of three full runs of 52.9, 53.8 and 53.3. It scored 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development. Against the leaderboard Scale published on 30 September 2026, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) score higher, while GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) score lower. Justee 3 is a legal AI product built on foundation models, not a model, so the comparison is not like for like. Scale has not reviewed or verified the result, and per-task scores are available on request.

An experiment, not a model contest
PRBench was built to test AI models, and Scale's leaderboard ranks models. Justee 3 is not a model. It is a legal AI product built on top of foundation models, with its own legal research, live web research and way of preparing answers to legal questions. Comparing it with the models on that list is not like for like, and we do not claim Justee 3 is a better model than any of them.
We ran PRBench to answer a different question: how much does a purpose-built legal AI system add on top of the models it uses? The leaderboard shows what models score on their own. The distance between those scores and Justee 3's is a rough measure of what the product layer adds, measured on a public benchmark with a grader we did not write.
That question matters to us because it is the work we do every day. Justee is built to help people understand legal documents, from NDAs and employment contracts to lease agreements, and much of the value of that product lies in the layer around the model: finding the right law, researching it live, and checking and structuring the answer. You can read more about who we are on our about page.
About the PRBench Legal benchmark
PRBench, the Professional Reasoning Benchmark, is Scale AI's test of real professional questions. Its legal set has 500 tasks written with lawyers who hold a JD; 47% are written as a non-expert user would ask them. Each task comes with an expert rubric of weighted criteria (a median of 17 per legal task, 9,092 in all) that rewards correct, useful points and penalises harmful or wrong advice. About 30% of PRBench tasks are multi-turn conversations, and the questions span many countries and 47 US jurisdictions.
An o4-mini judge checks every criterion against the answer. The task score is the weighted total clipped to the range 0 to 1, reported here times 100. Scale reports that the judge agrees with human experts 80.2% of the time, close to the 79.6% at which two experts agree with each other. The hard subset is the 250 hardest legal tasks.
In plain terms, a PRBench Legal score of 53.3 does not mean that 53% of answers were correct. It means that, on average, an answer earned a little over half of the weighted rubric points available, after deductions for harmful or wrong advice. The rubrics are demanding by design: they list the many points an expert would want to see, and few answers hit every one of them.
What we tested
Justee 3 is the third major version of Justee chat. We tested it in its deep analysis mode as a complete product: the same service our users call, including its legal research and live web research, without changes for the benchmark. Each task went in as a conversation, earlier turns included, and only Justee 3's final answer was graded. Justee 3 will be released to our users in batches in the first half of October 2026, on all paid plans.
Testing the product rather than a bare model was deliberate. People do not ask a model in isolation: they upload a contract to Justee's compliance review, ask follow-up questions, and expect the answer to reflect the law where they live. A benchmark run that stripped out the research would have measured something our users never see.
PRBench Legal results for Justee 3
The table below shows every run. Justee's PRBench Legal evaluation shows a mean of 53.3 across all 500 legal tasks, 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development.
Justee 3 scored 53.3 on all 500 PRBench Legal tasks, the mean of three full runs, and the three runs differ by at most 0.9 points.
The three runs differ by at most 0.9 points (standard deviation between runs 0.45). The standard error of the 500-task mean from task sampling is 0.8.
PRBench is public, and 228 of its 500 tasks were used at some point while we developed Justee 3's answers. We report the full set because it compares with the public leaderboard, and 51.3 on the 272 tasks never used as the more conservative estimate for questions Justee 3 has not seen. Justee's development split data shows a gap of 4.4 points between the tasks used during development (55.7) and those never used (51.3), which is why we lead our own caveats with the lower figure.
| Tasks | Run 1 | Run 2 | Run 3 | Mean |
|---|---|---|---|---|
| All 500 legal tasks | 52.9 | 53.8 | 53.3 | 53.3 |
| Hard subset (250) | 42.8 | 43.2 | 42.2 | 42.7 |
| Never used during development (272) | 51.2 | 51.6 | 51.0 | 51.3 |
| Used during development (228) | 54.9 | 56.4 | 55.9 | 55.7 |
Source: Justee evaluation, three full runs on one fixed build of Justee 3 (deep analysis mode), graded by Justee with Scale AI's unmodified PRBench harness (commit 7b2df0a) and o4-mini judge. Scores are mean-clipped rubric scores times 100. Not verified by Scale AI.
See what Justee finds in your contract
Upload a contract and get a clause-by-clause review, typically in about 2 minutes.
Context: Scale's public leaderboard
Scale's leaderboard lists AI models, each evaluated by Scale. Justee 3's number is for a product, evaluated by us with Scale's grader, and Scale has not reviewed it. Justee 3 includes its own legal research and web research; we do not know which leaderboard entries, if any, used tools. With those caveats, the table below shows the ten highest published entries and three more that Scale marks as new: Gemini 3.8 Flash, GPT 6 Astra and claude-opus-5-5.
Only the two Muse Spark 1.x models score higher, at 61.6 and 57.1. Scale ranks entries by their intervals: an entry drops one place for every model whose whole interval lies above its own. Taking the spread of our three runs as Justee 3's interval, that rule puts Justee 3 third, the place Scale also gives claude-fable 5, Muse Spark and claude-opus-4-6 (Non-Thinking). Justee 3's score is higher than theirs by less than one point. GPT 6 Astra (48.36), Gemini 3.8 Flash (48.90) and claude-opus-5-5 (47.58) are clearly behind.
On Scale's leaderboard of 30 September 2026, Justee 3's 53.3 on PRBench Legal sits behind only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05), and 4.94 points above GPT 6 Astra (48.36).
| Entry | Score (legal, full dataset) | Note |
|---|---|---|
| Muse Spark 1.3 | 61.56 | New |
| Muse Spark 1.1 | 57.05 | |
| Justee 3 | 53.3 | New; product, evaluated by Justee |
| claude-fable 5 | 52.56 | |
| Muse Spark | 52.29 | |
| claude-opus-4-6 (Non-Thinking) | 52.27 | |
| Fable 5.1 | 51.60 | New |
| gpt-5.6-sol (max) | 50.50 | |
| gpt-5-pro | 49.89 | |
| o3-pro | 49.67 | |
| gpt-5.1-thinking | 49.33 | |
| Gemini 3.8 Flash | 48.90 | New |
| GPT 6 Astra | 48.36 | New |
| claude-opus-5-5 | 47.58 | New |
Leaderboard values as published by Scale AI on 30 September 2026 (full dataset, legal), shown by Scale with its own intervals, for example 61.56±1.72. New marks the entries Scale labels as new, and Justee 3, which is new as well. Justee 3: evaluated by Justee with Scale's grader, not verified by Scale.

What PRBench does not check: legal references
Running the benchmark taught us something about it. PRBench's grader does not verify citations. The judge reads the answer and decides, criterion by criterion, whether it meets the rubric; it has no access to legal sources. An answer can meet a criterion such as “cites the governing statute” with the wrong section number, and an invented case loses points only if a rubric criterion happens to cover it.
At the same time, the rubrics reward naming specific authorities: in one of our early runs, answers that cited case law scored 49 on average against 39 for answers that did not. Together, that pulls a system towards citing more, whether or not each citation is right. We saw it happen: in one experiment, a configuration that scored higher on PRBench named nearly three times as many authorities per answer as our standard chat, and our checks could not confirm 27% of them, against 8% for the standard chat. We did not ship that configuration.
For people using a legal AI product, a wrong citation is among the most serious errors: they repeat it to a landlord, an employer or a court. So we check citations outside the benchmark, with an automated check that looks each sampled statute, regulation and case up in live sources and asks whether a source confirms it says what the answer uses it for. “Confirmed” means a live source supports the reference; it is not a lawyer's review. We call the share of sampled references that a live source confirms the Justee Citation Confirmation Rate.
We ran it on a random sample of 50 Justee 3 answers to questions never used in development: up to three sentences per answer that name an authority, 150 in all. Justee's citation check found that a live source confirmed 57% of them and contradicted 3% (5). For the other 40% the search found nothing either way; many of those were table cells or sentence fragments rather than full citations, and counting only sentences with a specific reference, 65% were confirmed and 31% could not be verified. We read all five contradicted ones: two are clear errors (a wrong section number and a wrong statutory-instrument number), two are debatable uses of real authorities, and one was a false alarm caused by a sentence cut in half.
Justee's citation check found that a live source confirmed 57% of 150 sampled references from 50 Justee 3 answers and contradicted 3%; on the same questions, the standard chat's Justee Citation Confirmation Rate was 73%, with 2% contradicted.
On the same 50 questions, our standard chat named about half as many authorities per answer, and more of them were confirmed (73%, with 2% contradicted). Deeper answers cite more, and more of what they cite needs checking: the same pull we saw in the benchmark.
None of this affects the PRBench Legal score Justee 3 received, which is exactly what Scale's grader gave.
How we ran it, and what we disclose
- Grader: Scale's PRBench harness, unmodified (commit 7b2df0a, MIT licence), with the o4-mini judge (snapshot o4-mini-2025-04-16) and the mean-clipped score. No judge call failed in any run.
- Data: the public legal set of 500 tasks (CC BY 4.0) and its 250-task hard subset.
- Conversations: earlier turns are sent as real chat history; only the final reply is graded.
- Date: Justee 3 was told the date was 30 September 2026.
- Build: all three runs used one fixed build of Justee 3, recorded with each run.
- Repetition: three full runs, and the mean is reported.
What we disclose
- Development use of public tasks: 228 of the 500 tasks were used during development. On the other 272 the score is 51.3.
- A repeated run: one run was interrupted by an infrastructure problem before grading and was repeated in full on the same build. Its answers were never graded, and no graded result was left out.
- Web research within a time limit: in 3 of 1,500 answers, web research returned nothing within its time limit. We kept those answers as the product gave them.
- Judge noise: the judge agrees with experts about 80% of the time. Three runs reduce run-to-run noise, not judge bias.
- Coverage: about half of the hard subset concerns jurisdictions outside those covered by Justee's own legal research index. Those answers rely on web research.

What this means if you use legal AI
Benchmarks such as PRBench Legal are useful because they are public, rubric-based and graded the same way for everyone. They are also easy to misread. Whatever tool you use, Justee or another, five questions help you judge any legal AI benchmark claim:
- Who ran the grader? A result graded by the vendor, even with an unmodified public grader, is not the same as one verified by the benchmark's owner. Ours is graded by Justee and has not been verified by Scale.
- Were the test questions seen during development? Public benchmarks leak into development. Ask for the score on unseen tasks; for Justee 3 on PRBench Legal that figure is 51.3.
- How many runs? One run can be lucky. We report three runs, which are at most 0.9 points apart.
- Model or product? A product with research tools and a bare model are answering under different conditions, so their scores measure different things.
- Are citations checked? PRBench does not check them, so a high score says nothing on its own about whether the authorities named are real and correctly used.
The last point is the one we would underline. Before you rely on any statute, regulation or case an AI names, look it up in a primary source. In the US, Cornell's Legal Information Institute and the official GovInfo collection publish federal statutes and regulations, and CourtListener lets you search court opinions. In the UK, legislation.gov.uk is the official home of Acts and statutory instruments. For organisations setting rules for AI use, the NIST AI Risk Management Framework is a practical starting point for documenting how a tool is tested and monitored.
Justee is designed to assist your judgment, not replace it. Use it to understand a severance agreement before a meeting, to compare two versions of a contract and see what changed, or to redact personal data before you share a document. If confidentiality matters, read about private mode and the document vault, and see our privacy FAQ. For anything you will rely on in a dispute, a filing or a negotiation, have a qualified lawyer review it.
Verification
Per-task scores for all three runs are available on request. We welcome an independent evaluation of Justee 3 on PRBench Legal. PRBench and its leaderboard are by Scale AI (scale.com/research/prbench). Justee 3's result was produced and graded by Justee with Scale's published grader and has not been verified by Scale.
PRBench rewards answers that name authorities, but its judge cannot check them. When a configuration that scored higher named nearly three times as many authorities and we could not confirm 27% of them, we did not ship it. For someone who repeats a citation to a landlord, an employer or a court, a wrong reference costs far more than a benchmark point.
That decision is why Justee runs a separate citation check alongside PRBench Legal. The benchmark measures how well an answer meets an expert rubric; the citation check measures whether the authorities it names hold up in live sources. Both numbers are reported here, and neither replaces a lawyer's review of anything you plan to rely on.
Frequently Asked Questions
What is PRBench Legal?
PRBench, the Professional Reasoning Benchmark, is Scale AI's public test of real professional questions. Its legal set has 500 tasks written with lawyers who hold a JD, each graded against an expert rubric of weighted criteria by an o4-mini judge. According to Scale, about 30% of tasks are multi-turn and 47% are phrased as a non-expert would ask them.
What did Justee 3 score on PRBench Legal?
Justee 3 scored 53.3 on all 500 legal tasks, the mean of three full runs (52.9, 53.8 and 53.3), graded by Justee with Scale AI's unmodified grader. It scored 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development, which is the more conservative estimate for unseen questions.
Why is Justee 3 placed third?
Scale ranks entries by their intervals: an entry drops one place for every model whose whole interval lies above its own. Taking the spread of our three runs as Justee 3's interval, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) sit wholly above it, which puts Justee 3 third, a place Scale also gives claude-fable 5, Muse Spark and claude-opus-4-6 (Non-Thinking).
Is Justee 3 a better model than GPT 6 Astra or Claude Opus 5.5?
We do not claim that. Justee 3 is a legal AI product built on foundation models, with its own legal research and live web research, while Scale's leaderboard ranks models. Justee 3's 53.3 is above GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) on Scale's published figures, but the comparison is not like for like. We treat the gap as a rough measure of what a product layer adds.
Does PRBench check whether legal citations are correct?
No. The judge decides whether an answer meets each rubric criterion without access to legal sources, so a wrong section number can still satisfy a criterion. That is why we check citations separately. In a sample of 150 references from 50 Justee 3 answers, a live source confirmed 57%, contradicted 3%, and found nothing either way for 40%.
Has Scale AI verified Justee 3's result?
No. The result was produced and graded by Justee using Scale's published grader, unmodified, and Scale has not reviewed or verified it. Per-task scores for all three runs are available on request, and we welcome an independent evaluation of Justee 3 on PRBench.
When can I use Justee 3?
Justee 3 is being released to users in batches in the first half of October 2026, on all paid plans. The results in this report were measured in its deep analysis mode. Justee's contract review, document comparison and PII redaction tools are available today and typically return a review in about 2 minutes per document.
Can AI replace a lawyer for legal questions?
No. Even the highest PRBench scores are well below a perfect score, and the benchmark does not check citations. Legal AI can help you understand a document, spot issues and prepare questions, but a qualified lawyer should review anything you rely on in a dispute, filing or negotiation.
Put Justee to work on your own documents
Review a contract, compare versions or redact personal data, typically in about 2 minutes.
About the author: Max Zaykov is the founder of Justee.ai and led the PRBench Legal evaluation reported here. This report describes an evaluation run by Justee; PRBench and its leaderboard are by Scale AI, which has not reviewed this result. Nothing in this article is legal advice.