Justee 3 posts the #3 score on PRBench Legal, ahead of GPT 6 Astra and Claude Opus 5.5

By Max Zaykov, Founder · October 2, 2026

Key Takeaways

  • Justee 3 scored 53.3 on all 500 PRBench Legal tasks, the mean of three full runs (52.9, 53.8 and 53.3), graded with Scale AI's unmodified grader.
  • Against Scale's public leaderboard of 30 September 2026, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) score higher; GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) score lower.
  • On the 272 tasks never used during development Justee 3 scored 51.3, the more conservative estimate, and 42.7 on the 250-task hard subset.
  • PRBench does not verify citations. In a separate check of 150 references from 50 answers, a live source confirmed 57% and contradicted 3%.
  • Justee 3 is a product, not a model, and Scale has not verified this result. It is released in batches in the first half of October 2026 on all paid plans.

Justee 3 posts the #3 score on PRBench Legal, Scale AI's public benchmark of 500 real legal questions, in our evaluation with Scale's published grader. Across three full runs it scored 53.3 on all 500 legal tasks (52.9, 53.8 and 53.3). On Scale's public leaderboard of AI models, published on 30 September 2026, only two Muse Spark models score higher, and GPT 6 Astra and Claude Opus 5.5 score lower.

This is an experiment in what a purpose-built legal AI product adds to the models it uses. Below we set out the benchmark, exactly what we tested, the full results with their run-to-run spread, how the number sits against Scale's leaderboard, and one thing PRBench does not measure that matters a great deal to anyone relying on a legal answer: whether the legal references are right. Justee 3 is released in batches in the first half of October 2026, on all paid plans.

PRBench Legal is the legal set of Scale AI's Professional Reasoning Benchmark: 500 tasks written with lawyers who hold a JD, each graded by an o4-mini judge against an expert rubric of weighted criteria. In Justee's evaluation, run with Scale's unmodified grader (commit 7b2df0a), Justee 3 scored 53.3 on all 500 tasks, the mean of three full runs of 52.9, 53.8 and 53.3. It scored 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development. Against the leaderboard Scale published on 30 September 2026, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) score higher, while GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) score lower. Justee 3 is a legal AI product built on foundation models, not a model, so the comparison is not like for like. Scale has not reviewed or verified the result, and per-task scores are available on request.

Justee 3 PRBench Legal results: 53.3 on all 500 tasks, 42.7 hard subset, 51.3 on unseen tasks
Justee 3 on PRBench Legal, mean of three full runs, graded by Justee with Scale AI's unmodified grader. Not verified by Scale.

An experiment, not a model contest

PRBench was built to test AI models, and Scale's leaderboard ranks models. Justee 3 is not a model. It is a legal AI product built on top of foundation models, with its own legal research, live web research and way of preparing answers to legal questions. Comparing it with the models on that list is not like for like, and we do not claim Justee 3 is a better model than any of them.

We ran PRBench to answer a different question: how much does a purpose-built legal AI system add on top of the models it uses? The leaderboard shows what models score on their own. The distance between those scores and Justee 3's is a rough measure of what the product layer adds, measured on a public benchmark with a grader we did not write.

That question matters to us because it is the work we do every day. Justee is built to help people understand legal documents, from NDAs and employment contracts to lease agreements, and much of the value of that product lies in the layer around the model: finding the right law, researching it live, and checking and structuring the answer. You can read more about who we are on our about page.

About the PRBench Legal benchmark

PRBench, the Professional Reasoning Benchmark, is Scale AI's test of real professional questions. Its legal set has 500 tasks written with lawyers who hold a JD; 47% are written as a non-expert user would ask them. Each task comes with an expert rubric of weighted criteria (a median of 17 per legal task, 9,092 in all) that rewards correct, useful points and penalises harmful or wrong advice. About 30% of PRBench tasks are multi-turn conversations, and the questions span many countries and 47 US jurisdictions.

An o4-mini judge checks every criterion against the answer. The task score is the weighted total clipped to the range 0 to 1, reported here times 100. Scale reports that the judge agrees with human experts 80.2% of the time, close to the 79.6% at which two experts agree with each other. The hard subset is the 250 hardest legal tasks.

In plain terms, a PRBench Legal score of 53.3 does not mean that 53% of answers were correct. It means that, on average, an answer earned a little over half of the weighted rubric points available, after deductions for harmful or wrong advice. The rubrics are demanding by design: they list the many points an expert would want to see, and few answers hit every one of them.

What we tested

Justee 3 is the third major version of Justee chat. We tested it in its deep analysis mode as a complete product: the same service our users call, including its legal research and live web research, without changes for the benchmark. Each task went in as a conversation, earlier turns included, and only Justee 3's final answer was graded. Justee 3 will be released to our users in batches in the first half of October 2026, on all paid plans.

Testing the product rather than a bare model was deliberate. People do not ask a model in isolation: they upload a contract to Justee's compliance review, ask follow-up questions, and expect the answer to reflect the law where they live. A benchmark run that stripped out the research would have measured something our users never see.

PRBench Legal results for Justee 3

The table below shows every run. Justee's PRBench Legal evaluation shows a mean of 53.3 across all 500 legal tasks, 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development.

Justee 3 scored 53.3 on all 500 PRBench Legal tasks, the mean of three full runs, and the three runs differ by at most 0.9 points.

The three runs differ by at most 0.9 points (standard deviation between runs 0.45). The standard error of the 500-task mean from task sampling is 0.8.

PRBench is public, and 228 of its 500 tasks were used at some point while we developed Justee 3's answers. We report the full set because it compares with the public leaderboard, and 51.3 on the 272 tasks never used as the more conservative estimate for questions Justee 3 has not seen. Justee's development split data shows a gap of 4.4 points between the tasks used during development (55.7) and those never used (51.3), which is why we lead our own caveats with the lower figure.

Justee 3 on PRBench Legal, by run
TasksRun 1Run 2Run 3Mean
All 500 legal tasks52.953.853.353.3
Hard subset (250)42.843.242.242.7
Never used during development (272)51.251.651.051.3
Used during development (228)54.956.455.955.7

Source: Justee evaluation, three full runs on one fixed build of Justee 3 (deep analysis mode), graded by Justee with Scale AI's unmodified PRBench harness (commit 7b2df0a) and o4-mini judge. Scores are mean-clipped rubric scores times 100. Not verified by Scale AI.

See what Justee finds in your contract

Upload a contract and get a clause-by-clause review, typically in about 2 minutes.

Review a contract

Context: Scale's public leaderboard

Scale's leaderboard lists AI models, each evaluated by Scale. Justee 3's number is for a product, evaluated by us with Scale's grader, and Scale has not reviewed it. Justee 3 includes its own legal research and web research; we do not know which leaderboard entries, if any, used tools. With those caveats, the table below shows the ten highest published entries and three more that Scale marks as new: Gemini 3.8 Flash, GPT 6 Astra and claude-opus-5-5.

Only the two Muse Spark 1.x models score higher, at 61.6 and 57.1. Scale ranks entries by their intervals: an entry drops one place for every model whose whole interval lies above its own. Taking the spread of our three runs as Justee 3's interval, that rule puts Justee 3 third, the place Scale also gives claude-fable 5, Muse Spark and claude-opus-4-6 (Non-Thinking). Justee 3's score is higher than theirs by less than one point. GPT 6 Astra (48.36), Gemini 3.8 Flash (48.90) and claude-opus-5-5 (47.58) are clearly behind.

On Scale's leaderboard of 30 September 2026, Justee 3's 53.3 on PRBench Legal sits behind only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05), and 4.94 points above GPT 6 Astra (48.36).
PRBench Legal: Scale's leaderboard with Justee 3 for context
EntryScore (legal, full dataset)Note
Muse Spark 1.361.56New
Muse Spark 1.157.05
Justee 353.3New; product, evaluated by Justee
claude-fable 552.56
Muse Spark52.29
claude-opus-4-6 (Non-Thinking)52.27
Fable 5.151.60New
gpt-5.6-sol (max)50.50
gpt-5-pro49.89
o3-pro49.67
gpt-5.1-thinking49.33
Gemini 3.8 Flash48.90New
GPT 6 Astra48.36New
claude-opus-5-547.58New

Leaderboard values as published by Scale AI on 30 September 2026 (full dataset, legal), shown by Scale with its own intervals, for example 61.56±1.72. New marks the entries Scale labels as new, and Justee 3, which is new as well. Justee 3: evaluated by Justee with Scale's grader, not verified by Scale.

PRBench Legal leaderboard chart: Justee 3 at 53.3 behind two Muse Spark models, ahead of GPT 6 Astra
Selected PRBench Legal leaderboard entries published by Scale AI on 30 September 2026, with Justee 3's own-graded result for context.

What PRBench does not check: legal references

Running the benchmark taught us something about it. PRBench's grader does not verify citations. The judge reads the answer and decides, criterion by criterion, whether it meets the rubric; it has no access to legal sources. An answer can meet a criterion such as “cites the governing statute” with the wrong section number, and an invented case loses points only if a rubric criterion happens to cover it.

At the same time, the rubrics reward naming specific authorities: in one of our early runs, answers that cited case law scored 49 on average against 39 for answers that did not. Together, that pulls a system towards citing more, whether or not each citation is right. We saw it happen: in one experiment, a configuration that scored higher on PRBench named nearly three times as many authorities per answer as our standard chat, and our checks could not confirm 27% of them, against 8% for the standard chat. We did not ship that configuration.

For people using a legal AI product, a wrong citation is among the most serious errors: they repeat it to a landlord, an employer or a court. So we check citations outside the benchmark, with an automated check that looks each sampled statute, regulation and case up in live sources and asks whether a source confirms it says what the answer uses it for. “Confirmed” means a live source supports the reference; it is not a lawyer's review. We call the share of sampled references that a live source confirms the Justee Citation Confirmation Rate.

We ran it on a random sample of 50 Justee 3 answers to questions never used in development: up to three sentences per answer that name an authority, 150 in all. Justee's citation check found that a live source confirmed 57% of them and contradicted 3% (5). For the other 40% the search found nothing either way; many of those were table cells or sentence fragments rather than full citations, and counting only sentences with a specific reference, 65% were confirmed and 31% could not be verified. We read all five contradicted ones: two are clear errors (a wrong section number and a wrong statutory-instrument number), two are debatable uses of real authorities, and one was a false alarm caused by a sentence cut in half.

Justee's citation check found that a live source confirmed 57% of 150 sampled references from 50 Justee 3 answers and contradicted 3%; on the same questions, the standard chat's Justee Citation Confirmation Rate was 73%, with 2% contradicted.

On the same 50 questions, our standard chat named about half as many authorities per answer, and more of them were confirmed (73%, with 2% contradicted). Deeper answers cite more, and more of what they cite needs checking: the same pull we saw in the benchmark.

None of this affects the PRBench Legal score Justee 3 received, which is exactly what Scale's grader gave.

How we ran it, and what we disclose

What we disclose

Checklist for reading a PRBench Legal or other legal AI benchmark claim and checking cited authorities
Five questions to ask of any legal AI benchmark result, and the step that matters most: checking the authorities an answer names.

What this means if you use legal AI

Benchmarks such as PRBench Legal are useful because they are public, rubric-based and graded the same way for everyone. They are also easy to misread. Whatever tool you use, Justee or another, five questions help you judge any legal AI benchmark claim:

  1. Who ran the grader? A result graded by the vendor, even with an unmodified public grader, is not the same as one verified by the benchmark's owner. Ours is graded by Justee and has not been verified by Scale.
  2. Were the test questions seen during development? Public benchmarks leak into development. Ask for the score on unseen tasks; for Justee 3 on PRBench Legal that figure is 51.3.
  3. How many runs? One run can be lucky. We report three runs, which are at most 0.9 points apart.
  4. Model or product? A product with research tools and a bare model are answering under different conditions, so their scores measure different things.
  5. Are citations checked? PRBench does not check them, so a high score says nothing on its own about whether the authorities named are real and correctly used.

The last point is the one we would underline. Before you rely on any statute, regulation or case an AI names, look it up in a primary source. In the US, Cornell's Legal Information Institute and the official GovInfo collection publish federal statutes and regulations, and CourtListener lets you search court opinions. In the UK, legislation.gov.uk is the official home of Acts and statutory instruments. For organisations setting rules for AI use, the NIST AI Risk Management Framework is a practical starting point for documenting how a tool is tested and monitored.

Justee is designed to assist your judgment, not replace it. Use it to understand a severance agreement before a meeting, to compare two versions of a contract and see what changed, or to redact personal data before you share a document. If confidentiality matters, read about private mode and the document vault, and see our privacy FAQ. For anything you will rely on in a dispute, a filing or a negotiation, have a qualified lawyer review it.

Verification

Per-task scores for all three runs are available on request. We welcome an independent evaluation of Justee 3 on PRBench Legal. PRBench and its leaderboard are by Scale AI (scale.com/research/prbench). Justee 3's result was produced and graded by Justee with Scale's published grader and has not been verified by Scale.

PRBench rewards answers that name authorities, but its judge cannot check them. When a configuration that scored higher named nearly three times as many authorities and we could not confirm 27% of them, we did not ship it. For someone who repeats a citation to a landlord, an employer or a court, a wrong reference costs far more than a benchmark point.

Max Zaykov, Founder, Justee.ai

That decision is why Justee runs a separate citation check alongside PRBench Legal. The benchmark measures how well an answer meets an expert rubric; the citation check measures whether the authorities it names hold up in live sources. Both numbers are reported here, and neither replaces a lawyer's review of anything you plan to rely on.

Frequently Asked Questions

What is PRBench Legal?

PRBench, the Professional Reasoning Benchmark, is Scale AI's public test of real professional questions. Its legal set has 500 tasks written with lawyers who hold a JD, each graded against an expert rubric of weighted criteria by an o4-mini judge. According to Scale, about 30% of tasks are multi-turn and 47% are phrased as a non-expert would ask them.

What did Justee 3 score on PRBench Legal?

Justee 3 scored 53.3 on all 500 legal tasks, the mean of three full runs (52.9, 53.8 and 53.3), graded by Justee with Scale AI's unmodified grader. It scored 42.7 on the 250-task hard subset and 51.3 on the 272 tasks never used during development, which is the more conservative estimate for unseen questions.

Why is Justee 3 placed third?

Scale ranks entries by their intervals: an entry drops one place for every model whose whole interval lies above its own. Taking the spread of our three runs as Justee 3's interval, only Muse Spark 1.3 (61.56) and Muse Spark 1.1 (57.05) sit wholly above it, which puts Justee 3 third, a place Scale also gives claude-fable 5, Muse Spark and claude-opus-4-6 (Non-Thinking).

Is Justee 3 a better model than GPT 6 Astra or Claude Opus 5.5?

We do not claim that. Justee 3 is a legal AI product built on foundation models, with its own legal research and live web research, while Scale's leaderboard ranks models. Justee 3's 53.3 is above GPT 6 Astra (48.36) and claude-opus-5-5 (47.58) on Scale's published figures, but the comparison is not like for like. We treat the gap as a rough measure of what a product layer adds.

Does PRBench check whether legal citations are correct?

No. The judge decides whether an answer meets each rubric criterion without access to legal sources, so a wrong section number can still satisfy a criterion. That is why we check citations separately. In a sample of 150 references from 50 Justee 3 answers, a live source confirmed 57%, contradicted 3%, and found nothing either way for 40%.

Has Scale AI verified Justee 3's result?

No. The result was produced and graded by Justee using Scale's published grader, unmodified, and Scale has not reviewed or verified it. Per-task scores for all three runs are available on request, and we welcome an independent evaluation of Justee 3 on PRBench.

When can I use Justee 3?

Justee 3 is being released to users in batches in the first half of October 2026, on all paid plans. The results in this report were measured in its deep analysis mode. Justee's contract review, document comparison and PII redaction tools are available today and typically return a review in about 2 minutes per document.

Can AI replace a lawyer for legal questions?

No. Even the highest PRBench scores are well below a perfect score, and the benchmark does not check citations. Legal AI can help you understand a document, spot issues and prepare questions, but a qualified lawyer should review anything you rely on in a dispute, filing or negotiation.

Put Justee to work on your own documents

Review a contract, compare versions or redact personal data, typically in about 2 minutes.

Get started

About the author: Max Zaykov is the founder of Justee.ai and led the PRBench Legal evaluation reported here. This report describes an evaluation run by Justee; PRBench and its leaderboard are by Scale AI, which has not reviewed this result. Nothing in this article is legal advice.