Skip to main content

Can Jev replace your LLM judge? Part 2: Testing harder answers

· 6 min read
Takaaki Yayoi
Databricks
Yuki Watanabe
Software Engineer at Databricks

In Part 1, we compared Jev with GPT, Claude, and DeepSeek as judges for MLflow QA answers. Jev matched all 30 human labels, with the lowest measured latency and estimated cost in that test. We wanted to see whether it would still catch errors in answers that were almost right.

Recap: What is Jev?​

Jev is TypeSafe's model for structured decisions. Given a question, relevant documentation, and a candidate answer, it returns the probability that the answer is correct, without a written explanation. In Part 1, we used that probability to turn each decision into a pass or fail.

Challenging Jev with a harder test set​

We built 12 questions, each with one correct answer and two subtly incorrect ones, in both English and Japanese (72 answers total). Some answers named the right API but explained it incorrectly.

We compared typesafe/jev-1.13 with gpt-oss-120b using the same criterion: was each answer factually and technically correct given a documentation excerpt? For Jev, a probability of at least 0.5 counted as “correct.” The incorrect answers included wrong API names, reversed source and target roles, missing permissions, and correct statements followed by a false qualification. For three questions, we also paraphrased the correct answers to see whether different wording would cause a rejection.

Result​

We ran the full dataset twice. Jev agreed with human labels on 64/72 answers in each run. GPT-OSS-120B agreed on all 72.

Jev and GPT-OSS-120B: agreement with human labels was 88.9% (64/72) and 100% (72/72); median latency was 0.20 and 1.44 seconds.

Jev scored 32/36 in each language in both runs. It accepted every correct answer, including the paraphrases, and caught errors such as a wrong retention period or a reversed yes/no conclusion. Its mistakes were almost correct answers with one material error. Here are three incorrect answers Jev accepted, quoted from the English test set:

A correct clause with its meaning reversed

Use WHEN NOT MATCHED THEN INSERT. It applies to target rows that have no matching row in the source.

The clause is right, but the explanation reverses the roles. It inserts source rows that have no match in the target.

A missing permission

CREATE SECRET and USE SCHEMA on main.default are enough.

For a user who does not own the schema, this omits the required USE CATALOG permission.

A correct explanation followed by a wrong command

Question: Does a DELETE on a table with deletion vectors immediately rewrite the Parquet files?

No. The rows are physically removed later when you run VACUUM ... APPLY (PURGE).

The “No” is correct, but the command is wrong. According to the documentation excerpt, rows are physically removed when files are rewritten, for example by OPTIMIZE or REORG TABLE ... APPLY (PURGE).

Speed and cost​

Jev's median latency was 0.20 seconds, compared with 1.44 seconds for GPT-OSS-120B: about seven times faster in this setup.

Per-judgment latency for Jev and GPT-OSS-120B, with medians of 0.20 and 1.44 seconds respectively

Two runs, with 144 measurements per judge. Black lines mark the medians. Labels translated from the original chart.

Jev's estimated inference cost was about $0.020 per 1,000 judgments, excluding credit-purchase fees. GPT-OSS produced an average of 224 output tokens per judgment, including reasoning.

Send uncertain judgments to another model​

Jev missed some harder answers, but its speed and cost are still appealing. Could its probabilities tell us which judgments need a second check?

In the two benchmark runs, every false acceptance had a probability of roughly 0.5–0.8, while every correct answer scored at least 0.91. We calculated what would have happened if judgments in the 0.2–0.8 range had been sent to GPT-OSS for a second opinion. Fewer than 20% of judgments would have needed that second check, and the combined verdicts would have matched all human labels in this dataset.

RunJudgments sent to GPT-OSSAgreement after routing
112/7272/72 (100%)
213/7272/72 (100%)

This range was chosen after looking at the same data, so treat it as a starting point rather than a rule. In repeated calls on the same inputs, one false acceptance scored 0.82 and would have fallen outside the range.

A similar approach is discussed in JEV-as-a-Judge: Accept When Confident, Escalate When Unsure.

Run Jev with make_judge​

Sneak peek: MLflow 3.17

MLflow 3.17 has not been released yet. It will support Jev in built-in scorers and custom judges through the typesafe:/ model URI. The example below previews that integration and will not run on the current stable release.

Set TYPESAFE_API_KEY in your environment, then define the same correctness check as a boolean judge:

import mlflow
from mlflow.genai.judges import make_judge

jev_correctness = make_judge(
name="jev_correctness",
model="typesafe:/jev-latest",
instructions=(
"Is the answer in {{ outputs }} factually and technically correct for "
"the question and documentation in {{ inputs }}? Respect any version "
"specified in the question. Mark it incorrect if any material claim "
"or instruction is wrong."
),
feedback_value_type=bool,
)

eval_data = [
{
"inputs": {
"question": "What is the default retention period for VACUUM?",
"context": "The default retention period for VACUUM is 7 days.",
},
"outputs": "The default retention period is 30 days.",
}
]

result = mlflow.genai.evaluate(data=eval_data, scorers=[jev_correctness])

MLflow records Jev's true/false verdict as jev_correctness and its probability in the feedback metadata under typesafe.probability. Keep human labels separate from the inputs sent to the judge, then compare the verdicts with those labels to measure agreement.

The benchmark used typesafe/jev-1.13; the make_judge example uses typesafe:/jev-latest and slightly different instructions, so its results may differ.

Try it on your answers​

Jev was fast and inexpensive in this test, but its mistakes matter: an answer can sound right and still give a user the wrong command or incomplete permissions. Include those cases in your judge evaluation.

Start with your existing correctness criterion and a small labeled dataset. Run Jev alongside your current judge, inspect the false acceptances, and test any routing threshold on new examples. The MLflow evaluation quickstart can help you set up the comparison.