Jerk CheckJudged by Jev

Method · October 6, 2026

How we tested

Does Jev judge the facts, or whoever is telling the story? We told every fight from both sides and checked.

The stories

We invented 500 everyday disputes inspired by 16 themes common on r/AmItheAsshole: weddings, roommates, in-laws, group trips, pets and more. No real post was read or copied. Every person and detail is made up.

Each dispute was designed with a right answer: one person at fault, the other at fault, both, or nobody, 125 of each. Each person then told it twice: once plainly, once spun in their own favor. That is 2,000 stories. The writers never saw the designed answer, and fresh auditors checked every pair of tellings against the facts until each pair was clean.

Then the file was frozen before any story was judged. Its SHA-256 is 909ea3e30af3341eb874e104dfb129a41f5fdf308cc64f02f03a7152c3b5ecc5. Nothing was edited or dropped after that.

The judging

Every story went through this site's own code, one Jev call each, with no live site traffic or stored verdicts. Jev never saw the other side's version.

For comparison, the same 2,000 stories went to a general chat model (gpt-6-luna, the model that writes this site's one-sentence explanations) with a one-line prompt asking for the verdict.

What we found

Among pairs where both tellings got a verdict, Jev named the same person at fault from both sides 87% of the time when told plainly and 88% when both sides spun it. The general chat model managed 78% and 74%. On spun stories, it moved toward the storyteller in 20% of pairs. Jev did in 6%.

When Jev does flip, it more often turns against the teller than for them: 52 times against and 10 for in plain tellings. Spun tellings were close to even, 28 against and 26 for.

Speed and cost

We timed both judges on a random sample of 200 of the 2,000 stories, in the same session within the same hour, with 4 calls in flight for each. Jev averaged 167 ms per verdict. OpenAI's gpt-6-luna averaged 2,288 ms, about 14 times longer. Jev was faster on all 200.

At list prices read on October 6, 2026, a Jev verdict costs about $0.00005 and a gpt-6-luna verdict about $0.00007, roughly a third less for Jev. The whole 2,000-story run came to about $0.10 for Jev (estimated from the sample) and $0.14 for gpt-6-luna (recorded).

The timings come from one Mac on one home network, and gpt-6-luna's ranged from under a second to 11.5 seconds, so the exact milliseconds will vary. Either price can change. Jev is cheaper on average, not on every call: gpt-6-luna was cheaper on 62 of the 200. On OpenAI's slower Batch tier, which cannot answer a live page, gpt-6-luna would cost less than Jev. And the cost compares the verdict call only: this site also asks gpt-6-luna to write each one-sentence explanation.

What it does not show

Jev is not always right. Its weak spot is fights where both people behaved badly: there it often blames just one of them. These are invented stories with designed answers, one comparison model and one prompt, so they say nothing about every chatbot or about real disputes.