Jev AI

Jev judge call vs dimension scores

tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

INDUSTRY
Model evaluation
CATEGORY
Evaluation & Benchmarking
KIND
SITE
HOST
agentjournal.dev
READ THE ORIGINAL

WHAT IT DECIDES

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a CHOICE + SCORE + NOUL decision. Read the source to confirm.

MORE IN Evaluation & Benchmarking