Jev AI

Jev judge call vs dimension scores

tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

领域
Model evaluation
分类
评测与基准
类型
SITE
托管于
agentjournal.dev
阅读原文

它做什么判断

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a CHOICE + SCORE + NOUL decision. Read the source to confirm.

更多评测与基准项目