Jev AI

Jev judge call vs dimension scores

tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

分野
Model evaluation
カテゴリ
評価とベンチマーク
種別
SITE
ホスト
agentjournal.dev
原文を読む

何を判断するか

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a CHOICE + SCORE + NOUL decision. Read the source to confirm.

評価とベンチマークの他のプロジェクト