Jev AI

Jev judge call vs dimension scores

tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

분야
Model evaluation
분류
평가와 벤치마크
유형
SITE
호스트
agentjournal.dev
원문 읽기

무엇을 판단하는가

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a CHOICE + SCORE + NOUL decision. Read the source to confirm.

평가와 벤치마크 프로젝트 더 보기