Jev AI

jev-acento

pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish `state` costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing `instructions` in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.

작성자
marcosmartinez
분야
Language evaluation
분류
평가와 벤치마크
유형
REPO
호스트
github.com
GITHUB에서 보기

무엇을 판단하는가

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

The community list's one-line summary does not say which question type this project uses. Read the source to find out — we would rather leave this blank than guess.

평가와 벤치마크 프로젝트 더 보기