Jev AI

jev-acento

pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish `state` costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing `instructions` in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.

作者
marcosmartinez
分野
Language evaluation
カテゴリ
評価とベンチマーク
種別
REPO
ホスト
github.com
GITHUB で見る

何を判断するか

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

The community list's one-line summary does not say which question type this project uses. Read the source to find out — we would rather leave this blank than guess.

評価とベンチマークの他のプロジェクト