Jev AI

jev-acento

pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish `state` costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing `instructions` in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.

OWNER
marcosmartinez
INDUSTRY
Language evaluation
CATEGORY
Evaluation & Benchmarking
KIND
REPO
HOST
github.com
VIEW ON GITHUB

WHAT IT DECIDES

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

The community list's one-line summary does not say which question type this project uses. Read the source to find out — we would rather leave this blank than guess.

MORE IN Evaluation & Benchmarking