Jev AI

jev-orderby-bench

measures whether a SQL ORDER BY over a Jev probability is defensible (pairwise inversion, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a pre-registered gate that jev-1.13.0 passes on 20 Newsgroups topics and fails four of six conditions on Amazon ESCI product relevance, and shows a DuckDB extension's default 40-row batching fails the ranking gate that one row per request passes.

作者
yodablocks
分野
Model evaluation
カテゴリ
評価とベンチマーク
種別
REPO
ホスト
github.com
GITHUB で見る

何を判断するか

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a SCORE + NOUL decision. Read the source to confirm.

SCORE 呼び出しの形

この質問型の一般的な骨格であって、このプロジェクトの実際のコードではない。

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=candidate,
    questions={
        "relevance": {
            "type": "score",
            "levels": ["low", "medium", "high"],
            "instructions": "How relevant is this to the query?",
        }
    },
)
if result["relevance"] >= 0.7:
    keep(candidate)
評価とベンチマークの他のプロジェクト