Jev AI

jev-ood-calibration

independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).

作者
scienthoon
分野
Model evaluation
カテゴリ
評価とベンチマーク
種別
REPO
ホスト
github.com
GITHUB で見る

何を判断するか

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a SCORE + NOUL decision. Read the source to confirm.

SCORE 呼び出しの形

この質問型の一般的な骨格であって、このプロジェクトの実際のコードではない。

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=candidate,
    questions={
        "relevance": {
            "type": "score",
            "levels": ["low", "medium", "high"],
            "instructions": "How relevant is this to the query?",
        }
    },
)
if result["relevance"] >= 0.7:
    keep(candidate)
評価とベンチマークの他のプロジェクト