Jev AI

jev-ood-calibration

independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).

작성자
scienthoon
분야
Model evaluation
분류
평가와 벤치마크
유형
REPO
호스트
github.com
GITHUB에서 보기

무엇을 판단하는가

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a SCORE + NOUL decision. Read the source to confirm.

SCORE 호출은 이렇게 생겼다

이 질문 유형의 일반적인 뼈대이지 이 프로젝트의 실제 코드가 아니다.

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=candidate,
    questions={
        "relevance": {
            "type": "score",
            "levels": ["low", "medium", "high"],
            "instructions": "How relevant is this to the query?",
        }
    },
)
if result["relevance"] >= 0.7:
    keep(candidate)
평가와 벤치마크 프로젝트 더 보기