Jev AI

jevcal

fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, reports how much traffic still has to escalate to an LLM, and fails CI when a model update breaks the locked thresholds.

OWNER
abhixhek
INDUSTRY
Model evaluation
CATEGORY
Evaluation & Benchmarking
KIND
REPO
HOST
github.com
VIEW ON GITHUB

WHAT IT DECIDES

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a SCORE + NOUL decision. Read the source to confirm.

THE SHAPE OF A SCORE CALL

A generic skeleton for this question type, not this project's actual code.

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=candidate,
    questions={
        "relevance": {
            "type": "score",
            "levels": ["low", "medium", "high"],
            "instructions": "How relevant is this to the query?",
        }
    },
)
if result["relevance"] >= 0.7:
    keep(candidate)
MORE IN Evaluation & Benchmarking