Jev AI

ASSAY-001

Independent pre-registered check of Jev calibration and type safety on Banking77 / CLINC150. Split verdict, full logs. Write-up: [donttrustme.ai](https://donttrustme.ai/assay-001.html)

OWNER
jourdanlabs
CATEGORY
Evaluation & Benchmarking
KIND
REPO
HOST
github.com
VIEW ON GITHUB

WHAT IT DECIDES

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a NOUL decision. Read the source to confirm.

THE SHAPE OF A NOUL CALL

A generic skeleton for this question type, not this project's actual code.

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=tool_call,
    questions={
        "is_risky": {
            "type": "noul",
            "instructions": "Could this call delete or overwrite user data?",
        }
    },
)
if result["is_risky"] and result["is_risky_confidence"] > 0.6:
    escalate_to_human(tool_call)
MORE IN Evaluation & Benchmarking