pytest-jev
a pytest plugin that asks one Jev Noul per plain-English claim about a reply (all claims in one request), passes a claim at p ≥ 0.8 and fails anything unsure, and adds Choice and Score checks; on its 12 example tests it matched Claude Sonnet 5's verdicts in 5.3 s vs 27.1 s at $0.00017 vs $0.0192 per run.
OWNER
allebee
INDUSTRY
LLM app testing
CATEGORY
Evaluation & Benchmarking
KIND
REPO
HOST
github.com
WHAT IT DECIDES
CHOICE
Picks one option from a fixed set
SCORE
Rates against ordered levels
NOUL
Answers a yes / no question
Inferred from this project's own one-line summary in the community list — it reads as a SCORE + NOUL decision. Read the source to confirm.
THE SHAPE OF A SCORE CALL
A generic skeleton for this question type, not this project's actual code.
from typesafe import TypeSafe
ts = TypeSafe()
result = ts.evaluate(
state=candidate,
questions={
"relevance": {
"type": "score",
"levels": ["low", "medium", "high"],
"instructions": "How relevant is this to the query?",
}
},
)
if result["relevance"] >= 0.7:
keep(candidate)MORE IN Evaluation & Benchmarking
Jev Web Analyzeranalyzes a public SaaS landing page as clean Markdown and asks Jev ten bounded `Choice` questions about first-visit understanding, returning inspectable findings for the first change to make.Jev Playgroundbenchmarks Jev against Luna, Haiku, and Gemini at choosing validated legal moves in explicit-state games, scoring decision quality and consistency across a sequence of moves.Jev vs Mistral and Gemini for event validationhead-to-head test of Jev against Mistral Small and Gemini Flash-Lite at validating local event listings.jev-research-evalreproducible eval harness plus field note for Jev Ultrafast research-browser tasks, with QC'd cases, a suite runner, and a report generator.