Jev AI

Jev vs GPT-4.1 on a synthetic survey

runs Jev and GPT-4.1 as the same 300 synthetic respondents over 24,596 paired Twin-2K-500 cells under criteria fixed in advance, finding that asking a yes/no item as `Noul` rather than `Choice` moves the result more than the gap between the two models, at a thirty-fourth of the cost. Write-up: [jjd-lab.github.io](https://jjd-lab.github.io/jev-synthetic-survey/)

作者
jjd-lab
分野
Survey research
カテゴリ
評価とベンチマーク
種別
REPO
ホスト
github.com
GITHUB で見る

何を判断するか

CHOICE

Picks one option from a fixed set

SCORE

Rates against ordered levels

NOUL

Answers a yes / no question

Inferred from this project's own one-line summary in the community list — it reads as a NOUL decision. Read the source to confirm.

NOUL 呼び出しの形

この質問型の一般的な骨格であって、このプロジェクトの実際のコードではない。

from typesafe import TypeSafe

ts = TypeSafe()
result = ts.evaluate(
    state=tool_call,
    questions={
        "is_risky": {
            "type": "noul",
            "instructions": "Could this call delete or overwrite user data?",
        }
    },
)
if result["is_risky"] and result["is_risky_confidence"] > 0.6:
    escalate_to_human(tool_call)
評価とベンチマークの他のプロジェクト