Jev 101
English ▾
Tools/Evaluation and calibration
Community

jev-eval

4esv/jev-eval · Evaluation and calibration

Independent Jev versus GPT-5.6 Terra comparison on three labeled classification tasks, reporting accuracy, calibration, latency, and cost.

Classification evaluationComparison
Read Chinese translation

在三个带标签的分类任务上,独立对比 Jev 与 GPT-5.6 Terra,报告准确率、校准、延迟和成本。

Results belong to each project's dataset, prompts, model version, and measurement setup. Inclusion means the evidence is inspectable, not that benchmarks were independently rerun.

Directory and research sources

Provenance is retained from the previous edition. This edition did not install or re-audit the project.

cobanov/awesome-jev →Pinned source evidence: README.md →Pinned source evidence: runners.py →