Independent Jev versus GPT-5.6 Terra comparison on three labeled classification tasks, reporting accuracy, calibration, latency, and cost.
Read Chinese translation
在三个带标签的分类任务上,独立对比 Jev 与 GPT-5.6 Terra,报告准确率、校准、延迟和成本。
Results belong to each project's dataset, prompts, model version, and measurement setup. Inclusion means the evidence is inspectable, not that benchmarks were independently rerun.
Directory and research sources
Provenance is retained from the previous edition. This edition did not install or re-audit the project.
cobanov/awesome-jev →Pinned source evidence: README.md →Pinned source evidence: runners.py →
