Jev 101
English ▾
Tools/Evaluation and calibration
Community

jev-benchmarks

AbdelStark/jev-benchmarks · Evaluation and calibration

Reproducible evaluation for calibration, selective risk, and latency.

CalibrationLatency
Read Chinese translation

用于校准、选择性风险与延迟的可复现评估。

Results belong to each project's dataset, prompts, model version, and measurement setup. Inclusion means the evidence is inspectable, not that benchmarks were independently rerun.

Directory and research sources

Provenance is retained from the previous edition. This edition did not install or re-audit the project.

cobanov/awesome-jev →