Independent synthetic-task study of Jev 1.13.0 framing sensitivity and failures, with raw responses and offline report checks.
Read Chinese translation
对 Jev 1.13.0 问题表述敏感性与失败模式的独立合成任务研究,附原始响应和离线报告检查。
Results belong to each project's dataset, prompts, model version, and measurement setup. Inclusion means the evidence is inspectable, not that benchmarks were independently rerun.
Directory and research sources
Provenance is retained from the previous edition. This edition did not install or re-audit the project.
cobanov/awesome-jev →Pinned source evidence: README.md →Pinned source evidence: behavior_study.py →
