reports/eval_27b_bundle.md
3.7 KB · 74 lines · markdown Raw
1 # Evaluation — export exports/jev-judge-qwen38-27b-v0.8
2
3 base: `/root/models/Qwen3.8-27B` · temperature: yes · max_seq_len 1024
4
5 ## Acceptance (DESIGN §8.3, main = test_set_30k)
6
7 **all rows**
8
9 | metric | target | value | B0 | pass |
10 |---|---|---|---|---|
11 | noul_auroc | >= 0.95 | 0.9961 | nan | ✅ |
12 | noul_brier_soft | <= 0.1 | 0.0013 | nan | ✅ |
13 | choice_top1 | >= 0.9 | 0.9041 | nan | ✅ |
14 | score_mae_expected | <= 0.35 | 0.0976 | nan | ✅ |
15 | kl | <= 0.15 | 0.0185 | nan | ✅ |
16 | ece | <= 0.03 | 0.0011 | nan | ✅ |
17
18 **excluding yuri_v1 exact-uniform placeholders (D1)**
19
20 | metric | target | value | B0 | pass |
21 |---|---|---|---|---|
22 | noul_auroc | >= 0.95 | 0.9961 | nan | ✅ |
23 | noul_brier_soft | <= 0.1 | 0.0016 | nan | ✅ |
24 | choice_top1 | >= 0.9 | 0.9041 | nan | ✅ |
25 | score_mae_expected | <= 0.35 | 0.0976 | nan | ✅ |
26 | kl | <= 0.15 | 0.0201 | nan | ✅ |
27 | ece | <= 0.03 | 0.0013 | nan | ✅ |
28
29 ## test_set_30k
30
31 all rows: `n=29955 · kl=0.0185 · js=0.0048 · top1=0.9220 · ece=0.0011 · mce=0.0043 · noul_auroc=0.9961 · noul_brier_soft=0.0013 · noul_brier_hard=0.0741 · score_mae_expected=0.0976 · score_rps=0.0077 · choice_top1=0.9041 · choice_kl=0.0364`
32
33 excluding yuri_v1 placeholders (2260 rows): `n=27695 · kl=0.0201 · js=0.0052 · top1=0.9220 · ece=0.0013 · mce=0.0043 · noul_auroc=0.9961 · noul_brier_soft=0.0016 · noul_brier_hard=0.0741 · score_mae_expected=0.0976 · score_rps=0.0077 · choice_top1=0.9041 · choice_kl=0.0364`
34
35 throughput: 14321 tok/s · 134.3 rows/s
36
37 ### test_set_30k by kind
38
39 | kind | n | kl | top1 | ece | noul_auroc | noul_brier_soft | score_mae_expected | choice_top1 |
40 |---|---|---|---|---|---|---|---|---|
41 | choice | 9199 | 0.036 | 0.904 | 0.003 | — | — | — | 0.904 |
42 | noul | 12229 | 0.003 | 0.966 | 0.002 | 0.996 | 0.001 | — | — |
43 | score | 8527 | 0.021 | 0.891 | 0.003 | — | — | 0.098 | — |
44
45 ### test_set_30k by source × kind
46
47 | source | kind | n | kl | top1 | ece | noul_auroc | noul_brier_soft | score_mae_expected | choice_top1 |
48 |---|---|---|---|---|---|---|---|---|---|
49 | openjev_v2 | choice | 887 | 0.146 | 0.885 | 0.020 | — | — | — | 0.885 |
50 | openjev_v2 | noul | 1432 | 0.003 | 0.999 | 0.002 | 1.000 | 0.001 | — | — |
51 | yuri_v1 | noul | 2260 | 0.000 | — | 0.004 | — | 0.000 | — | — |
52 | yuri_v3 | choice | 8312 | 0.025 | 0.906 | 0.003 | — | — | — | 0.906 |
53 | yuri_v3 | noul | 8537 | 0.004 | 0.960 | 0.001 | 0.995 | 0.002 | — | — |
54 | yuri_v3 | score | 8527 | 0.021 | 0.891 | 0.003 | — | — | 0.098 | — |
55
56 ### test_set_30k by family
57
58 | family | n | kl | top1 | ece | noul_auroc | noul_brier_soft | score_mae_expected | choice_top1 |
59 |---|---|---|---|---|---|---|---|---|
60 | agent | 2583 | 0.016 | 0.920 | 0.002 | 0.995 | 0.002 | 0.099 | 0.894 |
61 | biology | 1654 | 0.016 | 0.916 | 0.004 | 0.998 | 0.002 | 0.123 | 0.917 |
62 | business | 2352 | 0.016 | 0.939 | 0.003 | 0.997 | 0.001 | 0.071 | 0.922 |
63 | chemistry | 1333 | 0.013 | 0.945 | 0.004 | 0.998 | 0.002 | 0.088 | 0.948 |
64 | genomics | 1914 | 0.018 | 0.917 | 0.005 | 0.993 | 0.002 | 0.098 | 0.915 |
65 | knowledge | 4807 | 0.010 | 0.912 | 0.002 | 0.990 | 0.001 | 0.110 | 0.907 |
66 | medical | 2120 | 0.016 | 0.921 | 0.006 | 0.995 | 0.002 | 0.078 | 0.887 |
67 | openjev | 2319 | 0.058 | 0.956 | 0.008 | 1.000 | 0.001 | — | 0.885 |
68 | physics | 1656 | 0.016 | 0.909 | 0.003 | 0.995 | 0.001 | 0.097 | 0.908 |
69 | science | 1667 | 0.017 | 0.907 | 0.005 | 0.992 | 0.001 | 0.115 | 0.915 |
70 | spatial | 1647 | 0.017 | 0.913 | 0.005 | 0.991 | 0.002 | 0.094 | 0.880 |
71 | structured | 1668 | 0.018 | 0.911 | 0.004 | 0.989 | 0.003 | 0.123 | 0.886 |
72 | technical | 2581 | 0.017 | 0.923 | 0.004 | 0.995 | 0.002 | 0.074 | 0.893 |
73 | theology | 1654 | 0.016 | 0.910 | 0.004 | 0.996 | 0.002 | 0.120 | 0.928 |
74