eval/results.md
1.6 KB · 54 lines · markdown Raw
1 # RL Agent evaluation
2
3 Metrics after calibration. Zero-shot = task families held out of training entirely.
4
5 ## In-task
6
7 | task family | questions | accuracy | ECE | NLL |
8 |---|---|---|---|---|
9 | conversation outcomes | 3600 | 0.482 | 0.019 | 0.693 |
10 | email triage | 2691 | 0.732 | 0.017 | 0.595 |
11 | emotion and tone | 1825 | 0.906 | 0.018 | 0.238 |
12 | inference and fact checking | 3022 | 0.883 | 0.054 | 0.340 |
13 | instruction-following tasks | 600 | 0.878 | 0.046 | 0.302 |
14 | intent and routing | 1475 | 0.991 | 0.009 | 0.181 |
15 | moderation and safety | 2708 | 0.967 | 0.061 | 0.153 |
16 | reading comprehension | 770 | 0.847 | 0.083 | 0.409 |
17 | response quality scoring | 3146 | 0.581 | 0.023 | 1.009 |
18 | robustness checks | 744 | 0.851 | 0.108 | 1.058 |
19 | search relevance | 733 | 0.628 | 0.066 | 0.728 |
20 | sentiment and rating | 961 | 0.442 | 0.438 | 3.545 |
21 | topic classification | 749 | 0.939 | 0.029 | 0.196 |
22
23 Overall: accuracy 0.753, ECE 0.030, Brier 0.308, accuracy at 50% coverage 0.947
24
25 ## Zero-shot
26
27 | task family | questions | accuracy | ECE | NLL |
28 |---|---|---|---|---|
29 | emotion and tone | 600 | 0.583 | 0.318 | 1.976 |
30 | instruction-following tasks | 600 | 0.863 | 0.045 | 0.319 |
31 | moderation and safety | 600 | 0.797 | 0.171 | 1.415 |
32 | sentiment and rating | 600 | 0.362 | 0.291 | 1.798 |
33
34 Overall: accuracy 0.651, ECE 0.204, Brier 0.532, accuracy at 50% coverage 0.818
35
36 ## Latency
37
38 ```
39 {
40 "1_questions": {
41 "p50_ms": 38.4,
42 "p95_ms": 42.1
43 },
44 "10_questions": {
45 "p50_ms": 156.0,
46 "p95_ms": 158.4
47 },
48 "50_questions": {
49 "p50_ms": 721.4,
50 "p95_ms": 733.0
51 }
52 }
53 ```
54