reports/data_audit.md
2.9 KB · 71 lines · markdown Raw
1 # Data audit — `jev-distill-corpus-v3`
2
3 tokenizer: `/root/models/Qwen3.5-9B` · max_seq_len: 1024 · template: bare-v1 (whole-string tokenization)
4
5 ## Splits
6
7 | split | rows | schema violations | mean tok | p50 | p95 | p99 | max | > max_seq_len | time |
8 |---|---|---|---|---|---|---|---|---|---|
9 | train | 655,806 | 0 | 129.0 | 87 | 351 | 531 | 856 | 0 | 83s |
10 | validation | 14,111 | 0 | 128.8 | 87 | 339 | 530 | 699 | 0 | 2s |
11 | calibration | 13,766 | 0 | 128.2 | 87 | 354 | 525 | 699 | 0 | 2s |
12 | test | 14,261 | 0 | 128.5 | 87 | 345 | 527 | 716 | 0 | 1s |
13 | test_set_30k | 29,955 | 0 | 106.6 | 85 | 214 | 509 | 702 | 0 | 3s |
14 | ood | 13,058 | 0 | 265.0 | 178 | 553 | 662 | 734 | 0 | 2s |
15
16 train tokens / epoch ≈ **84.6M**; 2 epochs ≈ **169M**
17
18 ## D1 — exactly-uniform teacher labels by split × source × kind
19
20 | split | source | kind | rows | uniform | share |
21 |---|---|---|---|---|---|
22 | train | openjev_v2 | choice | 20,434 | 511 | 2.5% |
23 | train | openjev_v2 | noul | 54,192 | 0 | 0.0% |
24 | train | yuri_v1 | noul | 137,203 | 137,203 | 100.0% |
25 | train | yuri_v3 | choice | 143,223 | 0 | 0.0% |
26 | train | yuri_v3 | noul | 150,515 | 1,514 | 1.0% |
27 | train | yuri_v3 | score | 150,239 | 0 | 0.0% |
28 | validation | openjev_v2 | choice | 463 | 15 | 3.2% |
29 | validation | openjev_v2 | noul | 1,167 | 0 | 0.0% |
30 | validation | yuri_v1 | noul | 2,920 | 2,920 | 100.0% |
31 | validation | yuri_v3 | choice | 3,129 | 0 | 0.0% |
32 | validation | yuri_v3 | noul | 3,231 | 35 | 1.1% |
33 | validation | yuri_v3 | score | 3,201 | 0 | 0.0% |
34 | calibration | openjev_v2 | choice | 463 | 9 | 1.9% |
35 | calibration | openjev_v2 | noul | 1,109 | 0 | 0.0% |
36 | calibration | yuri_v1 | noul | 2,812 | 2,812 | 100.0% |
37 | calibration | yuri_v3 | choice | 2,955 | 0 | 0.0% |
38 | calibration | yuri_v3 | noul | 3,254 | 30 | 0.9% |
39 | calibration | yuri_v3 | score | 3,173 | 0 | 0.0% |
40 | test | openjev_v2 | choice | 457 | 8 | 1.8% |
41 | test | openjev_v2 | noul | 1,139 | 0 | 0.0% |
42 | test | yuri_v1 | noul | 2,951 | 2,951 | 100.0% |
43 | test | yuri_v3 | choice | 3,163 | 0 | 0.0% |
44 | test | yuri_v3 | noul | 3,208 | 37 | 1.2% |
45 | test | yuri_v3 | score | 3,343 | 0 | 0.0% |
46 | test_set_30k | openjev_v2 | choice | 887 | 29 | 3.3% |
47 | test_set_30k | openjev_v2 | noul | 1,432 | 0 | 0.0% |
48 | test_set_30k | yuri_v1 | noul | 2,260 | 2,260 | 100.0% |
49 | test_set_30k | yuri_v3 | choice | 8,312 | 0 | 0.0% |
50 | test_set_30k | yuri_v3 | noul | 8,537 | 91 | 1.1% |
51 | test_set_30k | yuri_v3 | score | 8,527 | 0 | 0.0% |
52 | ood | openjev_v2 | choice | 3,219 | 0 | 0.0% |
53 | ood | openjev_v2 | noul | 9,767 | 0 | 0.0% |
54 | ood | openjev_v2 | score | 72 | 0 | 0.0% |
55
56 ## Train mix
57
58 | kind | rows | share |
59 |---|---|---|
60 | noul | 341,910 | 52.1% |
61 | choice | 163,657 | 25.0% |
62 | score | 150,239 | 22.9% |
63
64 | source | rows | share |
65 |---|---|---|
66 | yuri_v3 | 443,977 | 67.7% |
67 | yuri_v1 | 137,203 | 20.9% |
68 | openjev_v2 | 74,626 | 11.4% |
69
70 choice n_options histogram (train): 2:910, 3:8,826, 4:142,002, 5:3,586, 6:511, 7:1,593, 8:8, 9:3,121, 16:3,100
71