README.md
52.5 KB · 785 lines · markdown Raw
1 ---
2 license: apache-2.0
3 base_model: Qwen/Qwen3.5-9B
4 base_model_relation: finetune
5 datasets:
6 - SargeDev/jev-distill-corpus-v3
7 language:
8 - en
9 library_name: transformers
10 pipeline_tag: text-classification
11 tags:
12 - system-one
13 - system-two
14 - blocks-of-experts
15 - typed-decisions
16 - decision-model
17 - calibrated-probabilities
18 - knowledge-distillation
19 - jev
20 - noul
21 - choice
22 - score
23 - lora
24 - qwen3_5
25 - text-generation
26 - dual-head
27 - vllm
28 - multimodal
29 - vision
30 - computer-use
31 - robotics
32 metrics:
33 - kl
34 - auroc
35 - brier
36 - ece
37 model-index:
38 - name: autotrust/JEV-9B (student of TypeSafe Jev 1.13)
39 results:
40 - task:
41 type: text-classification
42 name: typed decisions (noul / choice / score) — agreement with the TypeSafe Jev 1.13 teacher
43 dataset:
44 type: SargeDev/jev-distill-corpus-v3
45 name: jev-distill-corpus-v3 · test_set_30k
46 split: test_set_30k
47 metrics:
48 - type: kl_divergence
49 name: mean KL(target ‖ model), all test rows (25,376 of 29,955 targets are TypeSafe Jev 1.13 distributions)
50 value: 0.0210
51 - type: auroc
52 name: noul AUROC
53 value: 0.996
54 - type: brier
55 name: noul Brier (vs. target probability, all rows)
56 value: 0.0015
57 - type: mae
58 name: score expected-value MAE (0–5 scale)
59 value: 0.103
60 - type: ece
61 name: ECE (15 bins, after temperature)
62 value: 0.0007
63 - type: accuracy
64 name: choice top-1 agreement (all rows)
65 value: 0.898
66 - type: accuracy
67 name: choice top-1 agreement (decisive-target rows, top-2 gap ≥ 0.1)
68 value: 0.954
69 - task:
70 type: text-generation
71 name: code generation — System 2 path (base lm_head, adapter off)
72 dataset:
73 type: openai/openai_humaneval
74 name: HumanEval
75 split: test
76 metrics:
77 - type: pass@1
78 name: pass@1 (greedy, completion-style prompt)
79 value: 0.707
80 ---
81
82 # autotrust/JEV-9B
83
84 ### AutoTrust's first integrated System 1 + System 2 open model, built with the Blocks of Experts recipe
85
86 **Fast, calibrated System 1 decisions that are indistinguishable from the closed TypeSafe Jev 1.13 by KL, and
87 deliberate System 2 generation and reasoning from an untouched Qwen3.5-9B — one set of weights, one vLLM engine,
88 routed per request. The fastest model of the family: it answers a single decision in about a third of the time the
89 hosted API takes. Its successor, [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B), is closer to Jev,
90 transfers better to unseen tasks and has a stronger System 2.**
91
92 ## New (3 October 2026): JEV-9B can see — robot arm and computer use
93
94 JEV-9B now takes images. Every step below is **one System 1 decision**: camera image or screenshot in, a probability
95 for every action out, in a single forward pass (about 0.2 s on one GPU). Run it with `bash vl/serve.sh` (see
96 [Images: quick start](#images-quick-start)).
97
98 **Robot arm: pick and place from camera images.** The arm sees a top camera image; at every step System 1 answers two
99 questions (is the target left or right of the gripper, above or below it), and the arm moves accordingly, halving its
100 step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation).
101
102 <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/robot_arm_pick_place.mp4" controls autoplay loop muted playsinline width="100%"></video>
103
104 On 20 random scenes it completed the task 10 times; every cube it grasped ended in the tray, and every miss was a grasp
105 3–5 cm off target. About 165 ms per decision. Asking it to choose one of 8 motor commands directly did not work: this
106 model is a fast visual judge, not an end-to-end controller.
107
108 **Computer use: screenshot → which element to click.** A real browser (headless Chromium). Every clickable element gets
109 a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
110
111 <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_shop.mp4" controls autoplay loop muted playsinline width="100%"></video>
112
113 <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_settings.mp4" controls autoplay loop muted playsinline width="100%"></video>
114
115 <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_mail.mp4" controls autoplay loop muted playsinline width="100%"></video>
116
117 **95% of 60 random multi-step tasks completed** (shop, settings, mail; 3–7 clicks each), about 0.2 s per click. The
118 colour swatches and switches carry no text, so those clicks are decided from the screenshot alone. With the numbered
119 boxes only (no element text) it completed 37%. The failures skipped a step (the colour) and then checked out an empty cart.
120
121 Code for both demos: [`vl/demos/`](vl/demos). Image judging, briefly: VL-RewardBench 74.3%, AgentRewardBench AUROC 0.91,
122 zero-shot short-video recommendation from covers AUC 0.72 (details in [`reports/vl/`](reports/vl)).
123
124 ### Images: quick start
125
126 ```bash
127 hf download autotrust/JEV-9B --include "vl/*" --local-dir JEV-9B
128 bash JEV-9B/vl/serve.sh # downloads Qwen/Qwen3.5-9B (with its vision encoder) and serves both systems on :8000
129 ```
130
131 ```python
132 import base64, requests
133
134 def image(path):
135 return {"image": "data:image/png;base64," + base64.b64encode(open(path, "rb").read()).decode()}
136
137 r = requests.post("http://localhost:8000/v1/decide", json={
138 "kind": "choice",
139 "state": ["Top camera image:", image("scene.png"), "\nTask: put the red cube in the tray."],
140 "question": "Is the red cube to the left or to the right of the gripper?",
141 "options": ["left", "right"]}).json()
142 print(dict(zip(r["options"], r["probabilities"])))
143 ```
144
145 How it works: JEV-9B's language weights are bit-identical to Qwen3.5-9B's, so `vl/serve.sh` serves the unmodified
146 multimodal Qwen3.5-9B with JEV-9B's System 1 adapter (`vl/adapter_vllm`, the same weights with the layer names moved).
147 Text decisions match the text-only model (300 test decisions: largest probability difference 0.011). System 2 also reads
148 images. Keep `--max-num-seqs 8` (set in `serve.sh`); decisions over images are zero-shot.
149
150 ## At a glance
151
152 **Integrated System 1 + System 2, first generation.** JEV-9B is AutoTrust's first open model to serve both modes of
153 thinking from a single set of weights; the second generation is
154 [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B). *System 1* answers typed questions (`noul` yes/no ·
155 `choice` over 2–16 options · `score` on a 0–5 scale) in one forward pass and returns a calibrated probability
156 distribution. *System 2* is ordinary text generation with step-by-step reasoning (thinking mode). Both run on the
157 same backbone in the same engine, and a request chooses its system.
158
159 **Blocks of Experts recipe.** Rather than fine-tuning one monolithic model, the Blocks of Experts (BoE) recipe keeps a
160 strong pretrained model as a frozen expert block and adds a small, detachable expert block trained for one capability.
161 In JEV-9B the System 2 block is Qwen3.5-9B, bit-identical to the release; the System 1 block is 40.2 M trained
162 parameters (0.5 % of the backbone), trained in ≈ 3 hours on one B200. Because the blocks stay separate, adding
163 System 1 costs System 2 nothing: HumanEval is 70.7 % before and after, with all 164 completions byte-identical. Folding
164 the same block into the backbone instead would have cost 9 points (61.6 %).
165
166 **Indistinguishable from the closed original on System 1, by KL.** On the 25,376 held-out questions (53 domains)
167 whose labels are TypeSafe Jev 1.13's own output distributions, the mean KL divergence is **≈ 0.019** (0 = identical).
168 An observer who sees sampled decisions gains on average 0.019 nats of evidence per decision about which model produced
169 it, so it takes about 54 decisions to gather a single nat. The fidelity extends to the teacher's mistakes (see
170 [System 1: indistinguishable from TypeSafe Jev 1.13](#system-1-indistinguishable-from-typesafe-jev-113-by-kl)).
171 Among the open Jev reproductions we could find, only the JEV models publish this distribution-level measure
172 (see [How JEV-9B compares with other open Jev reproductions](#how-jev-9b-compares-with-other-open-jev-reproductions)).
173
174 **Faster than the hosted API.** On one B200, a single decision takes a median ≈ 90 ms, against 238–301 ms measured
175 independently for the hosted TypeSafe Jev 1.13 API, and one GPU sustains about 15× the decisions per second an
176 independent benchmark achieved against that API (see [Speed](#speed-vs-the-hosted-typesafe-jev-113)).
177
178 **The fast member of the family; JEV-27B is the closer one.** Same recipe, same API: JEV-9B is 2.6× faster than
179 JEV-27B on the same benchmark and its weights are a third of the size (18 GB vs 54 GB). JEV-27B lowers mean KL to Jev's
180 distributions from ≈ 0.019 to ≈ 0.017, more than halves KL on unseen task families (0.234 → 0.104), keeps 96 % instead
181 of 90 % of the teacher's accuracy on an independent 16-option benchmark, and scores 78.0 % instead of 70.7 % on
182 HumanEval (see [JEV-9B vs JEV-27B](#jev-9b-vs-jev-27b)).
183
184 > **Two models, two organisations.** **TypeSafe Jev 1.13** is the hosted, closed-source model made by TypeSafe AI; it
185 > is the *teacher* whose published output distributions this model was trained on. **autotrust/JEV-9B** (this
186 > repository) is an independent open-weights *student* built by AutoTrust AI from the Apache-2.0 corpus
187 > [`SargeDev/jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3). It is not
188 > affiliated with, endorsed by, or a product of TypeSafe AI, and shares no weights or code with it.
189
190 ## Headline results
191
192 System 1 numbers are on the held-out `test_set_30k` of `jev-distill-corpus-v3`. Its 29,955 rows come from three
193 sources: 25,376 rows labelled with TypeSafe Jev 1.13's own output distributions (`yuri_v3`), 2,319 Open-Jev rows with
194 programmatic ground-truth labels (`openjev_v2`), and 2,260 placeholder rows (`yuri_v1`). Rows marked *Jev-labelled* use
195 only the first group.
196
197 | | What is measured | autotrust/JEV-9B | How to read it |
198 |---|---|---|---|
199 | **System 1** | Mean KL divergence from TypeSafe Jev 1.13's distributions, Jev-labelled rows, 0 = identical | **≈ 0.019** | Indistinguishable from the teacher's decisions at this resolution: ≈ 54 sampled decisions to gather one nat of evidence |
200 | | Mean KL to all test targets (Jev, programmatic and placeholder labels) | **0.021** | The figure in the model index above |
201 | | Yes/no AUROC (`noul`), Jev-labelled rows | **0.994** | Ranks true vs. false almost perfectly (0.996 over all rows) |
202 | | Choice top-1 agreement with Jev, Jev-labelled rows | **90.2 %** | 95.4 % over all rows where the target's top two options differ by ≥ 0.1; on near ties any faithful copy agrees about half the time |
203 | | Rating error (`score`, 0–5 scale), mean absolute error of the expected rating | **0.103** | About one tenth of a rating step |
204 | | Expected calibration error | **0.0007** | A stated 80 % is an 80 %; fitted temperatures ≈ 1.00, no post-hoc correction needed |
205 | | KL to the programmatic labels of task families never seen in training (Open-Jev OOD split) | **0.234** | Transfer to new tasks; these labels are ground truth, not Jev's outputs. JEV-27B: 0.104 |
206 | | Independent benchmark with human gold labels, 16 options | **90 % of the teacher** (0.694 vs 0.769) | 94–97 % of the teacher at 2, 4 and 8 options; see [Benchmark highlights](#benchmark-highlights) |
207 | **System 2** | HumanEval pass@1, greedy | **70.7 %** | Identical to Qwen3.5-9B (116/164); all 164 completions byte-identical to the base model |
208 | **Speed** | Single decision, median, one B200 | **≈ 90 ms** | Hosted TypeSafe Jev 1.13, measured independently: 238 ms mean, 291–301 ms median |
209 | | Decisions per second on the independent benchmark, one B200 | **≈ 340** | ≈ 15× the 23 per second measured against the hosted API; see [Speed](#speed-vs-the-hosted-typesafe-jev-113) |
210 | | Batched, 128 decisions per batch | **2.5 ms** per decision | With vLLM: 205 decisions/s over HTTP at 256 concurrent clients, text generation ≈ 50× faster than the PyTorch path |
211 | **Efficiency** | Trained parameters | **40.2 M** (0.5 % of 7.9 B) | ≈ 3 B200-hours, 0.93 epoch ≈ 608 k rows |
212
213 ## JEV-9B vs JEV-27B
214
215 JEV-9B was AutoTrust's first integrated System 1 + System 2 model.
216 [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B), the second generation, uses the same recipe, code,
217 hyper-parameters, API and two-block packaging; only the backbone and memory settings changed. Both are evaluated on the
218 same held-out test set and the same independent benchmark.
219
220 <p align="center">
221 <a href="https://huggingface.co/autotrust/JEV-27B/blob/main/27b-2.jpg"><img src="https://huggingface.co/autotrust/JEV-27B/resolve/main/27b-2.jpg" alt="JEV family benchmark highlights: KL to TypeSafe Jev 1.13 by question type, accuracy as a percentage of Jev on an independent benchmark, HumanEval for the System 2 path, and speed against the hosted API, for JEV-9B (light bars) and JEV-27B (dark bars)" width="100%"></a>
222 <br>
223 <sub><b>JEV family benchmark highlights</b> (chart from the JEV-27B repository; light bars = JEV-9B). A · KL to Jev by question type · B · accuracy as % of Jev on an independent benchmark · C · System 2 unchanged · D · speed vs the hosted API · click to enlarge</sub>
224 </p>
225
226 | | **JEV-9B** | JEV-27B | JEV-27B vs JEV-9B |
227 |---|---|---|---|
228 | Backbone | Qwen3.5-9B | Qwen3.8-27B | |
229 | **System 1** — mean KL to TypeSafe Jev 1.13, Jev-labelled rows | ≈ 0.019 | **≈ 0.017** | ≈ −11 % |
230 | Mean KL to all test targets | 0.021 | **0.019** | −11 % |
231 | KL to ground-truth labels, unseen task families (OOD) | 0.234 | **0.104** | −56 % |
232 | Top-1 accuracy, unseen task families (OOD) | 0.918 | **0.942** | +2.4 pts |
233 | Choice top-1 agreement with Jev, Jev-labelled rows | 90.2 % | **90.5 %** | +0.3 pts |
234 | Rating error (`score` MAE, all Jev-labelled) | 0.103 | **0.098** | −5 % |
235 | Top-1 flips under option shuffle (test set) | 3.9 % | **2.9 %** | −1.0 pt |
236 | Yes/no AUROC (`noul`), Jev-labelled rows | 0.994 | **0.995** | +0.001 |
237 | Calibration error (ECE) | **0.0007** | 0.0009 | JEV-9B slightly lower; both below 0.001 |
238 | Independent benchmark, 16 options — % of teacher accuracy | 90 % | **96 %** | +6 pts |
239 | Independent benchmark — answers changed by option order alone (teacher: 7.0 %) | 11.5 % | **7.4 %** | JEV-27B is close to the teacher's 7.0 % |
240 | **System 2** — HumanEval pass@1 (greedy) | 70.7 % | **78.0 %** | +7.3 pts |
241 | Latency on one B200 — single request / batched | **≈ 90 ms / 2.5 ms** | 137 ms / 4.2 ms | JEV-9B is faster |
242 | Benchmark throughput — 14,400 decisions on one B200 | **42 s** | 110 s | JEV-9B is 2.6× faster |
243 | Download size (backbone + adapter) | **18 GB** | 54 GB | |
244 | Trained parameters / compute | 40.2 M / ≈ 3 B200-hours | 108.9 M / ≈ 9.2 B200-hours | |
245
246 On the fresh Hacker News, V2EX and community examples (illustrations, not a benchmark), JEV-9B got 92 of 96 decisions
247 right against 95 of 96 for JEV-27B. The difference is on the harder tasks: JEV-9B misses a TypeScript port that
248 breaks a "branded, range-checked integer" rule (0.33; JEV-27B 0.93) and flags a CEO wire-transfer fraud message with
249 less confidence (0.56; JEV-27B 0.84).
250
251 **Which to pick.** For routing, moderation, topic triage and short option lists, JEV-9B gives nearly the same answers
252 2.6× faster (14,400 benchmark decisions in 42 s vs 110 s on one B200) with a third of the weight memory. For long option lists
253 (more than about 8), unfamiliar task families, code-rule checks, fraud screening, or when the System 2 path matters, use
254 [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B).
255
256 ## How JEV-9B compares with other open Jev reproductions
257
258 Dozens of open reproductions of TypeSafe Jev appeared within weeks of its launch; the community
259 [Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) (formerly the Jev Reproductions
260 Tracker) evaluates 55 of them. Most are trained on human or programmatic gold labels, or on their own synthetic data, so
261 they aim to match or beat Jev's accuracy rather than reproduce its probabilities. "Closest to Jev" therefore depends on
262 how closeness is measured:
263
264 | measure of closeness to TypeSafe Jev 1.13 | published results (snapshot of 25 September 2026) | where JEV-9B stands |
265 |---|---|---|
266 | **Distribution level:** KL to Jev's own output distributions on held-out rows | JEV-27B ≈ 0.017 and JEV-9B ≈ 0.019 on 25,376 Jev-labelled rows. We found no other open reproduction that publishes this measure. | Second lowest published, after JEV-27B |
267 | **Accuracy relative to Jev** on [`decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure), 16 options, human gold labels | JEV-27B 96 % · JEV-9B 90 % · Laya 90 % · DeBERTa-v3-large zero-shot 90 % · DeBERTa-v3-base zero-shot 83 % · GLiClass-large 81 % · bge-large 73 % · gte-large 69 % | Level with the best of the other models measured there; JEV-27B is closer (JEV rows are AutoTrust re-runs of the same items; the others were run by the benchmark's author) |
268 | **Score parity on community leaderboards** | [Decision Index 0.2](https://huggingface.co/spaces/multimodalart/jev-decision-index): Jev 51.67, AutoJev-27B 50.94. [JevBench v1.4.2](https://github.com/fstandhartinger/jevbench): decider-4b v2 64.13, Jev 63.29, JevK5 62.04. [Open-Jev](https://zefan-cai.github.io/open-jev/benchmarks/) public JevBench subset: Jev 200/231, Open-Jev 27B v1.1 197/231 | Not yet evaluated |
269
270 On the evidence published today, the two JEV models are the closest open models to TypeSafe Jev 1.13 at the level of
271 output distributions, with JEV-9B second to JEV-27B. On the independent benchmark JEV-9B is level with the best of the
272 other models measured there, not ahead of them. It has not yet been run on the Decision Index or JevBench, where
273 AutoJev-27B scores within about one point of Jev and decider-4b v2 edges ahead of it, so we do not claim it is the
274 closest by every measure. Note that some reproductions report beating Jev on their own test sets (AutoJev-27B reports
275 84.60 % against Jev's 82.79 %); that is a different goal from reproducing Jev's behaviour.
276
277 *Not to be confused with AutoJev-27B (`denis-pplx/autojev-27b`), an unrelated Qwen3.8-27B decision model trained with
278 full-weight SFT on its own data.*
279
280 ## Speed vs the hosted TypeSafe Jev 1.13
281
282 TypeSafe does not publish Jev's size or hardware; it reports 70–500 ms end to end. Independent measurements, and ours:
283
284 | | TypeSafe Jev 1.13, hosted API | **JEV-9B, one B200** | JEV-27B, one B200 |
285 |---|---|---|---|
286 | One decision, single request | 238 ms mean over 29,600 calls ([`decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure)); 291–301 ms median on three workloads ([Open-Jev](https://zefan-cai.github.io/open-jev/benchmarks/)) | **≈ 90 ms** median (87 ms) | 137 ms median |
287 | Decisions per second on `decision-models-under-pressure` | 23, with 5 client workers and one question per call | **≈ 340** (14,400 in 42 s) | ≈ 130 (14,400 in 110 s) |
288 | Batched, 128 decisions per batch | — | **2.5 ms** per decision | 4.2 ms per decision |
289
290 So JEV-9B answers a single decision in roughly a third of the time (JEV-27B in roughly half), and one GPU sustains
291 about 15× (JEV-27B: about 6×) the throughput the benchmark's author achieved against the hosted API. Read these with
292 the caveats: our latencies are measured on the serving host with no network hop, while the hosted numbers include
293 internet, TLS and queueing; hosted throughput depends on client concurrency and the API's rate limits; Jev's latency is
294 roughly flat in the number of questions per request, so bundling questions narrows the throughput gap; and our figures
295 are self-reported while Jev's come from third parties. The two throughput runs use the same benchmark but not an
296 identical call set (ours stops at 16 options).
297
298 ## System 1: indistinguishable from TypeSafe Jev 1.13, by KL
299
300 **What the number means.** KL(Jev ‖ model) is the expected log-likelihood ratio, per sampled decision, between
301 TypeSafe Jev 1.13 and the student when the decision comes from Jev. On the 25,376 held-out rows whose targets are
302 Jev's own output distributions, the mean is ≈ 0.019 nats (computed from the per-slice values below, which are
303 published to three decimals): one decision carries almost no evidence about which of the two models produced it, and
304 an observer needs about 1 / KL ≈ 54 independent decisions to accumulate one nat (a likelihood ratio of about e ≈ 2.7 : 1).
305
306 For scale, Jev is not deterministic itself: an independent study found it changes its answer on 4.3 % of repeated,
307 identical 64-option calls, and it returns probabilities rounded to two decimals, which is the resolution of the
308 targets used here.
309
310 | Jev-labelled slice (`yuri_v3`, `test_set_30k`) | n | KL | ≈ decisions to gather one nat (1 / KL) |
311 |---|---|---|---|
312 | `noul` | 8,537 | 0.005 | ≈ 200 |
313 | `choice` | 8,312 | 0.028 | ≈ 36 |
314 | `score` | 8,527 | 0.023 | ≈ 43 |
315 | **all Jev-labelled rows** | **25,376** | **≈ 0.019** | **≈ 54** |
316
317 JEV-27B reaches ≈ 0.017 (≈ 60 decisions per nat) on the same rows.
318
319 The other test rows are not labelled by Jev and are not part of this claim: Open-Jev rows carry programmatic ground
320 truth (in-distribution KL 0.004 for `noul`, 0.176 for `choice`; 0.234 on the OOD split of unseen task families), and
321 the `yuri_v1` rows carry placeholder labels. No Jev-labelled out-of-distribution set exists in the corpus, so the claim
322 is established on the 53 training domains; outside them, the independent benchmark with human labels (90–97 % of Jev's
323 accuracy) is the best available evidence.
324
325 **Fidelity includes the teacher's mistakes.** On a poker spot where a solver always checks, TypeSafe Jev 1.13 shoves
326 with 0.62 in a published test; JEV-9B shoves too, with 0.70 (JEV-27B 0.63). A faithful copy of System 1 is also a
327 faithful copy of its blind spots. At 9 B the student also adds some of its own: on an independent benchmark 11.5 % of
328 its 16-option answers change when only the option order changes, against 7.0 % for the teacher (JEV-27B 7.4 %).
329
330 ## The Blocks of Experts recipe
331
332 ```
333 ┌── System 2 block: lm_head (248,320 × 4096) ───────► text generation and reasoning
334 Request ─► Router ─► Qwen3.5-9B backbone (frozen, bit-identical to the base) (adapter off; HumanEval 70.7 % = base)
335 per │
336 request └── + System 1 block: LoRA (40.1 M) + 24-slot head (98 k) ─► calibrated typed decision
337 (adapter on, decision path only) (one prefill pass; KL ≈ 0.019 to Jev)
338 ```
339
340 | block | what it is | parameters | trained? | used for |
341 |---|---|---|---|---|
342 | Backbone | `Qwen/Qwen3.5-9B` text tower (vision tower and MTP head dropped), bf16 | 7.9 B | no — bit-identical to the base | both systems |
343 | **System 2 block** | the original `lm_head` (248,320 × 4096) | part of the base | no | text generation and step-by-step reasoning |
344 | **System 1 block** | LoRA r=16 on the decoder projections + a 24-slot fp32 decision head initialised from `lm_head` rows | 40.1 M + 98 k | yes, ≈ 3 B200-hours | calibrated typed decisions |
345 | Router | per request: the vLLM LoRA module `jev-decision`, or `peft` adapter on/off | — | — | chooses the system |
346
347 **Why separate blocks rather than one merged fine-tune.** Folding the System 1 LoRA into the backbone would let a
348 single weight set serve both heads, but it costs generation quality: the merged backbone with the original `lm_head`
349 scores 61.6 % (101/164) on HumanEval against 70.7 % for the base, a 9-point drop, even though prose perplexity barely
350 moves (3.15 → 3.30). Keeping the backbone pristine and applying the System 1 block only on the decision path removes
351 that trade-off. For decision serving the adapter is merged *in memory* at start-up, so decision latency matches a
352 merged bundle.
353
354 **Why the recipe is this efficient.**
355
356 1. **Pretraining does most of the work; distillation sharpens.** The decision head is initialised from the backbone's
357 own `lm_head` rows for the verbalizer tokens (`false/true`, `0`–`5`, `A`–`P`), so at step 0 its output equals the
358 pretrained model's zero-shot restricted next-token distribution (verified to |Δp| < 1e-5; measured 8.6e-07). Before
359 seeing a single label it already agrees with the test targets on 53 % of `choice` questions with `noul` AUROC 0.82;
360 distillation takes it to 90 % / 0.996.
361 2. **Small trainable footprint.** 40.2 M parameters — 0.5 % of the backbone. Validation KL was already below 0.10 after
362 the first 64 k rows, test KL reached 0.028 after 0.49 epoch (≈ 1.7 B200-hours) and 0.021 after 0.93 epoch.
363 3. **Transfer to unseen tasks.** The pretrained backbone reads the *content* of a new task instead of matching surface
364 patterns of the training domains: KL 0.234 and top-1 0.918 against the programmatic labels of the OOD split. It
365 also reads real, long, structured states (prose, JSON game states, policy documents; up to 856 tokens in the corpus).
366 4. **Reads options, not positions.** With 30 % option-permutation augmentation, the top-1 flip rate under shuffled
367 `choice` options is 3.9 %; the same backbone before distillation flips 38 % of the time.
368 5. **Calibration falls out of the objective.** Distilling full teacher distributions with KL (plus an ordinal RPS term
369 for `score`) gives fitted temperatures of 1.002 / 0.984 / 1.012 and ECE 0.0007 with no post-hoc correction.
370 6. **It scales without code changes, and scale pays off.** The same code, hyper-parameters and packaging produced the
371 second-generation JEV-27B; only `model_path` and memory settings changed (the head-initialisation identity holds
372 there too, 4.5e-07). Going from 9 B to 27 B lowers KL to Jev from ≈ 0.019 to ≈ 0.017, halves OOD KL
373 (0.234 → 0.104), and raises the System 2 path from 70.7 % to 78.0 % on HumanEval.
374
375 ## Benchmark highlights
376
377 ### Independent benchmark: side by side with TypeSafe Jev 1.13
378
379 [`gazelle93/decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure) (published
380 25 Sep 2026) asks decision models to pick the right label for real texts from CLINC-150, MTOP, GoEmotions, DBpedia and
381 financial tweets under three kinds of pressure: more options, near-miss options, and shuffled option order. The labels
382 are human gold labels, none of this data is in our training set, and TypeSafe Jev 1.13's results are published with it.
383 We re-ran the same items with autotrust/JEV-9B and autotrust/JEV-27B, up to our 16-option limit.
384
385 | | TypeSafe Jev 1.13 (published) | **autotrust/JEV-9B** | autotrust/JEV-27B |
386 |---|---|---|---|
387 | Accuracy with 2 / 4 / 8 / 16 options (800 items, 4 domains) | 0.890 / 0.801 / 0.782 / 0.769 | **0.868 / 0.774 / 0.735 / 0.694** | 0.876 / 0.784 / 0.767 / 0.740 |
388 | 16 options — CLINC / DBpedia / GoEmotions / MTOP | 0.945 / 0.900 / 0.470 / 0.760 | **0.875 / 0.855 / 0.325 / 0.720** | 0.930 / 0.885 / 0.415 / 0.730 |
389 | 16 options, near-miss vs. unrelated wrong options (CLINC + MTOP, 400 items) | 0.912 vs 0.985 | **0.875 vs 0.975** | 0.907 vs 0.983 |
390 | Answers changed by shuffling the options alone (16 options, 5 orderings) | 7.0 % | **11.5 %** | 7.4 % |
391 | Time for 14,400 decisions on one B200 | — | **42 s** | 110 s |
392
393 On data it was never trained on, JEV-9B reaches 97 % of the teacher's accuracy with 2 and 4 options, 94 % with 8 and
394 90 % with 16: it falls behind faster than JEV-27B (96–98 %) as the option list grows, loses a little more on near-miss
395 options, and is more sensitive to option order than the teacher. Our run follows the benchmark's published method (gold
396 plus the first K−1 distractors of a pool, shuffled per item); the orderings are seeded differently, so compare
397 aggregates, not individual items.
398
399 ### Fresh examples (Hacker News and V2EX, 23–25 September 2026)
400
401 Expected answers were written by hand before the model was run. These are illustrations (≈ 110 decisions), not a
402 benchmark.
403
404 | task | autotrust/JEV-9B | autotrust/JEV-27B |
405 |---|---|---|
406 | Topic of 19 HN front-page stories (10 options) + "is it about AI?" | **38 / 38** | 38 / 38 |
407 | 12 comments from a heated HN thread: "insults or attacks someone?" + "what is it mainly doing?" (6 options) | **22 / 24** | 23 / 24 |
408 | 10 V2EX hot posts in **Chinese**: "contains a referral / invite code?" + "promotes a product or paid offer?" | **18 / 19** | 19 / 19 |
409 | Community use cases: code-rule checks in the style of `adhere`, injection filtering, ticket routing, phishing, code-review diffs, urgency scores | **14 / 15** | 15 / 15 |
410
411 | input | question | autotrust/JEV-9B |
412 |---|---|---|
413 | HN: "Two-tier encryption in the UK" | topic (10 options) | security & privacy · 0.87 |
414 | HN: "Using LLMs to trace alchemical knowledge and decode 17th century letters" | about AI? | P(true) = 0.88 |
415 | HN comment: "Please stop this. We've asked you before to observe the guidelines…" | what is it mainly doing? | moderating the discussion · 0.78 |
416 | V2EX: "一个不需要 gemini pro 的完全免费的注册 Muse 的方法 … 邀请码:…" | contains a referral / invite code? | P(true) = 1.00 |
417 | V2EX: "今天中秋节,还要加班的有吗?来报道下" | promotes a product or paid offer? | P(true) = 0.01 |
418 | Diff replacing a parameterised query with `"… WHERE id = " + request.args["id"]` | introduces a security vulnerability? | P(true) = 0.93 (0.17 for a variable rename) |
419 | "I'm not happy with the fit. What are my options here?" | asking for a refund? | P(true) = 0.17 (TypeSafe's docs report 0.22 for Jev on this exact text) |
420
421 Where it failed or wavered:
422
423 * **Code-rule check**: missed a TypeScript file that declares `const port: number = Number(process.env.PORT)` against
424 the rule "a port must be a branded, range-checked integer" (0.33); JEV-27B flags it (0.93).
425 * **Fraud screening**: a CEO wire-transfer (business-email-compromise) message is flagged, but only at 0.56 (JEV-27B 0.84).
426 * **Comment intent**: "Because it's not a real argument. It's a deflection people use." read as attacking another
427 commenter (0.66) rather than arguing a point.
428 * **Chinese promotion**: a V2EX post launching a paid HTTPS debugging tool was not flagged as promotional (0.44;
429 JEV-27B 0.77).
430 * **A poker spot with the nuts** (check or shove four times the pot; a solver checks 100 %): shoves with 0.70; the
431 teacher shoved with 0.62, so this mistake comes from the teacher.
432 * Counting ("more than 3 fruits?" / "more than 5?" for a list of 4: 0.84 / 0.37), date comparisons and an instruction
433 injected inside the state were handled correctly, but on a handful of examples only.
434
435 Per-example outputs and the benchmark aggregates are in `reports/realworld_9b.json` (the HN and V2EX inputs came from
436 their public APIs on 25 September 2026).
437
438 ## Quickstart with vLLM (recommended)
439
440 **One vLLM engine serves both systems from the same pristine weights.** Ordinary requests go through the base
441 `lm_head` (System 2, exactly Qwen3.5-9B); requests addressed to the LoRA module `jev-decision` go through the decision
442 head (System 1). `adapter_vllm/` contains the backbone LoRA plus the 24-slot decision head re-expressed as an `lm_head`
443 LoRA (only the 24 verbalizer rows change), so a typed decision is a single prefill step with `max_tokens=1`,
444 constrained to the option tokens and read back as log-probabilities.
445
446 ### 1 — Start the server (OpenAI-compatible)
447
448 ```bash
449 hf download autotrust/JEV-9B --local-dir JEV-9B # ~18 GB
450 vllm serve JEV-9B --served-model-name autotrust/JEV-9B \
451 --enable-lora --max-lora-rank 32 --lora-modules jev-decision=JEV-9B/adapter_vllm \
452 --logprobs-mode processed_logprobs --max-model-len 4096
453 ```
454
455 `--logprobs-mode processed_logprobs` is required: it makes the returned log-probabilities respect `allowed_token_ids`.
456 `--max-model-len 4096` is sized for decisions; raise it (for example to 16384) if System 2 requests will think at
457 length. Add `--enable-prefix-caching --mamba-cache-mode align` if you ask many questions about the same state.
458
459 ### 2 — System 2: generation and reasoning (the unmodified base model)
460
461 ```bash
462 curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{
463 "model": "autotrust/JEV-9B",
464 "messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
465 "max_tokens": 60, "chat_template_kwargs": {"enable_thinking": false}}'
466 ```
467
468 Set `"enable_thinking": true` for deliberate, step-by-step reasoning. This path is Qwen3.5-9B unchanged; see the
469 [Qwen3.5-9B model card](https://huggingface.co/Qwen/Qwen3.5-9B) for its reasoning benchmarks and recommended sampling
470 settings.
471
472 ### 3 — System 1: typed decisions (Python, only `requests` + two small JSON files)
473
474 ```python
475 import json, math, requests
476 from huggingface_hub import hf_hub_download
477
478 REPO, URL = "autotrust/JEV-9B", "http://localhost:8000"
479 dh = json.load(open(hf_hub_download(REPO, "adapter_vllm/decision_head.json"))) # bias + verbalizer token ids
480 T = json.load(open(hf_hub_download(REPO, "calibration.json")))["per_kind"] # per-kind temperatures
481
482 def decide(kind, state, question, options=None):
483 options = {"noul": ["false", "true"], "score": [str(i) for i in range(6)]}.get(kind, options)
484 lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
485 prompt = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
486 s = dh["slots"]["ranges"][kind][0]
487 ids = dh["verbalizer_ids"][s : s + len(options)] # the option tokens of this kind
488 r = requests.post(f"{URL}/v1/completions", json={
489 "model": "jev-decision", "prompt": prompt, "max_tokens": 1, "temperature": 1.0,
490 "logprobs": len(options), "allowed_token_ids": ids,
491 "add_special_tokens": False, "return_tokens_as_token_ids": True}).json()
492 lp = {int(k.split(":")[1]): v for k, v in r["choices"][0]["logprobs"]["top_logprobs"][0].items()}
493 z = [(lp.get(t, -1e9) + dh["bias"][s + i]) / T[kind] for i, t in enumerate(ids)] # + head bias, / temperature
494 e = [math.exp(x - max(z)) for x in z]
495 return {o: x / sum(e) for o, x in zip(options, e)}
496
497 print(decide("choice", "SKU AX-330 stock at 8% of safety level; supplier late twice this quarter.",
498 "Supplier response for this scenario.", ["issue_warning", "renegotiate", "dual_source", "maintain"]))
499 # ≈ {'issue_warning': 0.33, 'renegotiate': 0.14, 'dual_source': 0.53, 'maintain': 0.001}
500 print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
501 "Is the customer asking for a refund?"))
502 ```
503
504 Values can differ in the third decimal between runs: vLLM computes in bf16 and results depend slightly on which
505 requests are batched together. The head bias and the temperature are applied client-side; the log-softmax normaliser
506 that vLLM applies cancels out, so the result is exactly the decision head's calibrated distribution.
507
508 ### 4 — System 1 → System 2: confidence-gated escalation
509
510 Because both systems live in one engine, a common pattern is to let System 1 answer when it is confident and hand the
511 rest to System 2. This is a usage pattern, not a configuration we have benchmarked; pick the threshold on your own
512 validation data, and serve with a `--max-model-len` large enough for the reasoning budget.
513
514 ```python
515 def solve(state, question, options, threshold=0.90):
516 p = decide("choice", state, question, options) # System 1: one prefill pass
517 best = max(p, key=p.get)
518 if p[best] >= threshold:
519 return {"system": 1, "answer": best, "distribution": p}
520 prompt = (f"{state}\n\nQuestion: {question}\nOptions: " + "; ".join(options)
521 + "\nThink it through, then give exactly one option on the last line.")
522 r = requests.post(f"{URL}/v1/chat/completions", json={ # System 2: same engine, base lm_head
523 "model": "autotrust/JEV-9B",
524 "messages": [{"role": "user", "content": prompt}],
525 "max_tokens": 8192, "chat_template_kwargs": {"enable_thinking": True}}).json()
526 return {"system": 2, "reply": r["choices"][0]["message"]["content"], "system1_distribution": p}
527 ```
528
529 ### Offline / batch (Python API)
530
531 ```python
532 from vllm import LLM, SamplingParams
533 from vllm.lora.request import LoRARequest
534
535 llm = LLM("JEV-9B", enable_lora=True, max_lora_rank=32, logprobs_mode="processed_logprobs", max_model_len=4096)
536 decision = LoRARequest("jev-decision", 1, "JEV-9B/adapter_vllm")
537
538 gen = llm.generate(["..."], SamplingParams(temperature=0.0, max_tokens=256)) # System 2, no LoRA
539 dec = llm.generate([prompt], [SamplingParams(max_tokens=1, temperature=1.0, # System 1
540 allowed_token_ids=ids, logprobs=len(ids))],
541 lora_request=decision) # then + bias, / T as above
542 ```
543
544 Mixed batches work too: pass a per-request `lora_request` list (`None` for System 2, `decision` for System 1) and both
545 systems are served in the same `generate` call.
546
547 ### Measured on one B200
548
549 | workload | PyTorch path | **vLLM** |
550 |---|---|---|
551 | System 2 — 164 HumanEval completions (greedy, ≤ 384 new tokens) | 165 s | **3.3 s** (≈ 50×) |
552 | System 1 — offline batch, 29,955 test questions | 75 s (398 q/s) | 80 s (374 q/s) |
553 | System 1 over HTTP — 64 / 256 concurrent clients | — | 150 / 205 req/s |
554 | System 1 fidelity vs. the PyTorch path | test KL 0.0210 | test KL 0.0211; mean \|Δp\| 0.0008 over HTTP |
555 | Many questions about one state, `--enable-prefix-caching` | — | +14–20 % throughput |
556
557 Notes:
558 * The big win is on System 2: generation is ≈ 50× faster. A decision is a single prefill pass with no decoding, so at
559 9 B offline batch throughput is about the same as the PyTorch path (at 27 B vLLM is 1.7× faster); for System 1 vLLM
560 mainly buys serving: continuous batching under concurrency, an OpenAI-compatible API, and one engine for both systems.
561 * Prefix caching: this architecture mixes Gated DeltaNet and attention layers, and vLLM caches it in blocks of 528
562 tokens, so only shared prefixes longer than 528 tokens are reused. The template puts `[kind]` before `[state]`, so
563 only questions of the same kind share a prefix. On 293 real states × 7.7 yes/no questions each (≈ 480-token states),
564 prefix caching served 19.6 % of prompt tokens from cache (+14–20 % throughput) with identical outputs.
565 * Requires a vLLM build with Qwen3.5 (`qwen3_5`) support, LoRA on `lm_head`, `--logprobs-mode` and
566 `allowed_token_ids`; tested with a vLLM development build from September 2026. Start-up takes 3–8 minutes
567 (CUDA-graph capture with LoRA enabled).
568
569 ## What System 1 does
570
571 | kind | question | returns |
572 |---|---|---|
573 | `noul` | "Is this statement true?" | `[P(false), P(true)]` |
574 | `choice` | "Which of these 2–16 options?" | one probability per option, aligned with your `options` |
575 | `score` | "Where on this ordered 0–5 scale?" | a distribution over the six levels (+ expected score) |
576
577 ```
578 [kind] choice
579 [state] SKU AX-330 stock at 8% of safety level; supplier late twice this quarter.
580 [question] Supplier response for this scenario.
581 [options]
582 A) issue_warning
583 B) renegotiate
584 C) dual_source
585 D) maintain
586 [decision]:
587 ```
588
589 The template is tokenised as one string; the last token's final-norm hidden state goes through a
590 **linear fp32 head `H → 24 slots`** (`noul` → slots 0–1, `score` → 2–7, `choice` → 8–23). Inactive slots are masked,
591 a per-kind temperature is applied, and a softmax yields the distribution aligned with your `options`. One prefill
592 pass, no decoding.
593
594 ## Other ways to run it
595
596 ### Plain `transformers` + `peft`
597
598 ```python
599 import json, torch
600 from huggingface_hub import hf_hub_download
601 from peft import PeftModel
602 from safetensors.torch import load_file
603 from transformers import AutoModelForCausalLM, AutoTokenizer
604
605 repo = "autotrust/JEV-9B"
606 tok = AutoTokenizer.from_pretrained(repo)
607 base = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda") # == Qwen3.5-9B text model
608
609 # --- System 2: the pristine base model, no adapter ----------------------------------------------
610 msgs = [{"role": "user", "content": "In two sentences, what is safety stock?"}]
611 enc = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
612 out = base.generate(**enc, max_new_tokens=80)
613 print(tok.decode(out[0, enc["input_ids"].shape[1]:], skip_special_tokens=True))
614
615 # --- System 1: attach the LoRA adapter (merged here for speed) + the 24-slot head ---------------
616 model = PeftModel.from_pretrained(base, repo, subfolder="adapter").merge_and_unload()
617 head = load_file(hf_hub_download(repo, "head.safetensors"))
618 cfg = json.load(open(hf_hub_download(repo, "judge_config.json")))
619 temp = json.load(open(hf_hub_download(repo, "calibration.json")))["per_kind"]
620 W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
621
622 def decide(kind, state, question, options):
623 letters = "ABCDEFGHIJKLMNOP"
624 lines = options if kind != "choice" else [f"{letters[i]}) {o}" for i, o in enumerate(options)]
625 text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
626 ids = tok(text, return_tensors="pt", add_special_tokens=False).to("cuda")
627 with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
628 h = model.model(**ids).last_hidden_state[0, -1].float() # backbone only, last token
629 z = (W @ h + b) / temp[kind]
630 s, _ = cfg["slots"]["ranges"][kind]
631 p = torch.softmax(z[s : s + len(options)], 0)
632 return dict(zip(options, p.tolist()))
633
634 print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
635 "Is the customer asking for a refund?", ["false", "true"]))
636 # {'false': 0.009, 'true': 0.991}
637 ```
638
639 `options` are validated: `noul` must be `["false","true"]`, `score` must be `["0".."5"]`, `choice` takes 2–16
640 free-text options. Note that `merge_and_unload()` above changes the backbone for the rest of the process; to keep both
641 systems in one process, leave the adapter unmerged and run System 2 inside `with model.disable_adapter():`.
642
643 ## Evaluation details
644
645 ### Additional System 1 metrics (`test_set_30k`, temperature applied)
646
647 | metric | autotrust/JEV-9B |
648 |---|---|
649 | `noul` Brier score against the target probability, all rows (lower is better) | 0.0015 |
650 | `score` ranked probability score (lower is better) | 0.0085 |
651 | Fitted temperatures noul / choice / score | 1.002 / 0.984 / 1.012 |
652 | Top-1 flip rate when `choice` options are shuffled (1,000 rows × 4 permutations) | 3.9 % |
653 | Out-of-distribution split — top-1 agreement · `noul` AUROC | 0.918 · 0.989 |
654 | Throughput — batch of 128 requests on one B200 | 2.5 ms per decision (≈ 400 decisions/s) |
655 | Single request on one B200 (median) | 87 ms |
656
657 ### System 2 — no-degradation check (HumanEval, greedy pass@1, completion-style prompt)
658
659 | weights | pass@1 | note |
660 |---|---|---|
661 | Qwen3.5-9B (base) | 70.7 % (116/164) | same loader and protocol as below |
662 | **autotrust/JEV-9B — System 2 path (backbone + `lm_head`, adapter off)** | **70.7 % (116/164)** | all 164 completions byte-identical to the base model |
663 | System 1 LoRA folded into the backbone + base `lm_head` (*not shipped*) | 61.6 % (101/164) | why the blocks are kept separate |
664
665 ### Per source × primitive (`test_set_30k`, temperature applied)
666
667 | source | kind | n | KL | top-1 | ECE | noul AUROC | score MAE |
668 |---|---|---|---|---|---|---|---|
669 | yuri_v3 — synthetic operational scenarios, labelled by TypeSafe Jev 1.13 | noul | 8,537 | 0.005 | 0.961 | 0.001 | 0.994 | — |
670 | yuri_v3 | choice | 8,312 | 0.028 | 0.902 | 0.002 | — | — |
671 | yuri_v3 | score | 8,527 | 0.023 | 0.883 | 0.002 | — | 0.103 |
672 | openjev_v2 — Open-Jev programmatic tasks, ground-truth labels (not Jev) | noul | 1,432 | 0.004 | 0.998 | 0.003 | 1.000 | — |
673 | openjev_v2 | choice | 887 | 0.176 | 0.857 | 0.020 | — | — |
674 | yuri_v1 — placeholder `[0.5, 0.5]` labels (see Limitations) | noul | 2,260 | 0.000 | — | 0.005 | — | — |
675
676 OOD split (13,058 Open-Jev rows from task families not in training, programmatic labels): KL 0.234, top-1 0.918, noul
677 AUROC 0.989; `choice` KL 0.351 / top-1 0.837 (game-state decisions are the hardest slice).
678
679 Choice option-permutation consistency (1,000 rows × 4 random permutations): mean max |Δp| 0.024, p90 0.055, top-1
680 flip rate 3.9 %.
681
682 ### Training trajectory (most recent first)
683
684 Validation KL on a fixed 4 k-row subset; `test_set_30k` metrics after calibration.
685
686 | stage | rows seen | val KL | t30k KL | choice top-1 | score MAE | noul AUROC | ECE |
687 |---|---|---|---|---|---|---|---|
688 | **autotrust/JEV-9B v0.8.0 — released weights (4,750 steps ≈ 0.93 epoch, LR annealed to ≈ 0.07×)** | 608 k | **0.019** | **0.0210** | **0.898** | **0.103** | **0.996** | **0.0007** |
689 | v0.7.0 (step 2000 + 500-step LR cool-down) | 320 k | 0.026 | 0.0276 | 0.884 | 0.119 | 0.994 | 0.0014 |
690 | step 2000 | 256 k | 0.0325 | 0.037 | 0.865 | 0.143 | 0.992 | 0.004 |
691 | step 1500 | 192 k | 0.038 | 0.040 | 0.861 | 0.151 | 0.991 | 0.0025 |
692 | step 500 | 64 k | 0.094 | 0.081 | 0.817 | 0.224 | 0.977 | 0.022 |
693 | untrained backbone with the initialised head (reference point, not the model) | 0 | 0.485 | 0.510 | 0.532 | 1.130 | 0.824 | 0.094 |
694
695 Annealing matters: the v0.7.0 cool-down (500 steps, lr ×0.9 → ×0.02 from step 2000) lowered KL by 25 %; continuing on
696 the unseen remainder of the epoch with the learning rate decayed to ≈ 0.07× (v0.8.0) lowered it by another 24 % and
697 added 1.4 points of choice agreement. JEV-27B folds this into a single cosine schedule.
698
699 ## Training details
700
701 | item | value |
702 |---|---|
703 | teacher / data | `SargeDev/jev-distill-corpus-v3` (740,957 rows; `train` 655,806) with three streams: `yuri_v3` (498,010 rows, TypeSafe Jev 1.13 full output distributions via OpenRouter), `openjev_v2` (94,801 rows, Open-Jev programmatic labels, CC0), `yuri_v1` (148,154 rows, placeholder labels, down-weighted) |
704 | backbone | `Qwen/Qwen3.5-9B` text tower only (vision tower and MTP head dropped), bf16, frozen |
705 | System 1 block (trainable) | LoRA r=16, α=32, dropout 0.05 on `in_proj_qkv, in_proj_z, out_proj, q/k/v/o_proj, gate/up/down_proj` (40.1 M, shipped unmerged in `adapter/`) + 24-slot head (98 k, fp32, initialised from `lm_head` rows) |
706 | System 2 block | the original `lm_head`, not trained |
707 | loss | KL(target ‖ model) over active slots + 0.5 · RPS (ranked probability score) for `score` |
708 | augmentation | 30 % random permutation of `choice` options (targets permuted consistently) |
709 | batching | 128 rows / step, kind-stratified (≥ 1/6 per primitive), length-bucketed, micro-batches capped at 24 k padded tokens, gradient checkpointing |
710 | optimiser | AdamW (fused), β=(0.9, 0.98), lr head 2e-4 / LoRA 1e-4, cosine, warmup 3 %, grad-clip 1.0 for 2,500 steps; then continued on the unseen remainder of the epoch (fresh AdamW state, warmup 2 %, lr ×0.9 → cosine) and stopped after 2,250 more steps at lr ≈ ×0.07 — 4,750 steps ≈ 0.93 epoch in total |
711 | label hygiene | `yuri_v1` rows carry exact-uniform `[0.5, 0.5]` placeholder labels (137,203 rows, 100 %); down-weighted ×0.05 in training and excluded from temperature fitting |
712 | calibration | per-kind scalar temperature (L-BFGS on the `calibration` split, 10,954 rows): noul 1.002 · choice 0.984 · score 1.012 |
713 | compute | 1× NVIDIA B200 (183 GB); ≈ 1.4 h (2,500 steps) + ≈ 1.5 h (2,250 steps) ≈ 3 GPU-hours; ≈ 7–9 k tokens/s |
714 | software | torch 2.13 + cu130, transformers 5.16, peft 0.21, flash-linear-attention 0.5.2 |
715
716 ## Limitations
717
718 * **System 1 mirrors TypeSafe Jev 1.13, including its mistakes.** This is a distillation, not an independent judge:
719 where the teacher was wrong or uncalibrated, so is autotrust/JEV-9B. Published evaluations of the teacher show it is
720 unreliable for multi-hop reasoning, arithmetic, dates, counting and adversarial inputs, and the student inherits all
721 of that. Confirmed on fresh inputs: the poker shove (0.70 vs the teacher's 0.62).
722 * **At 9 B it adds some blind spots of its own.** 11.5 % of its 16-option answers change with option order alone
723 (teacher 7.0 %), it keeps 90 % rather than 96 % of the teacher's accuracy at 16 options, and it missed a code-rule
724 violation that JEV-27B catches. Prefer [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B) for long option
725 lists, unfamiliar task families and code-rule checks.
726 * **The two systems are integrated in serving, not in knowledge.** System 1 cannot explain its decisions, and System 2
727 is the unmodified base model: it knows nothing about the decisions it is packaged with and was not trained to agree
728 with System 1. If you escalate from System 1 to System 2, expect them to disagree sometimes.
729 * **"Indistinguishable" is a KL statement on Jev-labelled rows from the 53 training domains.** The corpus has no
730 Jev-labelled out-of-distribution set; the OOD figures (KL 0.234, 0.351 for game-state `choice`) are measured against
731 programmatic ground truth, and on the independent benchmark the student reaches 90–97 % of Jev's accuracy, not 100 %.
732 * **Speed comparisons with the hosted API are not like for like.** Our timings exclude network time; the hosted
733 figures are third-party measurements that include it and depend on client concurrency and rate limits.
734 * **Choice agreement is capped by teacher ambiguity.** The teacher's `choice` distributions are soft (median top-1
735 probability 0.70). On the 14 % of rows where the teacher's top two options are within 0.1 of each other, argmax
736 agreement is near chance for *any* faithful mimic (0.46 where the gap is < 0.05). On teacher-decisive rows agreement
737 is 0.954, and the student's argmax captures 97.7 % of the teacher probability mass a perfect mimic could (0.693 vs
738 0.709).
739 * **Fixed option sets.** `noul` and `score` accept only their canonical options; `choice` accepts 2–16 options. Inputs
740 longer than 1,024 tokens are truncated (state only, head 60 % / tail 40 %) at serving unless you raise the limit.
741 * **English-centric.** The corpus is English; multilingual behaviour is inherited from the backbone and was not
742 systematically measured (the Chinese V2EX examples above are illustrations only).
743 * **Placeholder labels in the corpus.** The `yuri_v1` memory-relevance stream is 100 % exact-uniform `[0.5, 0.5]` —
744 those rows teach nothing about relevance. The model outputs ≈ 0.5 on them by design; do not use it for
745 memory-relevance scoring without further training.
746 * **Not for high-stakes decisions.** Use confidence gating: act automatically only above a threshold you validated on
747 your own data, and route the rest to System 2, a stronger model, or a human.
748
749 ## Files
750
751 ```
752 model-0000{1..5}-of-00005.safetensors Qwen3.5-9B text backbone incl. lm_head — bit-identical to the base model
753 (bf16; GDN A_log / gated-norm weights fp32 as in the original), 17.9 GB
754 model.safetensors.index.json · config.json
755 adapter/ System 1 LoRA (peft format, r=16, 40.1 M params, 154 MB) — apply only for decisions
756 head.safetensors 24-slot decision head (fp32): proj.weight [24, 4096], proj.bias [24]
757 judge_config.json slot layout, verbalizer token ids, template version, weights_mode=unmerged, provenance
758 calibration.json per-kind temperatures (+ fit diagnostics)
759 adapter_vllm/ the same adapter for vLLM: backbone LoRA (zero-padded to r=32) + decision head as an
760 lm_head LoRA, plus decision_head.json (head bias, verbalizer token ids)
761 tokenizer.json · tokenizer_config.json · chat_template.jinja
762 reports/ evaluation reports: test-set evaluation, bundle checks, HumanEval per-problem
763 results, vLLM measurements, real-world tests, training-milestone reviews
764 vl/ vision: serve.sh (multimodal Qwen3.5-9B + System 1), serve_decide.py (POST /v1/decide),
765 adapter_vllm/ (layer names for the multimodal model), calibration.json, demos/
766 videos/ robot-arm and computer-use demo videos
767 reports/vl/ image evaluations and demo results
768 ```
769
770 ## License and acknowledgements
771
772 Weights: **Apache-2.0** (base model `Qwen/Qwen3.5-9B` is Apache-2.0; training corpus
773 `SargeDev/jev-distill-corpus-v3` is Apache-2.0, its `openjev_v2` stream additionally CC0). The System One framing and
774 the `noul` / `choice` / `score` primitives originate with TypeSafe AI's Jev; autotrust/JEV-9B is an independent
775 student model trained on public data and shares no weights, code or affiliation with TypeSafe AI.
776
777 ```bibtex
778 @misc{autotrust_jev9b_2026,
779 title = {autotrust/JEV-9B: the first integrated System 1 + System 2 open model built with the Blocks of Experts recipe (Qwen3.5-9B; System 1 distilled from TypeSafe Jev 1.13)},
780 author = {{AutoTrust AI}},
781 year = {2026},
782 url = {https://huggingface.co/autotrust/JEV-9B}
783 }
784 ```
785