README.md
| 1 | --- |
| 2 | license: apache-2.0 |
| 3 | base_model: Qwen/Qwen3.5-9B |
| 4 | base_model_relation: finetune |
| 5 | datasets: |
| 6 | - SargeDev/jev-distill-corpus-v3 |
| 7 | language: |
| 8 | - en |
| 9 | library_name: transformers |
| 10 | pipeline_tag: text-classification |
| 11 | tags: |
| 12 | - system-one |
| 13 | - system-two |
| 14 | - blocks-of-experts |
| 15 | - typed-decisions |
| 16 | - decision-model |
| 17 | - calibrated-probabilities |
| 18 | - knowledge-distillation |
| 19 | - jev |
| 20 | - noul |
| 21 | - choice |
| 22 | - score |
| 23 | - lora |
| 24 | - qwen3_5 |
| 25 | - text-generation |
| 26 | - dual-head |
| 27 | - vllm |
| 28 | - multimodal |
| 29 | - vision |
| 30 | - computer-use |
| 31 | - robotics |
| 32 | metrics: |
| 33 | - kl |
| 34 | - auroc |
| 35 | - brier |
| 36 | - ece |
| 37 | model-index: |
| 38 | - name: autotrust/JEV-9B (student of TypeSafe Jev 1.13) |
| 39 | results: |
| 40 | - task: |
| 41 | type: text-classification |
| 42 | name: typed decisions (noul / choice / score) — agreement with the TypeSafe Jev 1.13 teacher |
| 43 | dataset: |
| 44 | type: SargeDev/jev-distill-corpus-v3 |
| 45 | name: jev-distill-corpus-v3 · test_set_30k |
| 46 | split: test_set_30k |
| 47 | metrics: |
| 48 | - type: kl_divergence |
| 49 | name: mean KL(target ‖ model), all test rows (25,376 of 29,955 targets are TypeSafe Jev 1.13 distributions) |
| 50 | value: 0.0210 |
| 51 | - type: auroc |
| 52 | name: noul AUROC |
| 53 | value: 0.996 |
| 54 | - type: brier |
| 55 | name: noul Brier (vs. target probability, all rows) |
| 56 | value: 0.0015 |
| 57 | - type: mae |
| 58 | name: score expected-value MAE (0–5 scale) |
| 59 | value: 0.103 |
| 60 | - type: ece |
| 61 | name: ECE (15 bins, after temperature) |
| 62 | value: 0.0007 |
| 63 | - type: accuracy |
| 64 | name: choice top-1 agreement (all rows) |
| 65 | value: 0.898 |
| 66 | - type: accuracy |
| 67 | name: choice top-1 agreement (decisive-target rows, top-2 gap ≥ 0.1) |
| 68 | value: 0.954 |
| 69 | - task: |
| 70 | type: text-generation |
| 71 | name: code generation — System 2 path (base lm_head, adapter off) |
| 72 | dataset: |
| 73 | type: openai/openai_humaneval |
| 74 | name: HumanEval |
| 75 | split: test |
| 76 | metrics: |
| 77 | - type: pass@1 |
| 78 | name: pass@1 (greedy, completion-style prompt) |
| 79 | value: 0.707 |
| 80 | --- |
| 81 | |
| 82 | # autotrust/JEV-9B |
| 83 | |
| 84 | ### AutoTrust's first integrated System 1 + System 2 open model, built with the Blocks of Experts recipe |
| 85 | |
| 86 | **Fast, calibrated System 1 decisions that are indistinguishable from the closed TypeSafe Jev 1.13 by KL, and |
| 87 | deliberate System 2 generation and reasoning from an untouched Qwen3.5-9B — one set of weights, one vLLM engine, |
| 88 | routed per request. The fastest model of the family: it answers a single decision in about a third of the time the |
| 89 | hosted API takes. Its successor, [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B), is closer to Jev, |
| 90 | transfers better to unseen tasks and has a stronger System 2.** |
| 91 | |
| 92 | ## New (3 October 2026): JEV-9B can see — robot arm and computer use |
| 93 | |
| 94 | JEV-9B now takes images. Every step below is **one System 1 decision**: camera image or screenshot in, a probability |
| 95 | for every action out, in a single forward pass (about 0.2 s on one GPU). Run it with `bash vl/serve.sh` (see |
| 96 | [Images: quick start](#images-quick-start)). |
| 97 | |
| 98 | **Robot arm: pick and place from camera images.** The arm sees a top camera image; at every step System 1 answers two |
| 99 | questions (is the target left or right of the gripper, above or below it), and the arm moves accordingly, halving its |
| 100 | step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation). |
| 101 | |
| 102 | <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/robot_arm_pick_place.mp4" controls autoplay loop muted playsinline width="100%"></video> |
| 103 | |
| 104 | On 20 random scenes it completed the task 10 times; every cube it grasped ended in the tray, and every miss was a grasp |
| 105 | 3–5 cm off target. About 165 ms per decision. Asking it to choose one of 8 motor commands directly did not work: this |
| 106 | model is a fast visual judge, not an end-to-end controller. |
| 107 | |
| 108 | **Computer use: screenshot → which element to click.** A real browser (headless Chromium). Every clickable element gets |
| 109 | a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats. |
| 110 | |
| 111 | <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_shop.mp4" controls autoplay loop muted playsinline width="100%"></video> |
| 112 | |
| 113 | <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_settings.mp4" controls autoplay loop muted playsinline width="100%"></video> |
| 114 | |
| 115 | <video src="https://huggingface.co/autotrust/JEV-9B/resolve/main/videos/computer_use_mail.mp4" controls autoplay loop muted playsinline width="100%"></video> |
| 116 | |
| 117 | **95% of 60 random multi-step tasks completed** (shop, settings, mail; 3–7 clicks each), about 0.2 s per click. The |
| 118 | colour swatches and switches carry no text, so those clicks are decided from the screenshot alone. With the numbered |
| 119 | boxes only (no element text) it completed 37%. The failures skipped a step (the colour) and then checked out an empty cart. |
| 120 | |
| 121 | Code for both demos: [`vl/demos/`](vl/demos). Image judging, briefly: VL-RewardBench 74.3%, AgentRewardBench AUROC 0.91, |
| 122 | zero-shot short-video recommendation from covers AUC 0.72 (details in [`reports/vl/`](reports/vl)). |
| 123 | |
| 124 | ### Images: quick start |
| 125 | |
| 126 | ```bash |
| 127 | hf download autotrust/JEV-9B --include "vl/*" --local-dir JEV-9B |
| 128 | bash JEV-9B/vl/serve.sh # downloads Qwen/Qwen3.5-9B (with its vision encoder) and serves both systems on :8000 |
| 129 | ``` |
| 130 | |
| 131 | ```python |
| 132 | import base64, requests |
| 133 | |
| 134 | def image(path): |
| 135 | return {"image": "data:image/png;base64," + base64.b64encode(open(path, "rb").read()).decode()} |
| 136 | |
| 137 | r = requests.post("http://localhost:8000/v1/decide", json={ |
| 138 | "kind": "choice", |
| 139 | "state": ["Top camera image:", image("scene.png"), "\nTask: put the red cube in the tray."], |
| 140 | "question": "Is the red cube to the left or to the right of the gripper?", |
| 141 | "options": ["left", "right"]}).json() |
| 142 | print(dict(zip(r["options"], r["probabilities"]))) |
| 143 | ``` |
| 144 | |
| 145 | How it works: JEV-9B's language weights are bit-identical to Qwen3.5-9B's, so `vl/serve.sh` serves the unmodified |
| 146 | multimodal Qwen3.5-9B with JEV-9B's System 1 adapter (`vl/adapter_vllm`, the same weights with the layer names moved). |
| 147 | Text decisions match the text-only model (300 test decisions: largest probability difference 0.011). System 2 also reads |
| 148 | images. Keep `--max-num-seqs 8` (set in `serve.sh`); decisions over images are zero-shot. |
| 149 | |
| 150 | ## At a glance |
| 151 | |
| 152 | **Integrated System 1 + System 2, first generation.** JEV-9B is AutoTrust's first open model to serve both modes of |
| 153 | thinking from a single set of weights; the second generation is |
| 154 | [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B). *System 1* answers typed questions (`noul` yes/no · |
| 155 | `choice` over 2–16 options · `score` on a 0–5 scale) in one forward pass and returns a calibrated probability |
| 156 | distribution. *System 2* is ordinary text generation with step-by-step reasoning (thinking mode). Both run on the |
| 157 | same backbone in the same engine, and a request chooses its system. |
| 158 | |
| 159 | **Blocks of Experts recipe.** Rather than fine-tuning one monolithic model, the Blocks of Experts (BoE) recipe keeps a |
| 160 | strong pretrained model as a frozen expert block and adds a small, detachable expert block trained for one capability. |
| 161 | In JEV-9B the System 2 block is Qwen3.5-9B, bit-identical to the release; the System 1 block is 40.2 M trained |
| 162 | parameters (0.5 % of the backbone), trained in ≈ 3 hours on one B200. Because the blocks stay separate, adding |
| 163 | System 1 costs System 2 nothing: HumanEval is 70.7 % before and after, with all 164 completions byte-identical. Folding |
| 164 | the same block into the backbone instead would have cost 9 points (61.6 %). |
| 165 | |
| 166 | **Indistinguishable from the closed original on System 1, by KL.** On the 25,376 held-out questions (53 domains) |
| 167 | whose labels are TypeSafe Jev 1.13's own output distributions, the mean KL divergence is **≈ 0.019** (0 = identical). |
| 168 | An observer who sees sampled decisions gains on average 0.019 nats of evidence per decision about which model produced |
| 169 | it, so it takes about 54 decisions to gather a single nat. The fidelity extends to the teacher's mistakes (see |
| 170 | [System 1: indistinguishable from TypeSafe Jev 1.13](#system-1-indistinguishable-from-typesafe-jev-113-by-kl)). |
| 171 | Among the open Jev reproductions we could find, only the JEV models publish this distribution-level measure |
| 172 | (see [How JEV-9B compares with other open Jev reproductions](#how-jev-9b-compares-with-other-open-jev-reproductions)). |
| 173 | |
| 174 | **Faster than the hosted API.** On one B200, a single decision takes a median ≈ 90 ms, against 238–301 ms measured |
| 175 | independently for the hosted TypeSafe Jev 1.13 API, and one GPU sustains about 15× the decisions per second an |
| 176 | independent benchmark achieved against that API (see [Speed](#speed-vs-the-hosted-typesafe-jev-113)). |
| 177 | |
| 178 | **The fast member of the family; JEV-27B is the closer one.** Same recipe, same API: JEV-9B is 2.6× faster than |
| 179 | JEV-27B on the same benchmark and its weights are a third of the size (18 GB vs 54 GB). JEV-27B lowers mean KL to Jev's |
| 180 | distributions from ≈ 0.019 to ≈ 0.017, more than halves KL on unseen task families (0.234 → 0.104), keeps 96 % instead |
| 181 | of 90 % of the teacher's accuracy on an independent 16-option benchmark, and scores 78.0 % instead of 70.7 % on |
| 182 | HumanEval (see [JEV-9B vs JEV-27B](#jev-9b-vs-jev-27b)). |
| 183 | |
| 184 | > **Two models, two organisations.** **TypeSafe Jev 1.13** is the hosted, closed-source model made by TypeSafe AI; it |
| 185 | > is the *teacher* whose published output distributions this model was trained on. **autotrust/JEV-9B** (this |
| 186 | > repository) is an independent open-weights *student* built by AutoTrust AI from the Apache-2.0 corpus |
| 187 | > [`SargeDev/jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3). It is not |
| 188 | > affiliated with, endorsed by, or a product of TypeSafe AI, and shares no weights or code with it. |
| 189 | |
| 190 | ## Headline results |
| 191 | |
| 192 | System 1 numbers are on the held-out `test_set_30k` of `jev-distill-corpus-v3`. Its 29,955 rows come from three |
| 193 | sources: 25,376 rows labelled with TypeSafe Jev 1.13's own output distributions (`yuri_v3`), 2,319 Open-Jev rows with |
| 194 | programmatic ground-truth labels (`openjev_v2`), and 2,260 placeholder rows (`yuri_v1`). Rows marked *Jev-labelled* use |
| 195 | only the first group. |
| 196 | |
| 197 | | | What is measured | autotrust/JEV-9B | How to read it | |
| 198 | |---|---|---|---| |
| 199 | | **System 1** | Mean KL divergence from TypeSafe Jev 1.13's distributions, Jev-labelled rows, 0 = identical | **≈ 0.019** | Indistinguishable from the teacher's decisions at this resolution: ≈ 54 sampled decisions to gather one nat of evidence | |
| 200 | | | Mean KL to all test targets (Jev, programmatic and placeholder labels) | **0.021** | The figure in the model index above | |
| 201 | | | Yes/no AUROC (`noul`), Jev-labelled rows | **0.994** | Ranks true vs. false almost perfectly (0.996 over all rows) | |
| 202 | | | Choice top-1 agreement with Jev, Jev-labelled rows | **90.2 %** | 95.4 % over all rows where the target's top two options differ by ≥ 0.1; on near ties any faithful copy agrees about half the time | |
| 203 | | | Rating error (`score`, 0–5 scale), mean absolute error of the expected rating | **0.103** | About one tenth of a rating step | |
| 204 | | | Expected calibration error | **0.0007** | A stated 80 % is an 80 %; fitted temperatures ≈ 1.00, no post-hoc correction needed | |
| 205 | | | KL to the programmatic labels of task families never seen in training (Open-Jev OOD split) | **0.234** | Transfer to new tasks; these labels are ground truth, not Jev's outputs. JEV-27B: 0.104 | |
| 206 | | | Independent benchmark with human gold labels, 16 options | **90 % of the teacher** (0.694 vs 0.769) | 94–97 % of the teacher at 2, 4 and 8 options; see [Benchmark highlights](#benchmark-highlights) | |
| 207 | | **System 2** | HumanEval pass@1, greedy | **70.7 %** | Identical to Qwen3.5-9B (116/164); all 164 completions byte-identical to the base model | |
| 208 | | **Speed** | Single decision, median, one B200 | **≈ 90 ms** | Hosted TypeSafe Jev 1.13, measured independently: 238 ms mean, 291–301 ms median | |
| 209 | | | Decisions per second on the independent benchmark, one B200 | **≈ 340** | ≈ 15× the 23 per second measured against the hosted API; see [Speed](#speed-vs-the-hosted-typesafe-jev-113) | |
| 210 | | | Batched, 128 decisions per batch | **2.5 ms** per decision | With vLLM: 205 decisions/s over HTTP at 256 concurrent clients, text generation ≈ 50× faster than the PyTorch path | |
| 211 | | **Efficiency** | Trained parameters | **40.2 M** (0.5 % of 7.9 B) | ≈ 3 B200-hours, 0.93 epoch ≈ 608 k rows | |
| 212 | |
| 213 | ## JEV-9B vs JEV-27B |
| 214 | |
| 215 | JEV-9B was AutoTrust's first integrated System 1 + System 2 model. |
| 216 | [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B), the second generation, uses the same recipe, code, |
| 217 | hyper-parameters, API and two-block packaging; only the backbone and memory settings changed. Both are evaluated on the |
| 218 | same held-out test set and the same independent benchmark. |
| 219 | |
| 220 | <p align="center"> |
| 221 | <a href="https://huggingface.co/autotrust/JEV-27B/blob/main/27b-2.jpg"><img src="https://huggingface.co/autotrust/JEV-27B/resolve/main/27b-2.jpg" alt="JEV family benchmark highlights: KL to TypeSafe Jev 1.13 by question type, accuracy as a percentage of Jev on an independent benchmark, HumanEval for the System 2 path, and speed against the hosted API, for JEV-9B (light bars) and JEV-27B (dark bars)" width="100%"></a> |
| 222 | <br> |
| 223 | <sub><b>JEV family benchmark highlights</b> (chart from the JEV-27B repository; light bars = JEV-9B). A · KL to Jev by question type · B · accuracy as % of Jev on an independent benchmark · C · System 2 unchanged · D · speed vs the hosted API · click to enlarge</sub> |
| 224 | </p> |
| 225 | |
| 226 | | | **JEV-9B** | JEV-27B | JEV-27B vs JEV-9B | |
| 227 | |---|---|---|---| |
| 228 | | Backbone | Qwen3.5-9B | Qwen3.8-27B | | |
| 229 | | **System 1** — mean KL to TypeSafe Jev 1.13, Jev-labelled rows | ≈ 0.019 | **≈ 0.017** | ≈ −11 % | |
| 230 | | Mean KL to all test targets | 0.021 | **0.019** | −11 % | |
| 231 | | KL to ground-truth labels, unseen task families (OOD) | 0.234 | **0.104** | −56 % | |
| 232 | | Top-1 accuracy, unseen task families (OOD) | 0.918 | **0.942** | +2.4 pts | |
| 233 | | Choice top-1 agreement with Jev, Jev-labelled rows | 90.2 % | **90.5 %** | +0.3 pts | |
| 234 | | Rating error (`score` MAE, all Jev-labelled) | 0.103 | **0.098** | −5 % | |
| 235 | | Top-1 flips under option shuffle (test set) | 3.9 % | **2.9 %** | −1.0 pt | |
| 236 | | Yes/no AUROC (`noul`), Jev-labelled rows | 0.994 | **0.995** | +0.001 | |
| 237 | | Calibration error (ECE) | **0.0007** | 0.0009 | JEV-9B slightly lower; both below 0.001 | |
| 238 | | Independent benchmark, 16 options — % of teacher accuracy | 90 % | **96 %** | +6 pts | |
| 239 | | Independent benchmark — answers changed by option order alone (teacher: 7.0 %) | 11.5 % | **7.4 %** | JEV-27B is close to the teacher's 7.0 % | |
| 240 | | **System 2** — HumanEval pass@1 (greedy) | 70.7 % | **78.0 %** | +7.3 pts | |
| 241 | | Latency on one B200 — single request / batched | **≈ 90 ms / 2.5 ms** | 137 ms / 4.2 ms | JEV-9B is faster | |
| 242 | | Benchmark throughput — 14,400 decisions on one B200 | **42 s** | 110 s | JEV-9B is 2.6× faster | |
| 243 | | Download size (backbone + adapter) | **18 GB** | 54 GB | | |
| 244 | | Trained parameters / compute | 40.2 M / ≈ 3 B200-hours | 108.9 M / ≈ 9.2 B200-hours | | |
| 245 | |
| 246 | On the fresh Hacker News, V2EX and community examples (illustrations, not a benchmark), JEV-9B got 92 of 96 decisions |
| 247 | right against 95 of 96 for JEV-27B. The difference is on the harder tasks: JEV-9B misses a TypeScript port that |
| 248 | breaks a "branded, range-checked integer" rule (0.33; JEV-27B 0.93) and flags a CEO wire-transfer fraud message with |
| 249 | less confidence (0.56; JEV-27B 0.84). |
| 250 | |
| 251 | **Which to pick.** For routing, moderation, topic triage and short option lists, JEV-9B gives nearly the same answers |
| 252 | 2.6× faster (14,400 benchmark decisions in 42 s vs 110 s on one B200) with a third of the weight memory. For long option lists |
| 253 | (more than about 8), unfamiliar task families, code-rule checks, fraud screening, or when the System 2 path matters, use |
| 254 | [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B). |
| 255 | |
| 256 | ## How JEV-9B compares with other open Jev reproductions |
| 257 | |
| 258 | Dozens of open reproductions of TypeSafe Jev appeared within weeks of its launch; the community |
| 259 | [Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index) (formerly the Jev Reproductions |
| 260 | Tracker) evaluates 55 of them. Most are trained on human or programmatic gold labels, or on their own synthetic data, so |
| 261 | they aim to match or beat Jev's accuracy rather than reproduce its probabilities. "Closest to Jev" therefore depends on |
| 262 | how closeness is measured: |
| 263 | |
| 264 | | measure of closeness to TypeSafe Jev 1.13 | published results (snapshot of 25 September 2026) | where JEV-9B stands | |
| 265 | |---|---|---| |
| 266 | | **Distribution level:** KL to Jev's own output distributions on held-out rows | JEV-27B ≈ 0.017 and JEV-9B ≈ 0.019 on 25,376 Jev-labelled rows. We found no other open reproduction that publishes this measure. | Second lowest published, after JEV-27B | |
| 267 | | **Accuracy relative to Jev** on [`decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure), 16 options, human gold labels | JEV-27B 96 % · JEV-9B 90 % · Laya 90 % · DeBERTa-v3-large zero-shot 90 % · DeBERTa-v3-base zero-shot 83 % · GLiClass-large 81 % · bge-large 73 % · gte-large 69 % | Level with the best of the other models measured there; JEV-27B is closer (JEV rows are AutoTrust re-runs of the same items; the others were run by the benchmark's author) | |
| 268 | | **Score parity on community leaderboards** | [Decision Index 0.2](https://huggingface.co/spaces/multimodalart/jev-decision-index): Jev 51.67, AutoJev-27B 50.94. [JevBench v1.4.2](https://github.com/fstandhartinger/jevbench): decider-4b v2 64.13, Jev 63.29, JevK5 62.04. [Open-Jev](https://zefan-cai.github.io/open-jev/benchmarks/) public JevBench subset: Jev 200/231, Open-Jev 27B v1.1 197/231 | Not yet evaluated | |
| 269 | |
| 270 | On the evidence published today, the two JEV models are the closest open models to TypeSafe Jev 1.13 at the level of |
| 271 | output distributions, with JEV-9B second to JEV-27B. On the independent benchmark JEV-9B is level with the best of the |
| 272 | other models measured there, not ahead of them. It has not yet been run on the Decision Index or JevBench, where |
| 273 | AutoJev-27B scores within about one point of Jev and decider-4b v2 edges ahead of it, so we do not claim it is the |
| 274 | closest by every measure. Note that some reproductions report beating Jev on their own test sets (AutoJev-27B reports |
| 275 | 84.60 % against Jev's 82.79 %); that is a different goal from reproducing Jev's behaviour. |
| 276 | |
| 277 | *Not to be confused with AutoJev-27B (`denis-pplx/autojev-27b`), an unrelated Qwen3.8-27B decision model trained with |
| 278 | full-weight SFT on its own data.* |
| 279 | |
| 280 | ## Speed vs the hosted TypeSafe Jev 1.13 |
| 281 | |
| 282 | TypeSafe does not publish Jev's size or hardware; it reports 70–500 ms end to end. Independent measurements, and ours: |
| 283 | |
| 284 | | | TypeSafe Jev 1.13, hosted API | **JEV-9B, one B200** | JEV-27B, one B200 | |
| 285 | |---|---|---|---| |
| 286 | | One decision, single request | 238 ms mean over 29,600 calls ([`decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure)); 291–301 ms median on three workloads ([Open-Jev](https://zefan-cai.github.io/open-jev/benchmarks/)) | **≈ 90 ms** median (87 ms) | 137 ms median | |
| 287 | | Decisions per second on `decision-models-under-pressure` | 23, with 5 client workers and one question per call | **≈ 340** (14,400 in 42 s) | ≈ 130 (14,400 in 110 s) | |
| 288 | | Batched, 128 decisions per batch | — | **2.5 ms** per decision | 4.2 ms per decision | |
| 289 | |
| 290 | So JEV-9B answers a single decision in roughly a third of the time (JEV-27B in roughly half), and one GPU sustains |
| 291 | about 15× (JEV-27B: about 6×) the throughput the benchmark's author achieved against the hosted API. Read these with |
| 292 | the caveats: our latencies are measured on the serving host with no network hop, while the hosted numbers include |
| 293 | internet, TLS and queueing; hosted throughput depends on client concurrency and the API's rate limits; Jev's latency is |
| 294 | roughly flat in the number of questions per request, so bundling questions narrows the throughput gap; and our figures |
| 295 | are self-reported while Jev's come from third parties. The two throughput runs use the same benchmark but not an |
| 296 | identical call set (ours stops at 16 options). |
| 297 | |
| 298 | ## System 1: indistinguishable from TypeSafe Jev 1.13, by KL |
| 299 | |
| 300 | **What the number means.** KL(Jev ‖ model) is the expected log-likelihood ratio, per sampled decision, between |
| 301 | TypeSafe Jev 1.13 and the student when the decision comes from Jev. On the 25,376 held-out rows whose targets are |
| 302 | Jev's own output distributions, the mean is ≈ 0.019 nats (computed from the per-slice values below, which are |
| 303 | published to three decimals): one decision carries almost no evidence about which of the two models produced it, and |
| 304 | an observer needs about 1 / KL ≈ 54 independent decisions to accumulate one nat (a likelihood ratio of about e ≈ 2.7 : 1). |
| 305 | |
| 306 | For scale, Jev is not deterministic itself: an independent study found it changes its answer on 4.3 % of repeated, |
| 307 | identical 64-option calls, and it returns probabilities rounded to two decimals, which is the resolution of the |
| 308 | targets used here. |
| 309 | |
| 310 | | Jev-labelled slice (`yuri_v3`, `test_set_30k`) | n | KL | ≈ decisions to gather one nat (1 / KL) | |
| 311 | |---|---|---|---| |
| 312 | | `noul` | 8,537 | 0.005 | ≈ 200 | |
| 313 | | `choice` | 8,312 | 0.028 | ≈ 36 | |
| 314 | | `score` | 8,527 | 0.023 | ≈ 43 | |
| 315 | | **all Jev-labelled rows** | **25,376** | **≈ 0.019** | **≈ 54** | |
| 316 | |
| 317 | JEV-27B reaches ≈ 0.017 (≈ 60 decisions per nat) on the same rows. |
| 318 | |
| 319 | The other test rows are not labelled by Jev and are not part of this claim: Open-Jev rows carry programmatic ground |
| 320 | truth (in-distribution KL 0.004 for `noul`, 0.176 for `choice`; 0.234 on the OOD split of unseen task families), and |
| 321 | the `yuri_v1` rows carry placeholder labels. No Jev-labelled out-of-distribution set exists in the corpus, so the claim |
| 322 | is established on the 53 training domains; outside them, the independent benchmark with human labels (90–97 % of Jev's |
| 323 | accuracy) is the best available evidence. |
| 324 | |
| 325 | **Fidelity includes the teacher's mistakes.** On a poker spot where a solver always checks, TypeSafe Jev 1.13 shoves |
| 326 | with 0.62 in a published test; JEV-9B shoves too, with 0.70 (JEV-27B 0.63). A faithful copy of System 1 is also a |
| 327 | faithful copy of its blind spots. At 9 B the student also adds some of its own: on an independent benchmark 11.5 % of |
| 328 | its 16-option answers change when only the option order changes, against 7.0 % for the teacher (JEV-27B 7.4 %). |
| 329 | |
| 330 | ## The Blocks of Experts recipe |
| 331 | |
| 332 | ``` |
| 333 | ┌── System 2 block: lm_head (248,320 × 4096) ───────► text generation and reasoning |
| 334 | Request ─► Router ─► Qwen3.5-9B backbone (frozen, bit-identical to the base) (adapter off; HumanEval 70.7 % = base) |
| 335 | per │ |
| 336 | request └── + System 1 block: LoRA (40.1 M) + 24-slot head (98 k) ─► calibrated typed decision |
| 337 | (adapter on, decision path only) (one prefill pass; KL ≈ 0.019 to Jev) |
| 338 | ``` |
| 339 | |
| 340 | | block | what it is | parameters | trained? | used for | |
| 341 | |---|---|---|---|---| |
| 342 | | Backbone | `Qwen/Qwen3.5-9B` text tower (vision tower and MTP head dropped), bf16 | 7.9 B | no — bit-identical to the base | both systems | |
| 343 | | **System 2 block** | the original `lm_head` (248,320 × 4096) | part of the base | no | text generation and step-by-step reasoning | |
| 344 | | **System 1 block** | LoRA r=16 on the decoder projections + a 24-slot fp32 decision head initialised from `lm_head` rows | 40.1 M + 98 k | yes, ≈ 3 B200-hours | calibrated typed decisions | |
| 345 | | Router | per request: the vLLM LoRA module `jev-decision`, or `peft` adapter on/off | — | — | chooses the system | |
| 346 | |
| 347 | **Why separate blocks rather than one merged fine-tune.** Folding the System 1 LoRA into the backbone would let a |
| 348 | single weight set serve both heads, but it costs generation quality: the merged backbone with the original `lm_head` |
| 349 | scores 61.6 % (101/164) on HumanEval against 70.7 % for the base, a 9-point drop, even though prose perplexity barely |
| 350 | moves (3.15 → 3.30). Keeping the backbone pristine and applying the System 1 block only on the decision path removes |
| 351 | that trade-off. For decision serving the adapter is merged *in memory* at start-up, so decision latency matches a |
| 352 | merged bundle. |
| 353 | |
| 354 | **Why the recipe is this efficient.** |
| 355 | |
| 356 | 1. **Pretraining does most of the work; distillation sharpens.** The decision head is initialised from the backbone's |
| 357 | own `lm_head` rows for the verbalizer tokens (`false/true`, `0`–`5`, `A`–`P`), so at step 0 its output equals the |
| 358 | pretrained model's zero-shot restricted next-token distribution (verified to |Δp| < 1e-5; measured 8.6e-07). Before |
| 359 | seeing a single label it already agrees with the test targets on 53 % of `choice` questions with `noul` AUROC 0.82; |
| 360 | distillation takes it to 90 % / 0.996. |
| 361 | 2. **Small trainable footprint.** 40.2 M parameters — 0.5 % of the backbone. Validation KL was already below 0.10 after |
| 362 | the first 64 k rows, test KL reached 0.028 after 0.49 epoch (≈ 1.7 B200-hours) and 0.021 after 0.93 epoch. |
| 363 | 3. **Transfer to unseen tasks.** The pretrained backbone reads the *content* of a new task instead of matching surface |
| 364 | patterns of the training domains: KL 0.234 and top-1 0.918 against the programmatic labels of the OOD split. It |
| 365 | also reads real, long, structured states (prose, JSON game states, policy documents; up to 856 tokens in the corpus). |
| 366 | 4. **Reads options, not positions.** With 30 % option-permutation augmentation, the top-1 flip rate under shuffled |
| 367 | `choice` options is 3.9 %; the same backbone before distillation flips 38 % of the time. |
| 368 | 5. **Calibration falls out of the objective.** Distilling full teacher distributions with KL (plus an ordinal RPS term |
| 369 | for `score`) gives fitted temperatures of 1.002 / 0.984 / 1.012 and ECE 0.0007 with no post-hoc correction. |
| 370 | 6. **It scales without code changes, and scale pays off.** The same code, hyper-parameters and packaging produced the |
| 371 | second-generation JEV-27B; only `model_path` and memory settings changed (the head-initialisation identity holds |
| 372 | there too, 4.5e-07). Going from 9 B to 27 B lowers KL to Jev from ≈ 0.019 to ≈ 0.017, halves OOD KL |
| 373 | (0.234 → 0.104), and raises the System 2 path from 70.7 % to 78.0 % on HumanEval. |
| 374 | |
| 375 | ## Benchmark highlights |
| 376 | |
| 377 | ### Independent benchmark: side by side with TypeSafe Jev 1.13 |
| 378 | |
| 379 | [`gazelle93/decision-models-under-pressure`](https://github.com/gazelle93/decision-models-under-pressure) (published |
| 380 | 25 Sep 2026) asks decision models to pick the right label for real texts from CLINC-150, MTOP, GoEmotions, DBpedia and |
| 381 | financial tweets under three kinds of pressure: more options, near-miss options, and shuffled option order. The labels |
| 382 | are human gold labels, none of this data is in our training set, and TypeSafe Jev 1.13's results are published with it. |
| 383 | We re-ran the same items with autotrust/JEV-9B and autotrust/JEV-27B, up to our 16-option limit. |
| 384 | |
| 385 | | | TypeSafe Jev 1.13 (published) | **autotrust/JEV-9B** | autotrust/JEV-27B | |
| 386 | |---|---|---|---| |
| 387 | | Accuracy with 2 / 4 / 8 / 16 options (800 items, 4 domains) | 0.890 / 0.801 / 0.782 / 0.769 | **0.868 / 0.774 / 0.735 / 0.694** | 0.876 / 0.784 / 0.767 / 0.740 | |
| 388 | | 16 options — CLINC / DBpedia / GoEmotions / MTOP | 0.945 / 0.900 / 0.470 / 0.760 | **0.875 / 0.855 / 0.325 / 0.720** | 0.930 / 0.885 / 0.415 / 0.730 | |
| 389 | | 16 options, near-miss vs. unrelated wrong options (CLINC + MTOP, 400 items) | 0.912 vs 0.985 | **0.875 vs 0.975** | 0.907 vs 0.983 | |
| 390 | | Answers changed by shuffling the options alone (16 options, 5 orderings) | 7.0 % | **11.5 %** | 7.4 % | |
| 391 | | Time for 14,400 decisions on one B200 | — | **42 s** | 110 s | |
| 392 | |
| 393 | On data it was never trained on, JEV-9B reaches 97 % of the teacher's accuracy with 2 and 4 options, 94 % with 8 and |
| 394 | 90 % with 16: it falls behind faster than JEV-27B (96–98 %) as the option list grows, loses a little more on near-miss |
| 395 | options, and is more sensitive to option order than the teacher. Our run follows the benchmark's published method (gold |
| 396 | plus the first K−1 distractors of a pool, shuffled per item); the orderings are seeded differently, so compare |
| 397 | aggregates, not individual items. |
| 398 | |
| 399 | ### Fresh examples (Hacker News and V2EX, 23–25 September 2026) |
| 400 | |
| 401 | Expected answers were written by hand before the model was run. These are illustrations (≈ 110 decisions), not a |
| 402 | benchmark. |
| 403 | |
| 404 | | task | autotrust/JEV-9B | autotrust/JEV-27B | |
| 405 | |---|---|---| |
| 406 | | Topic of 19 HN front-page stories (10 options) + "is it about AI?" | **38 / 38** | 38 / 38 | |
| 407 | | 12 comments from a heated HN thread: "insults or attacks someone?" + "what is it mainly doing?" (6 options) | **22 / 24** | 23 / 24 | |
| 408 | | 10 V2EX hot posts in **Chinese**: "contains a referral / invite code?" + "promotes a product or paid offer?" | **18 / 19** | 19 / 19 | |
| 409 | | Community use cases: code-rule checks in the style of `adhere`, injection filtering, ticket routing, phishing, code-review diffs, urgency scores | **14 / 15** | 15 / 15 | |
| 410 | |
| 411 | | input | question | autotrust/JEV-9B | |
| 412 | |---|---|---| |
| 413 | | HN: "Two-tier encryption in the UK" | topic (10 options) | security & privacy · 0.87 | |
| 414 | | HN: "Using LLMs to trace alchemical knowledge and decode 17th century letters" | about AI? | P(true) = 0.88 | |
| 415 | | HN comment: "Please stop this. We've asked you before to observe the guidelines…" | what is it mainly doing? | moderating the discussion · 0.78 | |
| 416 | | V2EX: "一个不需要 gemini pro 的完全免费的注册 Muse 的方法 … 邀请码:…" | contains a referral / invite code? | P(true) = 1.00 | |
| 417 | | V2EX: "今天中秋节,还要加班的有吗?来报道下" | promotes a product or paid offer? | P(true) = 0.01 | |
| 418 | | Diff replacing a parameterised query with `"… WHERE id = " + request.args["id"]` | introduces a security vulnerability? | P(true) = 0.93 (0.17 for a variable rename) | |
| 419 | | "I'm not happy with the fit. What are my options here?" | asking for a refund? | P(true) = 0.17 (TypeSafe's docs report 0.22 for Jev on this exact text) | |
| 420 | |
| 421 | Where it failed or wavered: |
| 422 | |
| 423 | * **Code-rule check**: missed a TypeScript file that declares `const port: number = Number(process.env.PORT)` against |
| 424 | the rule "a port must be a branded, range-checked integer" (0.33); JEV-27B flags it (0.93). |
| 425 | * **Fraud screening**: a CEO wire-transfer (business-email-compromise) message is flagged, but only at 0.56 (JEV-27B 0.84). |
| 426 | * **Comment intent**: "Because it's not a real argument. It's a deflection people use." read as attacking another |
| 427 | commenter (0.66) rather than arguing a point. |
| 428 | * **Chinese promotion**: a V2EX post launching a paid HTTPS debugging tool was not flagged as promotional (0.44; |
| 429 | JEV-27B 0.77). |
| 430 | * **A poker spot with the nuts** (check or shove four times the pot; a solver checks 100 %): shoves with 0.70; the |
| 431 | teacher shoved with 0.62, so this mistake comes from the teacher. |
| 432 | * Counting ("more than 3 fruits?" / "more than 5?" for a list of 4: 0.84 / 0.37), date comparisons and an instruction |
| 433 | injected inside the state were handled correctly, but on a handful of examples only. |
| 434 | |
| 435 | Per-example outputs and the benchmark aggregates are in `reports/realworld_9b.json` (the HN and V2EX inputs came from |
| 436 | their public APIs on 25 September 2026). |
| 437 | |
| 438 | ## Quickstart with vLLM (recommended) |
| 439 | |
| 440 | **One vLLM engine serves both systems from the same pristine weights.** Ordinary requests go through the base |
| 441 | `lm_head` (System 2, exactly Qwen3.5-9B); requests addressed to the LoRA module `jev-decision` go through the decision |
| 442 | head (System 1). `adapter_vllm/` contains the backbone LoRA plus the 24-slot decision head re-expressed as an `lm_head` |
| 443 | LoRA (only the 24 verbalizer rows change), so a typed decision is a single prefill step with `max_tokens=1`, |
| 444 | constrained to the option tokens and read back as log-probabilities. |
| 445 | |
| 446 | ### 1 — Start the server (OpenAI-compatible) |
| 447 | |
| 448 | ```bash |
| 449 | hf download autotrust/JEV-9B --local-dir JEV-9B # ~18 GB |
| 450 | vllm serve JEV-9B --served-model-name autotrust/JEV-9B \ |
| 451 | --enable-lora --max-lora-rank 32 --lora-modules jev-decision=JEV-9B/adapter_vllm \ |
| 452 | --logprobs-mode processed_logprobs --max-model-len 4096 |
| 453 | ``` |
| 454 | |
| 455 | `--logprobs-mode processed_logprobs` is required: it makes the returned log-probabilities respect `allowed_token_ids`. |
| 456 | `--max-model-len 4096` is sized for decisions; raise it (for example to 16384) if System 2 requests will think at |
| 457 | length. Add `--enable-prefix-caching --mamba-cache-mode align` if you ask many questions about the same state. |
| 458 | |
| 459 | ### 2 — System 2: generation and reasoning (the unmodified base model) |
| 460 | |
| 461 | ```bash |
| 462 | curl -s localhost:8000/v1/chat/completions -H 'content-type: application/json' -d '{ |
| 463 | "model": "autotrust/JEV-9B", |
| 464 | "messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}], |
| 465 | "max_tokens": 60, "chat_template_kwargs": {"enable_thinking": false}}' |
| 466 | ``` |
| 467 | |
| 468 | Set `"enable_thinking": true` for deliberate, step-by-step reasoning. This path is Qwen3.5-9B unchanged; see the |
| 469 | [Qwen3.5-9B model card](https://huggingface.co/Qwen/Qwen3.5-9B) for its reasoning benchmarks and recommended sampling |
| 470 | settings. |
| 471 | |
| 472 | ### 3 — System 1: typed decisions (Python, only `requests` + two small JSON files) |
| 473 | |
| 474 | ```python |
| 475 | import json, math, requests |
| 476 | from huggingface_hub import hf_hub_download |
| 477 | |
| 478 | REPO, URL = "autotrust/JEV-9B", "http://localhost:8000" |
| 479 | dh = json.load(open(hf_hub_download(REPO, "adapter_vllm/decision_head.json"))) # bias + verbalizer token ids |
| 480 | T = json.load(open(hf_hub_download(REPO, "calibration.json")))["per_kind"] # per-kind temperatures |
| 481 | |
| 482 | def decide(kind, state, question, options=None): |
| 483 | options = {"noul": ["false", "true"], "score": [str(i) for i in range(6)]}.get(kind, options) |
| 484 | lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)] |
| 485 | prompt = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:" |
| 486 | s = dh["slots"]["ranges"][kind][0] |
| 487 | ids = dh["verbalizer_ids"][s : s + len(options)] # the option tokens of this kind |
| 488 | r = requests.post(f"{URL}/v1/completions", json={ |
| 489 | "model": "jev-decision", "prompt": prompt, "max_tokens": 1, "temperature": 1.0, |
| 490 | "logprobs": len(options), "allowed_token_ids": ids, |
| 491 | "add_special_tokens": False, "return_tokens_as_token_ids": True}).json() |
| 492 | lp = {int(k.split(":")[1]): v for k, v in r["choices"][0]["logprobs"]["top_logprobs"][0].items()} |
| 493 | z = [(lp.get(t, -1e9) + dh["bias"][s + i]) / T[kind] for i, t in enumerate(ids)] # + head bias, / temperature |
| 494 | e = [math.exp(x - max(z)) for x in z] |
| 495 | return {o: x / sum(e) for o, x in zip(options, e)} |
| 496 | |
| 497 | print(decide("choice", "SKU AX-330 stock at 8% of safety level; supplier late twice this quarter.", |
| 498 | "Supplier response for this scenario.", ["issue_warning", "renegotiate", "dual_source", "maintain"])) |
| 499 | # ≈ {'issue_warning': 0.33, 'renegotiate': 0.14, 'dual_source': 0.53, 'maintain': 0.001} |
| 500 | print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.", |
| 501 | "Is the customer asking for a refund?")) |
| 502 | ``` |
| 503 | |
| 504 | Values can differ in the third decimal between runs: vLLM computes in bf16 and results depend slightly on which |
| 505 | requests are batched together. The head bias and the temperature are applied client-side; the log-softmax normaliser |
| 506 | that vLLM applies cancels out, so the result is exactly the decision head's calibrated distribution. |
| 507 | |
| 508 | ### 4 — System 1 → System 2: confidence-gated escalation |
| 509 | |
| 510 | Because both systems live in one engine, a common pattern is to let System 1 answer when it is confident and hand the |
| 511 | rest to System 2. This is a usage pattern, not a configuration we have benchmarked; pick the threshold on your own |
| 512 | validation data, and serve with a `--max-model-len` large enough for the reasoning budget. |
| 513 | |
| 514 | ```python |
| 515 | def solve(state, question, options, threshold=0.90): |
| 516 | p = decide("choice", state, question, options) # System 1: one prefill pass |
| 517 | best = max(p, key=p.get) |
| 518 | if p[best] >= threshold: |
| 519 | return {"system": 1, "answer": best, "distribution": p} |
| 520 | prompt = (f"{state}\n\nQuestion: {question}\nOptions: " + "; ".join(options) |
| 521 | + "\nThink it through, then give exactly one option on the last line.") |
| 522 | r = requests.post(f"{URL}/v1/chat/completions", json={ # System 2: same engine, base lm_head |
| 523 | "model": "autotrust/JEV-9B", |
| 524 | "messages": [{"role": "user", "content": prompt}], |
| 525 | "max_tokens": 8192, "chat_template_kwargs": {"enable_thinking": True}}).json() |
| 526 | return {"system": 2, "reply": r["choices"][0]["message"]["content"], "system1_distribution": p} |
| 527 | ``` |
| 528 | |
| 529 | ### Offline / batch (Python API) |
| 530 | |
| 531 | ```python |
| 532 | from vllm import LLM, SamplingParams |
| 533 | from vllm.lora.request import LoRARequest |
| 534 | |
| 535 | llm = LLM("JEV-9B", enable_lora=True, max_lora_rank=32, logprobs_mode="processed_logprobs", max_model_len=4096) |
| 536 | decision = LoRARequest("jev-decision", 1, "JEV-9B/adapter_vllm") |
| 537 | |
| 538 | gen = llm.generate(["..."], SamplingParams(temperature=0.0, max_tokens=256)) # System 2, no LoRA |
| 539 | dec = llm.generate([prompt], [SamplingParams(max_tokens=1, temperature=1.0, # System 1 |
| 540 | allowed_token_ids=ids, logprobs=len(ids))], |
| 541 | lora_request=decision) # then + bias, / T as above |
| 542 | ``` |
| 543 | |
| 544 | Mixed batches work too: pass a per-request `lora_request` list (`None` for System 2, `decision` for System 1) and both |
| 545 | systems are served in the same `generate` call. |
| 546 | |
| 547 | ### Measured on one B200 |
| 548 | |
| 549 | | workload | PyTorch path | **vLLM** | |
| 550 | |---|---|---| |
| 551 | | System 2 — 164 HumanEval completions (greedy, ≤ 384 new tokens) | 165 s | **3.3 s** (≈ 50×) | |
| 552 | | System 1 — offline batch, 29,955 test questions | 75 s (398 q/s) | 80 s (374 q/s) | |
| 553 | | System 1 over HTTP — 64 / 256 concurrent clients | — | 150 / 205 req/s | |
| 554 | | System 1 fidelity vs. the PyTorch path | test KL 0.0210 | test KL 0.0211; mean \|Δp\| 0.0008 over HTTP | |
| 555 | | Many questions about one state, `--enable-prefix-caching` | — | +14–20 % throughput | |
| 556 | |
| 557 | Notes: |
| 558 | * The big win is on System 2: generation is ≈ 50× faster. A decision is a single prefill pass with no decoding, so at |
| 559 | 9 B offline batch throughput is about the same as the PyTorch path (at 27 B vLLM is 1.7× faster); for System 1 vLLM |
| 560 | mainly buys serving: continuous batching under concurrency, an OpenAI-compatible API, and one engine for both systems. |
| 561 | * Prefix caching: this architecture mixes Gated DeltaNet and attention layers, and vLLM caches it in blocks of 528 |
| 562 | tokens, so only shared prefixes longer than 528 tokens are reused. The template puts `[kind]` before `[state]`, so |
| 563 | only questions of the same kind share a prefix. On 293 real states × 7.7 yes/no questions each (≈ 480-token states), |
| 564 | prefix caching served 19.6 % of prompt tokens from cache (+14–20 % throughput) with identical outputs. |
| 565 | * Requires a vLLM build with Qwen3.5 (`qwen3_5`) support, LoRA on `lm_head`, `--logprobs-mode` and |
| 566 | `allowed_token_ids`; tested with a vLLM development build from September 2026. Start-up takes 3–8 minutes |
| 567 | (CUDA-graph capture with LoRA enabled). |
| 568 | |
| 569 | ## What System 1 does |
| 570 | |
| 571 | | kind | question | returns | |
| 572 | |---|---|---| |
| 573 | | `noul` | "Is this statement true?" | `[P(false), P(true)]` | |
| 574 | | `choice` | "Which of these 2–16 options?" | one probability per option, aligned with your `options` | |
| 575 | | `score` | "Where on this ordered 0–5 scale?" | a distribution over the six levels (+ expected score) | |
| 576 | |
| 577 | ``` |
| 578 | [kind] choice |
| 579 | [state] SKU AX-330 stock at 8% of safety level; supplier late twice this quarter. |
| 580 | [question] Supplier response for this scenario. |
| 581 | [options] |
| 582 | A) issue_warning |
| 583 | B) renegotiate |
| 584 | C) dual_source |
| 585 | D) maintain |
| 586 | [decision]: |
| 587 | ``` |
| 588 | |
| 589 | The template is tokenised as one string; the last token's final-norm hidden state goes through a |
| 590 | **linear fp32 head `H → 24 slots`** (`noul` → slots 0–1, `score` → 2–7, `choice` → 8–23). Inactive slots are masked, |
| 591 | a per-kind temperature is applied, and a softmax yields the distribution aligned with your `options`. One prefill |
| 592 | pass, no decoding. |
| 593 | |
| 594 | ## Other ways to run it |
| 595 | |
| 596 | ### Plain `transformers` + `peft` |
| 597 | |
| 598 | ```python |
| 599 | import json, torch |
| 600 | from huggingface_hub import hf_hub_download |
| 601 | from peft import PeftModel |
| 602 | from safetensors.torch import load_file |
| 603 | from transformers import AutoModelForCausalLM, AutoTokenizer |
| 604 | |
| 605 | repo = "autotrust/JEV-9B" |
| 606 | tok = AutoTokenizer.from_pretrained(repo) |
| 607 | base = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda") # == Qwen3.5-9B text model |
| 608 | |
| 609 | # --- System 2: the pristine base model, no adapter ---------------------------------------------- |
| 610 | msgs = [{"role": "user", "content": "In two sentences, what is safety stock?"}] |
| 611 | enc = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda") |
| 612 | out = base.generate(**enc, max_new_tokens=80) |
| 613 | print(tok.decode(out[0, enc["input_ids"].shape[1]:], skip_special_tokens=True)) |
| 614 | |
| 615 | # --- System 1: attach the LoRA adapter (merged here for speed) + the 24-slot head --------------- |
| 616 | model = PeftModel.from_pretrained(base, repo, subfolder="adapter").merge_and_unload() |
| 617 | head = load_file(hf_hub_download(repo, "head.safetensors")) |
| 618 | cfg = json.load(open(hf_hub_download(repo, "judge_config.json"))) |
| 619 | temp = json.load(open(hf_hub_download(repo, "calibration.json")))["per_kind"] |
| 620 | W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda() |
| 621 | |
| 622 | def decide(kind, state, question, options): |
| 623 | letters = "ABCDEFGHIJKLMNOP" |
| 624 | lines = options if kind != "choice" else [f"{letters[i]}) {o}" for i, o in enumerate(options)] |
| 625 | text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:" |
| 626 | ids = tok(text, return_tensors="pt", add_special_tokens=False).to("cuda") |
| 627 | with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16): |
| 628 | h = model.model(**ids).last_hidden_state[0, -1].float() # backbone only, last token |
| 629 | z = (W @ h + b) / temp[kind] |
| 630 | s, _ = cfg["slots"]["ranges"][kind] |
| 631 | p = torch.softmax(z[s : s + len(options)], 0) |
| 632 | return dict(zip(options, p.tolist())) |
| 633 | |
| 634 | print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.", |
| 635 | "Is the customer asking for a refund?", ["false", "true"])) |
| 636 | # {'false': 0.009, 'true': 0.991} |
| 637 | ``` |
| 638 | |
| 639 | `options` are validated: `noul` must be `["false","true"]`, `score` must be `["0".."5"]`, `choice` takes 2–16 |
| 640 | free-text options. Note that `merge_and_unload()` above changes the backbone for the rest of the process; to keep both |
| 641 | systems in one process, leave the adapter unmerged and run System 2 inside `with model.disable_adapter():`. |
| 642 | |
| 643 | ## Evaluation details |
| 644 | |
| 645 | ### Additional System 1 metrics (`test_set_30k`, temperature applied) |
| 646 | |
| 647 | | metric | autotrust/JEV-9B | |
| 648 | |---|---| |
| 649 | | `noul` Brier score against the target probability, all rows (lower is better) | 0.0015 | |
| 650 | | `score` ranked probability score (lower is better) | 0.0085 | |
| 651 | | Fitted temperatures noul / choice / score | 1.002 / 0.984 / 1.012 | |
| 652 | | Top-1 flip rate when `choice` options are shuffled (1,000 rows × 4 permutations) | 3.9 % | |
| 653 | | Out-of-distribution split — top-1 agreement · `noul` AUROC | 0.918 · 0.989 | |
| 654 | | Throughput — batch of 128 requests on one B200 | 2.5 ms per decision (≈ 400 decisions/s) | |
| 655 | | Single request on one B200 (median) | 87 ms | |
| 656 | |
| 657 | ### System 2 — no-degradation check (HumanEval, greedy pass@1, completion-style prompt) |
| 658 | |
| 659 | | weights | pass@1 | note | |
| 660 | |---|---|---| |
| 661 | | Qwen3.5-9B (base) | 70.7 % (116/164) | same loader and protocol as below | |
| 662 | | **autotrust/JEV-9B — System 2 path (backbone + `lm_head`, adapter off)** | **70.7 % (116/164)** | all 164 completions byte-identical to the base model | |
| 663 | | System 1 LoRA folded into the backbone + base `lm_head` (*not shipped*) | 61.6 % (101/164) | why the blocks are kept separate | |
| 664 | |
| 665 | ### Per source × primitive (`test_set_30k`, temperature applied) |
| 666 | |
| 667 | | source | kind | n | KL | top-1 | ECE | noul AUROC | score MAE | |
| 668 | |---|---|---|---|---|---|---|---| |
| 669 | | yuri_v3 — synthetic operational scenarios, labelled by TypeSafe Jev 1.13 | noul | 8,537 | 0.005 | 0.961 | 0.001 | 0.994 | — | |
| 670 | | yuri_v3 | choice | 8,312 | 0.028 | 0.902 | 0.002 | — | — | |
| 671 | | yuri_v3 | score | 8,527 | 0.023 | 0.883 | 0.002 | — | 0.103 | |
| 672 | | openjev_v2 — Open-Jev programmatic tasks, ground-truth labels (not Jev) | noul | 1,432 | 0.004 | 0.998 | 0.003 | 1.000 | — | |
| 673 | | openjev_v2 | choice | 887 | 0.176 | 0.857 | 0.020 | — | — | |
| 674 | | yuri_v1 — placeholder `[0.5, 0.5]` labels (see Limitations) | noul | 2,260 | 0.000 | — | 0.005 | — | — | |
| 675 | |
| 676 | OOD split (13,058 Open-Jev rows from task families not in training, programmatic labels): KL 0.234, top-1 0.918, noul |
| 677 | AUROC 0.989; `choice` KL 0.351 / top-1 0.837 (game-state decisions are the hardest slice). |
| 678 | |
| 679 | Choice option-permutation consistency (1,000 rows × 4 random permutations): mean max |Δp| 0.024, p90 0.055, top-1 |
| 680 | flip rate 3.9 %. |
| 681 | |
| 682 | ### Training trajectory (most recent first) |
| 683 | |
| 684 | Validation KL on a fixed 4 k-row subset; `test_set_30k` metrics after calibration. |
| 685 | |
| 686 | | stage | rows seen | val KL | t30k KL | choice top-1 | score MAE | noul AUROC | ECE | |
| 687 | |---|---|---|---|---|---|---|---| |
| 688 | | **autotrust/JEV-9B v0.8.0 — released weights (4,750 steps ≈ 0.93 epoch, LR annealed to ≈ 0.07×)** | 608 k | **0.019** | **0.0210** | **0.898** | **0.103** | **0.996** | **0.0007** | |
| 689 | | v0.7.0 (step 2000 + 500-step LR cool-down) | 320 k | 0.026 | 0.0276 | 0.884 | 0.119 | 0.994 | 0.0014 | |
| 690 | | step 2000 | 256 k | 0.0325 | 0.037 | 0.865 | 0.143 | 0.992 | 0.004 | |
| 691 | | step 1500 | 192 k | 0.038 | 0.040 | 0.861 | 0.151 | 0.991 | 0.0025 | |
| 692 | | step 500 | 64 k | 0.094 | 0.081 | 0.817 | 0.224 | 0.977 | 0.022 | |
| 693 | | untrained backbone with the initialised head (reference point, not the model) | 0 | 0.485 | 0.510 | 0.532 | 1.130 | 0.824 | 0.094 | |
| 694 | |
| 695 | Annealing matters: the v0.7.0 cool-down (500 steps, lr ×0.9 → ×0.02 from step 2000) lowered KL by 25 %; continuing on |
| 696 | the unseen remainder of the epoch with the learning rate decayed to ≈ 0.07× (v0.8.0) lowered it by another 24 % and |
| 697 | added 1.4 points of choice agreement. JEV-27B folds this into a single cosine schedule. |
| 698 | |
| 699 | ## Training details |
| 700 | |
| 701 | | item | value | |
| 702 | |---|---| |
| 703 | | teacher / data | `SargeDev/jev-distill-corpus-v3` (740,957 rows; `train` 655,806) with three streams: `yuri_v3` (498,010 rows, TypeSafe Jev 1.13 full output distributions via OpenRouter), `openjev_v2` (94,801 rows, Open-Jev programmatic labels, CC0), `yuri_v1` (148,154 rows, placeholder labels, down-weighted) | |
| 704 | | backbone | `Qwen/Qwen3.5-9B` text tower only (vision tower and MTP head dropped), bf16, frozen | |
| 705 | | System 1 block (trainable) | LoRA r=16, α=32, dropout 0.05 on `in_proj_qkv, in_proj_z, out_proj, q/k/v/o_proj, gate/up/down_proj` (40.1 M, shipped unmerged in `adapter/`) + 24-slot head (98 k, fp32, initialised from `lm_head` rows) | |
| 706 | | System 2 block | the original `lm_head`, not trained | |
| 707 | | loss | KL(target ‖ model) over active slots + 0.5 · RPS (ranked probability score) for `score` | |
| 708 | | augmentation | 30 % random permutation of `choice` options (targets permuted consistently) | |
| 709 | | batching | 128 rows / step, kind-stratified (≥ 1/6 per primitive), length-bucketed, micro-batches capped at 24 k padded tokens, gradient checkpointing | |
| 710 | | optimiser | AdamW (fused), β=(0.9, 0.98), lr head 2e-4 / LoRA 1e-4, cosine, warmup 3 %, grad-clip 1.0 for 2,500 steps; then continued on the unseen remainder of the epoch (fresh AdamW state, warmup 2 %, lr ×0.9 → cosine) and stopped after 2,250 more steps at lr ≈ ×0.07 — 4,750 steps ≈ 0.93 epoch in total | |
| 711 | | label hygiene | `yuri_v1` rows carry exact-uniform `[0.5, 0.5]` placeholder labels (137,203 rows, 100 %); down-weighted ×0.05 in training and excluded from temperature fitting | |
| 712 | | calibration | per-kind scalar temperature (L-BFGS on the `calibration` split, 10,954 rows): noul 1.002 · choice 0.984 · score 1.012 | |
| 713 | | compute | 1× NVIDIA B200 (183 GB); ≈ 1.4 h (2,500 steps) + ≈ 1.5 h (2,250 steps) ≈ 3 GPU-hours; ≈ 7–9 k tokens/s | |
| 714 | | software | torch 2.13 + cu130, transformers 5.16, peft 0.21, flash-linear-attention 0.5.2 | |
| 715 | |
| 716 | ## Limitations |
| 717 | |
| 718 | * **System 1 mirrors TypeSafe Jev 1.13, including its mistakes.** This is a distillation, not an independent judge: |
| 719 | where the teacher was wrong or uncalibrated, so is autotrust/JEV-9B. Published evaluations of the teacher show it is |
| 720 | unreliable for multi-hop reasoning, arithmetic, dates, counting and adversarial inputs, and the student inherits all |
| 721 | of that. Confirmed on fresh inputs: the poker shove (0.70 vs the teacher's 0.62). |
| 722 | * **At 9 B it adds some blind spots of its own.** 11.5 % of its 16-option answers change with option order alone |
| 723 | (teacher 7.0 %), it keeps 90 % rather than 96 % of the teacher's accuracy at 16 options, and it missed a code-rule |
| 724 | violation that JEV-27B catches. Prefer [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B) for long option |
| 725 | lists, unfamiliar task families and code-rule checks. |
| 726 | * **The two systems are integrated in serving, not in knowledge.** System 1 cannot explain its decisions, and System 2 |
| 727 | is the unmodified base model: it knows nothing about the decisions it is packaged with and was not trained to agree |
| 728 | with System 1. If you escalate from System 1 to System 2, expect them to disagree sometimes. |
| 729 | * **"Indistinguishable" is a KL statement on Jev-labelled rows from the 53 training domains.** The corpus has no |
| 730 | Jev-labelled out-of-distribution set; the OOD figures (KL 0.234, 0.351 for game-state `choice`) are measured against |
| 731 | programmatic ground truth, and on the independent benchmark the student reaches 90–97 % of Jev's accuracy, not 100 %. |
| 732 | * **Speed comparisons with the hosted API are not like for like.** Our timings exclude network time; the hosted |
| 733 | figures are third-party measurements that include it and depend on client concurrency and rate limits. |
| 734 | * **Choice agreement is capped by teacher ambiguity.** The teacher's `choice` distributions are soft (median top-1 |
| 735 | probability 0.70). On the 14 % of rows where the teacher's top two options are within 0.1 of each other, argmax |
| 736 | agreement is near chance for *any* faithful mimic (0.46 where the gap is < 0.05). On teacher-decisive rows agreement |
| 737 | is 0.954, and the student's argmax captures 97.7 % of the teacher probability mass a perfect mimic could (0.693 vs |
| 738 | 0.709). |
| 739 | * **Fixed option sets.** `noul` and `score` accept only their canonical options; `choice` accepts 2–16 options. Inputs |
| 740 | longer than 1,024 tokens are truncated (state only, head 60 % / tail 40 %) at serving unless you raise the limit. |
| 741 | * **English-centric.** The corpus is English; multilingual behaviour is inherited from the backbone and was not |
| 742 | systematically measured (the Chinese V2EX examples above are illustrations only). |
| 743 | * **Placeholder labels in the corpus.** The `yuri_v1` memory-relevance stream is 100 % exact-uniform `[0.5, 0.5]` — |
| 744 | those rows teach nothing about relevance. The model outputs ≈ 0.5 on them by design; do not use it for |
| 745 | memory-relevance scoring without further training. |
| 746 | * **Not for high-stakes decisions.** Use confidence gating: act automatically only above a threshold you validated on |
| 747 | your own data, and route the rest to System 2, a stronger model, or a human. |
| 748 | |
| 749 | ## Files |
| 750 | |
| 751 | ``` |
| 752 | model-0000{1..5}-of-00005.safetensors Qwen3.5-9B text backbone incl. lm_head — bit-identical to the base model |
| 753 | (bf16; GDN A_log / gated-norm weights fp32 as in the original), 17.9 GB |
| 754 | model.safetensors.index.json · config.json |
| 755 | adapter/ System 1 LoRA (peft format, r=16, 40.1 M params, 154 MB) — apply only for decisions |
| 756 | head.safetensors 24-slot decision head (fp32): proj.weight [24, 4096], proj.bias [24] |
| 757 | judge_config.json slot layout, verbalizer token ids, template version, weights_mode=unmerged, provenance |
| 758 | calibration.json per-kind temperatures (+ fit diagnostics) |
| 759 | adapter_vllm/ the same adapter for vLLM: backbone LoRA (zero-padded to r=32) + decision head as an |
| 760 | lm_head LoRA, plus decision_head.json (head bias, verbalizer token ids) |
| 761 | tokenizer.json · tokenizer_config.json · chat_template.jinja |
| 762 | reports/ evaluation reports: test-set evaluation, bundle checks, HumanEval per-problem |
| 763 | results, vLLM measurements, real-world tests, training-milestone reviews |
| 764 | vl/ vision: serve.sh (multimodal Qwen3.5-9B + System 1), serve_decide.py (POST /v1/decide), |
| 765 | adapter_vllm/ (layer names for the multimodal model), calibration.json, demos/ |
| 766 | videos/ robot-arm and computer-use demo videos |
| 767 | reports/vl/ image evaluations and demo results |
| 768 | ``` |
| 769 | |
| 770 | ## License and acknowledgements |
| 771 | |
| 772 | Weights: **Apache-2.0** (base model `Qwen/Qwen3.5-9B` is Apache-2.0; training corpus |
| 773 | `SargeDev/jev-distill-corpus-v3` is Apache-2.0, its `openjev_v2` stream additionally CC0). The System One framing and |
| 774 | the `noul` / `choice` / `score` primitives originate with TypeSafe AI's Jev; autotrust/JEV-9B is an independent |
| 775 | student model trained on public data and shares no weights, code or affiliation with TypeSafe AI. |
| 776 | |
| 777 | ```bibtex |
| 778 | @misc{autotrust_jev9b_2026, |
| 779 | title = {autotrust/JEV-9B: the first integrated System 1 + System 2 open model built with the Blocks of Experts recipe (Qwen3.5-9B; System 1 distilled from TypeSafe Jev 1.13)}, |
| 780 | author = {{AutoTrust AI}}, |
| 781 | year = {2026}, |
| 782 | url = {https://huggingface.co/autotrust/JEV-9B} |
| 783 | } |
| 784 | ``` |
| 785 | |