README.md
11.9 KB · 240 lines · markdown Raw
1 ---
2 base_model: Qwen/Qwen3.8-27B
3 base_model_relation: quantized
4 pipeline_tag: text-generation
5 tags:
6 - mtp
7 - multi-token-prediction
8 - speculative-decoding
9 - vision
10 - image
11 - multimodal
12 - image-text-to-text
13 - text-generation-inference
14 - qwen
15 - qwen3.8
16 - qwen38
17 - 27b
18 - pi
19 - coding
20 - coder
21 - code-generation
22 - agent
23 - tool-use
24 - tool-calling
25 - function-calling
26 - reasoning
27 - conversational
28 - sft
29 - grpo
30 - reinforcement-learning
31 - gguf
32 - llama.cpp
33 - quantized
34 - imatrix
35 - iquant
36 license: apache-2.0
37 library_name: gguf
38 ---
39
40 <img src="assets/benchmark-hero-gpqa-v2.png" alt="Qwen3.8-27B-pi and Base: Terminal-Bench 2.1 xhigh preview and Pi GPQA Diamond xhigh attempt pass rate; SciCode Pi/S8 subproblem scores" width="100%" />
41
42 <p align="center">
43 <a href="https://huggingface.co/bytkim/Qwen3.8-27B-pi">BF16</a> ·
44 <a href="https://huggingface.co/bytkim/Qwen3.8-27B-pi-FP8">FP8</a> ·
45 <strong>GGUF</strong>
46 </p>
47
48 <p align="center">
49 <a href="#quickstart">Quickstart</a> ·
50 <a href="#gguf-performance--release-features">Benchmarks</a> ·
51 <a href="#advanced">Advanced</a> ·
52 <a href="#license">License</a>
53 </p>
54
55
56 ## Built for pi
57
58 **Qwen3.8-27B-pi** builds on [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) for coding work in the Pi agent harness. It is tailored to the loop of reading a repository, editing files, running tools, and responding to their feedback. The aim is more completed work with less generated text, while retaining the base model's familiar interface.
59
60 Fine-tuned for the Pi agent harness, the model is designed to work through coding tasks as an iterative process: inspect the code, make a change, check the result, and use tool feedback to guide the next step. Adjustable reasoning effort lets you balance responsiveness with deeper problem-solving, from focused edits to more involved debugging and implementation work.
61
62 <img src="assets/benchmark-by-reasoning-gpqa-v1.png" alt="Qwen3.8-27B-pi versus Base across low, medium, and xhigh reasoning: Terminal-Bench 2.1 score, turns, tool calls, reasoning tokens, and total output tokens; Pi GPQA Diamond overall attempt pass rates; SciCode Pi/S8 subproblem scores" width="100%" />
63
64 ## Qwen3.8-27B-pi Highlights
65
66 - **Built for Pi:** Fine-tuned for the everyday coding loop—reading repositories, editing files, running tools, and working through feedback in the Pi harness—with an emphasis on turning plans into working, checked implementations.
67 - **Curated coding experience:** Supervised fine-tuning on filtered, successful Pi sessions helps the model learn complete coding workflows, rather than isolated answers or code snippets, including how to adapt to existing environments and check results against task requirements.
68 - **Task quality matters:** Development also included reviewing, repairing, and exploring more demanding coding tasks—with a focus on clear requirements and checks that distinguish working solutions from incorrect ones.
69 - **Refined through reinforcement learning:** A second training stage pairs verified task outcomes with a custom reasoning-efficiency reward built on GRPO, encouraging successful low- and medium-effort solutions to reason more economically while leaving xhigh focused on correctness.
70 - **Available in practical formats:** Choose a deployment format that fits your hardware, including GGUF quantizations calibrated using complete Pi coding sessions.
71
72 Pi’s development followed a staged process, from trajectory curation and supervised fine-tuning to reinforcement learning and checkpoint evaluation. Each stage was assessed against the same practical goal: helping the model complete useful work in Pi while managing the resources it spends. Checkpoints were compared using actual agent outcomes alongside generated tokens, tool calls, and completion time—not training loss alone.
73
74 ## GGUF Performance & Release Features
75
76 <img src="assets/gguf-quantization-benchmark-comparison.png" alt="GGUF quantization benchmark comparison at medium reasoning, with Pi FP8 reference results for Terminal-Bench 2.1, GPQA Diamond, and SciCode" width="100%" />
77
78 Choose from a range of GGUF sizes to fit your hardware, including compact IQ (“importance-aware”) options designed to preserve quality at lower memory use. Calibration on complete Pi coding sessions helps guide compression toward the parts of the model most important to those workflows.
79
80 ## Performance by reasoning level
81
82 <img src="assets/terminal-bench-score-tokens.png" alt="Terminal-Bench 2.1 task success versus mean output tokens for Pi and Base at low, medium, and xhigh reasoning" width="100%" />
83
84 > Curated Pi sessions and reinforcement learning emphasized completed, checked coding work. Pi shows a steadier rise in completion from low to xhigh than Base, with fewer output tokens at every matched effort level. In these selected results, Pi’s medium setting matches Base’s xhigh completion rate with about 41% fewer output tokens.
85
86 <img src="assets/gpqa-score-tokens.png" alt="GPQA Diamond attempt pass rate versus mean output tokens for Pi FP8 and Base at low, medium, and xhigh reasoning; y-axis 75–90 percent" width="100%" />
87
88 > Pi’s RL stage encouraged economical low/medium reasoning while keeping xhigh focused on correctness. The graph shows a smoother rise in Pi’s attempt success as effort increases, while Base peaks at medium. Pi achieves the highest xhigh score, but Base retains the medium-effort edge.
89
90 <img src="assets/scicode-score-tokens.png" alt="SciCode passed subproblems out of 337 versus mean output tokens for Pi FP8 and Base at low, medium, and xhigh reasoning; y-axis 35–50 percent" width="100%" />
91
92 > Pi’s development prioritized successful solutions alongside resource use—not shorter responses alone. Both models improve with higher effort, but Pi solves more subproblems at every level. At xhigh, Pi scores higher with about 23% fewer output tokens; at medium, its higher score requires more tokens.
93
94 ## Quickstart
95
96 Download the Q4_K_M model, vision projector, and compact Q4_0 MTP head. These files are used by both quickstarts below.
97
98 ```bash
99 hf download bytkim/Qwen3.8-27B-pi-GGUF \
100 --include "Qwen3.8-27B-pi-Q4_K_M.gguf" \
101 "mmproj-Qwen3.8-27B-pi-BF16.gguf" \
102 "mtp/mtp-Qwen3.8-27B-pi-Q4_0.gguf" \
103 --local-dir models
104 ```
105
106 ### Thinking mode
107
108 ```bash
109 llama-server \
110 --model models/Qwen3.8-27B-pi-Q4_K_M.gguf \
111 --mmproj models/mmproj-Qwen3.8-27B-pi-BF16.gguf \
112 --model-draft models/mtp/mtp-Qwen3.8-27B-pi-Q4_0.gguf \
113 --spec-type draft-mtp \
114 --spec-draft-n-max 3 \
115 --ctx-size 262144 \
116 --parallel 1 \
117 --temp 1.0 \
118 --top-k 20 \
119 --min-p 0.0
120 ```
121
122 > Sampling is configured explicitly where llama.cpp’s defaults differ from Qwen’s recommended thinking settings. Unlike vLLM, llama.cpp does not automatically load this repository’s `generation_config.json`; API requests can override these server defaults.
123
124 **Choose your MTP head:** the quickstart uses Q4_0. To use BF16 or Q8_0 instead, download that head and change `--model-draft`.
125
126 | MTP head | Download size | Draft-3 | Draft-6 | Draft-8 |
127 |---|---:|---:|---:|---:|
128 | [`mtp-Qwen3.8-27B-pi-BF16.gguf`](mtp/mtp-Qwen3.8-27B-pi-BF16.gguf) | 5.95 GB | +24–35% | +29–39% | +13–32% |
129 | [`mtp-Qwen3.8-27B-pi-Q8_0.gguf`](mtp/mtp-Qwen3.8-27B-pi-Q8_0.gguf) | 3.16 GB | +48–57% | +52–66% | +42–52% |
130 | [`mtp-Qwen3.8-27B-pi-Q4_0.gguf`](mtp/mtp-Qwen3.8-27B-pi-Q4_0.gguf) | 2.01 GB | +50–58% | +60–72% | +48–55% |
131
132 *Decode speedup vs. no MTP on a 3-task subset using Q4_K_M, RTX PRO 6000, 262K context, and parallel 1.*
133
134 The MTP head does not need to match your main model’s quantization. To run without MTP, omit `--model-draft`, `--spec-type`, and `--spec-draft-n-max`.
135
136 ### Non-thinking mode
137
138 For direct responses without a thinking section, use this command instead:
139
140 ```bash
141 llama-server \
142 --model models/Qwen3.8-27B-pi-Q4_K_M.gguf \
143 --mmproj models/mmproj-Qwen3.8-27B-pi-BF16.gguf \
144 --model-draft models/mtp/mtp-Qwen3.8-27B-pi-Q4_0.gguf \
145 --spec-type draft-mtp \
146 --spec-draft-n-max 3 \
147 --ctx-size 262144 \
148 --parallel 1 \
149 --chat-template-kwargs '{"enable_thinking":false}' \
150 --temp 0.7 \
151 --top-p 0.8 \
152 --top-k 20 \
153 --min-p 0.0 \
154 --presence-penalty 1.5
155 ```
156
157 > This command disables thinking by default and applies Qwen’s recommended non-thinking sampling settings. Turning thinking off alone does not change sampling.
158
159 ## Advanced
160
161 ### DFlash2 speculative decoding
162
163 #### Download the DFlash2 companion
164
165 Download one mirrored [DFlash2 companion](dflash2/) before starting the server. This command downloads the 1.14 GB Q4_K_M draft used below, without downloading the other variants.
166
167 ```bash
168 hf download bytkim/Qwen3.8-27B-pi-GGUF \
169 --include "dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf" \
170 --local-dir models
171 ```
172
173 For Q8_0 or BF16, replace the filename in both the download command and `--model-draft`. Download your main model separately as shown in Quickstart.
174
175 > DFlash2 companion weights are mirrored unchanged from [Inco AI](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF), with attribution and license included in [`dflash2/`](dflash2/).
176
177 ```bash
178 llama-server \
179 --model models/Qwen3.8-27B-pi-Q4_K_M.gguf \
180 --model-draft models/dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
181 --spec-type draft-dflash \
182 --spec-draft-n-max 7 \
183 --ctx-size 262144 \
184 --parallel 1 \
185 --temp 1.0 \
186 --top-p 0.95 \
187 --top-k 20 \
188 --min-p 0.0
189 ```
190
191 **Choose your DFlash2 companion:** substitute its path in `--model-draft`.
192
193 | DFlash2 companion | Download size | Option |
194 |---|---:|---|
195 | [`Qwen3.8-27B-DFlash2-BF16.gguf`](dflash2/Qwen3.8-27B-DFlash2-BF16.gguf) | 3.86 GB | Original · unquantized |
196 | [`Qwen3.8-27B-DFlash2-Q8_0.gguf`](dflash2/Qwen3.8-27B-DFlash2-Q8_0.gguf) | 2.06 GB | Higher precision |
197 | [`Qwen3.8-27B-DFlash2-Q4_K_M.gguf`](dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf) | 1.14 GB | Default · smallest download |
198
199 The DFlash2 companion does not need to match your main model’s quantization. Use it instead of the MTP companion, with a recent llama.cpp build that supports DFlash2. Generation speed and available context depend on your hardware and workload.
200
201 ### Flash Attention and KV-cache formats
202
203 llama.cpp supports `f32`, `f16`, `bf16`, `q8_0`, `q5_0`, `q5_1`, `q4_0`, `q4_1`, and `iq4_nl` as cache-format options.
204
205 Add these options to either Quickstart command for Q8 cache with Flash Attention:
206
207 ```bash
208 --flash-attn on \
209 --cache-type-k q8_0 \
210 --cache-type-v q8_0
211 ```
212
213 Flash Attention accepts `on`, `off`, or `auto` (default). Quantized V cache requires Flash Attention; `auto` enables it when quantized V is requested, but the backend must support the combination. Flash Attention is an attention implementation, not DFlash2 speculative decoding.
214
215 ### KV-cache memory
216
217 Smaller cache formats reduce memory use, leaving more room for longer conversations. K and V can use different formats; the estimates below use the same format for both at **256K context** (`--ctx-size 262144 --parallel 1`).
218
219 | Cache format | K (GiB) | V (GiB) | Total (GiB) |
220 |---|---:|---:|---:|
221 | `f32` | 16 | 16 | 32 |
222 | `f16` / `bf16` | 8 | 8 | 16 |
223 | `q8_0` | 4.25 | 4.25 | 8.50 |
224 | `q5_1` | 3 | 3 | 6 |
225 | `q5_0` | 2.75 | 2.75 | 5.50 |
226 | `q4_1` | 2.50 | 2.50 | 5 |
227 | `q4_0` / `iq4_nl` | 2.25 | 2.25 | 4.50 |
228
229 Start with `f16`, or try `q8_0` to save memory. Shorter contexts use proportionally less cache; these estimates exclude model weights and other runtime memory. Check quality when using smaller formats.
230
231 [llama.cpp cache options](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
232
233 ## License
234
235 Qwen3.8-27B-pi is fine-tuned from [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B), whose base-model weights are licensed under the Apache License 2.0. See [LICENSE](LICENSE) for the full terms and [NOTICE](NOTICE) for upstream attribution and Pi modification details. Mirrored DFlash2 companions retain their license and provenance in [`dflash2/`](dflash2/).
236
237 ## Acknowledgements
238
239 Built on Qwen3.8-27B from the Qwen team and adapted for the Pi agent harness. Thanks to the open-source training, inference, quantization, and evaluation projects—and dataset contributors—that supported its development.
240