dflash2/UPSTREAM_README.md
3.5 KB · 124 lines · markdown Raw
1 ---
2 license: apache-2.0
3 library_name: llama.cpp
4 pipeline_tag: text-generation
5 base_model:
6 - Qwen/Qwen3.8-27B
7 inference: false
8 tags:
9 - gguf
10 - dflash2
11 - speculative-decoding
12 - draft-model
13 - llama.cpp
14 ---
15
16 # Qwen3.8-27B-DFlash2-GGUF
17
18 [Blog](https://inco.ai/blog/dflash2/) | [GitHub](https://github.com/z-lab/dflash)
19
20 This repository contains GGUF conversions of
21 [`incoai/Qwen3.8-27B-DFlash2`](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2),
22 the DFlash 2 draft model for
23 [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B).
24 It is not a standalone language model: it runs inside a speculative
25 decoding server and drafts tokens for the target model to verify. The
26 checkpoints are also mirrored at
27 [`z-lab/Qwen3.8-27B-DFlash2-GGUF`](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2-GGUF).
28
29 DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts
30 a whole block of tokens in a single pass and keeps the top candidates at
31 every position. A lightweight selector then traces one coherent path through
32 them. Two-tap dynamic convolutions in the backbone keep the draft from
33 decaying toward the end of the block. Decoding is lossless: greedy output
34 matches the target model exactly, and sampling preserves its distribution.
35
36 <div align="center">
37 <img src="assets/dflash2-figure.png" alt="DFlash 2: parallel block drafting with a candidate path selector" width="100%">
38 </div>
39
40 | File | Size |
41 | :--- | ---: |
42 | `Qwen3.8-27B-DFlash2-Q4_K_M.gguf` | 1.1 GB |
43 | `Qwen3.8-27B-DFlash2-Q8_0.gguf` | 2.0 GB |
44 | `Qwen3.8-27B-DFlash2-BF16.gguf` | 3.8 GB |
45
46 ## Quick Start
47
48 Build [llama.cpp](https://github.com/ggml-org/llama.cpp) with DFlash 2
49 support ([PR #27342](https://github.com/ggml-org/llama.cpp/pull/27342)):
50
51 ```bash
52 git clone https://github.com/ggml-org/llama.cpp.git
53 cd llama.cpp
54 git fetch origin pull/27342/head:pr-27342
55 git switch pr-27342
56
57 # NVIDIA CUDA
58 cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
59 cmake --build build -j
60
61 # Apple Silicon
62 cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
63 cmake --build build -j
64 ```
65
66 Then serve:
67
68 ```bash
69 ./build/bin/llama-server \
70 -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
71 -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \
72 --spec-type draft-dflash \
73 --spec-draft-n-max 7
74 ```
75
76 See the [blog post](https://inco.ai/blog/dflash2/) for other engines and
77 more details.
78
79 ## Evaluation
80
81 - Target: [`ggml-org/Qwen3.8-27B-GGUF`](https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF), `Q4_K_M`
82 - Sampling: Qwen3.8's officially recommended parameters (temperature 1.0, top-p 0.95, top-k 20), with `xhigh` reasoning effort
83 - Maximum new tokens: 2048
84 - Prompts: the first eight GSM8K test examples
85
86 ### Acceptance Length
87
88 Acceptance length is the per-request mean of completion tokens divided by
89 verification steps. Higher is better.
90
91 | Draft GGUF | Acceptance Length |
92 | :--- | ---: |
93 | BF16 | 5.28 |
94 | Q8_0 | 5.13 |
95 | Q4_K_M | 5.39 |
96
97 Full evaluations of the base checkpoint are on the
98 [main model card](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2).
99
100 ## Citation
101
102 If you find DFlash 2 useful, please cite:
103
104 ```bibtex
105 @misc{inco2026dflash2,
106 title = {{DFlash 2: Keep Drafting Parallel}},
107 author = {{Inco AI}},
108 year = {2026},
109 month = {August},
110 url = {https://inco.ai/blog/dflash2/}
111 }
112 ```
113
114 Please also cite the original DFlash paper:
115
116 ```bibtex
117 @inproceedings{chen2026dflash,
118 title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
119 author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
120 booktitle = {International Conference on Machine Learning (ICML)},
121 year = {2026}
122 }
123 ```
124