dflash2/UPSTREAM_README.md
| 1 | --- |
| 2 | license: apache-2.0 |
| 3 | library_name: llama.cpp |
| 4 | pipeline_tag: text-generation |
| 5 | base_model: |
| 6 | - Qwen/Qwen3.8-27B |
| 7 | inference: false |
| 8 | tags: |
| 9 | - gguf |
| 10 | - dflash2 |
| 11 | - speculative-decoding |
| 12 | - draft-model |
| 13 | - llama.cpp |
| 14 | --- |
| 15 | |
| 16 | # Qwen3.8-27B-DFlash2-GGUF |
| 17 | |
| 18 | [Blog](https://inco.ai/blog/dflash2/) | [GitHub](https://github.com/z-lab/dflash) |
| 19 | |
| 20 | This repository contains GGUF conversions of |
| 21 | [`incoai/Qwen3.8-27B-DFlash2`](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2), |
| 22 | the DFlash 2 draft model for |
| 23 | [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). |
| 24 | It is not a standalone language model: it runs inside a speculative |
| 25 | decoding server and drafts tokens for the target model to verify. The |
| 26 | checkpoints are also mirrored at |
| 27 | [`z-lab/Qwen3.8-27B-DFlash2-GGUF`](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2-GGUF). |
| 28 | |
| 29 | DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts |
| 30 | a whole block of tokens in a single pass and keeps the top candidates at |
| 31 | every position. A lightweight selector then traces one coherent path through |
| 32 | them. Two-tap dynamic convolutions in the backbone keep the draft from |
| 33 | decaying toward the end of the block. Decoding is lossless: greedy output |
| 34 | matches the target model exactly, and sampling preserves its distribution. |
| 35 | |
| 36 | <div align="center"> |
| 37 | <img src="assets/dflash2-figure.png" alt="DFlash 2: parallel block drafting with a candidate path selector" width="100%"> |
| 38 | </div> |
| 39 | |
| 40 | | File | Size | |
| 41 | | :--- | ---: | |
| 42 | | `Qwen3.8-27B-DFlash2-Q4_K_M.gguf` | 1.1 GB | |
| 43 | | `Qwen3.8-27B-DFlash2-Q8_0.gguf` | 2.0 GB | |
| 44 | | `Qwen3.8-27B-DFlash2-BF16.gguf` | 3.8 GB | |
| 45 | |
| 46 | ## Quick Start |
| 47 | |
| 48 | Build [llama.cpp](https://github.com/ggml-org/llama.cpp) with DFlash 2 |
| 49 | support ([PR #27342](https://github.com/ggml-org/llama.cpp/pull/27342)): |
| 50 | |
| 51 | ```bash |
| 52 | git clone https://github.com/ggml-org/llama.cpp.git |
| 53 | cd llama.cpp |
| 54 | git fetch origin pull/27342/head:pr-27342 |
| 55 | git switch pr-27342 |
| 56 | |
| 57 | # NVIDIA CUDA |
| 58 | cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON |
| 59 | cmake --build build -j |
| 60 | |
| 61 | # Apple Silicon |
| 62 | cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON |
| 63 | cmake --build build -j |
| 64 | ``` |
| 65 | |
| 66 | Then serve: |
| 67 | |
| 68 | ```bash |
| 69 | ./build/bin/llama-server \ |
| 70 | -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \ |
| 71 | -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \ |
| 72 | --spec-type draft-dflash \ |
| 73 | --spec-draft-n-max 7 |
| 74 | ``` |
| 75 | |
| 76 | See the [blog post](https://inco.ai/blog/dflash2/) for other engines and |
| 77 | more details. |
| 78 | |
| 79 | ## Evaluation |
| 80 | |
| 81 | - Target: [`ggml-org/Qwen3.8-27B-GGUF`](https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF), `Q4_K_M` |
| 82 | - Sampling: Qwen3.8's officially recommended parameters (temperature 1.0, top-p 0.95, top-k 20), with `xhigh` reasoning effort |
| 83 | - Maximum new tokens: 2048 |
| 84 | - Prompts: the first eight GSM8K test examples |
| 85 | |
| 86 | ### Acceptance Length |
| 87 | |
| 88 | Acceptance length is the per-request mean of completion tokens divided by |
| 89 | verification steps. Higher is better. |
| 90 | |
| 91 | | Draft GGUF | Acceptance Length | |
| 92 | | :--- | ---: | |
| 93 | | BF16 | 5.28 | |
| 94 | | Q8_0 | 5.13 | |
| 95 | | Q4_K_M | 5.39 | |
| 96 | |
| 97 | Full evaluations of the base checkpoint are on the |
| 98 | [main model card](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2). |
| 99 | |
| 100 | ## Citation |
| 101 | |
| 102 | If you find DFlash 2 useful, please cite: |
| 103 | |
| 104 | ```bibtex |
| 105 | @misc{inco2026dflash2, |
| 106 | title = {{DFlash 2: Keep Drafting Parallel}}, |
| 107 | author = {{Inco AI}}, |
| 108 | year = {2026}, |
| 109 | month = {August}, |
| 110 | url = {https://inco.ai/blog/dflash2/} |
| 111 | } |
| 112 | ``` |
| 113 | |
| 114 | Please also cite the original DFlash paper: |
| 115 | |
| 116 | ```bibtex |
| 117 | @inproceedings{chen2026dflash, |
| 118 | title = {{DFlash: Block Diffusion for Flash Speculative Decoding}}, |
| 119 | author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian}, |
| 120 | booktitle = {International Conference on Machine Learning (ICML)}, |
| 121 | year = {2026} |
| 122 | } |
| 123 | ``` |
| 124 | |