dflash2/README.md
| 1 | # DFlash2 companion models |
| 2 | |
| 3 | Unmodified official Qwen3.8-27B DFlash2 GGUF drafts from [Inco AI / Z Lab](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF/tree/51962825493a48b846b40126d35c799ac4093ad0). These are companion draft models, not standalone language models or Pi-fine-tuned weights. |
| 4 | |
| 5 | | Draft | Download size | |
| 6 | |---|---:| |
| 7 | | [Qwen3.8-27B-DFlash2-BF16.gguf](Qwen3.8-27B-DFlash2-BF16.gguf) | 3.86 GB | |
| 8 | | [Qwen3.8-27B-DFlash2-Q4_K_M.gguf](Qwen3.8-27B-DFlash2-Q4_K_M.gguf) | 1.14 GB | |
| 9 | | [Qwen3.8-27B-DFlash2-Q8_0.gguf](Qwen3.8-27B-DFlash2-Q8_0.gguf) | 2.06 GB | |
| 10 | |
| 11 | To try a downloaded draft with a compatible llama.cpp build, replace the main Quickstart's MTP options with: |
| 12 | |
| 13 | ```bash |
| 14 | --model-draft models/dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf \ |
| 15 | --spec-type draft-dflash \ |
| 16 | --spec-draft-n-max 7 |
| 17 | ``` |
| 18 | |
| 19 | The files are mirrored unchanged; mirroring does not constitute runtime validation with Pi. Download size is not total runtime memory. |
| 20 | |
| 21 | Licensed under [Apache License 2.0](LICENSE). See [source.json](source.json) for the pinned source revision and SHA-256 checksums, and [UPSTREAM_README.md](UPSTREAM_README.md) for the original attribution and citations. Upstream's PR-branch installation directions are historical; DFlash2 support has since merged into llama.cpp. |
| 22 | |