dflash2/README.md
1.2 KB · 22 lines · markdown Raw
1 # DFlash2 companion models
2
3 Unmodified official Qwen3.8-27B DFlash2 GGUF drafts from [Inco AI / Z Lab](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF/tree/51962825493a48b846b40126d35c799ac4093ad0). These are companion draft models, not standalone language models or Pi-fine-tuned weights.
4
5 | Draft | Download size |
6 |---|---:|
7 | [Qwen3.8-27B-DFlash2-BF16.gguf](Qwen3.8-27B-DFlash2-BF16.gguf) | 3.86 GB |
8 | [Qwen3.8-27B-DFlash2-Q4_K_M.gguf](Qwen3.8-27B-DFlash2-Q4_K_M.gguf) | 1.14 GB |
9 | [Qwen3.8-27B-DFlash2-Q8_0.gguf](Qwen3.8-27B-DFlash2-Q8_0.gguf) | 2.06 GB |
10
11 To try a downloaded draft with a compatible llama.cpp build, replace the main Quickstart's MTP options with:
12
13 ```bash
14 --model-draft models/dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
15 --spec-type draft-dflash \
16 --spec-draft-n-max 7
17 ```
18
19 The files are mirrored unchanged; mirroring does not constitute runtime validation with Pi. Download size is not total runtime memory.
20
21 Licensed under [Apache License 2.0](LICENSE). See [source.json](source.json) for the pinned source revision and SHA-256 checksums, and [UPSTREAM_README.md](UPSTREAM_README.md) for the original attribution and citations. Upstream's PR-branch installation directions are historical; DFlash2 support has since merged into llama.cpp.
22