inference/README.md
| 1 | # Inference code for DeepSeek models |
| 2 | |
| 3 | First convert huggingface model weight files to the format of this project. |
| 4 | ```bash |
| 5 | export EXPERTS=256 |
| 6 | export MP=4 |
| 7 | export CONFIG=config.json |
| 8 | python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} --n-experts ${EXPERTS} --model-parallel ${MP} |
| 9 | ``` |
| 10 | |
| 11 | Then chat with DeepSeek model at will! |
| 12 | ```bash |
| 13 | torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --interactive |
| 14 | ``` |
| 15 | |
| 16 | Or batch inference from file. |
| 17 | ```bash |
| 18 | torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE} |
| 19 | ``` |
| 20 | |
| 21 | Or multi nodes inference. |
| 22 | ```bash |
| 23 | torchrun --nnodes ${NODES} --nproc-per-node $((MP / NODES)) --node-rank $RANK --master-addr $ADDR generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE} |
| 24 | ``` |
| 25 | |
| 26 | If you want to use fp8, just remove `"expert_dtype": "fp4"` in `config.json` and specify `--expert-dtype fp8` in `convert.py`. |
| 27 | |