inference/README.md
951 B · 27 lines · markdown Raw
1 # Inference code for DeepSeek models
2
3 First convert huggingface model weight files to the format of this project.
4 ```bash
5 export EXPERTS=256
6 export MP=4
7 export CONFIG=config.json
8 python convert.py --hf-ckpt-path ${HF_CKPT_PATH} --save-path ${SAVE_PATH} --n-experts ${EXPERTS} --model-parallel ${MP}
9 ```
10
11 Then chat with DeepSeek model at will!
12 ```bash
13 torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --interactive
14 ```
15
16 Or batch inference from file.
17 ```bash
18 torchrun --nproc-per-node ${MP} generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}
19 ```
20
21 Or multi nodes inference.
22 ```bash
23 torchrun --nnodes ${NODES} --nproc-per-node $((MP / NODES)) --node-rank $RANK --master-addr $ADDR generate.py --ckpt-path ${SAVE_PATH} --config ${CONFIG} --input-file ${FILE}
24 ```
25
26 If you want to use fp8, just remove `"expert_dtype": "fp4"` in `config.json` and specify `--expert-dtype fp8` in `convert.py`.
27