encoding/README.md
| 1 | # DeepSeek-V4 Encoding |
| 2 | |
| 3 | This document describes the prompt encoding format used by DeepSeek-V4 series models. The encoding handles multi-turn conversations, tool calling, extended thinking (reasoning), and quick instruction tasks. |
| 4 | |
| 5 | A self-contained reference implementation is provided in `encoding_dsv4.py`. |
| 6 | |
| 7 | ## Quick Start |
| 8 | |
| 9 | ```python |
| 10 | from encoding_dsv4 import encode_messages, parse_message_from_completion_text |
| 11 | |
| 12 | # Encode a conversation |
| 13 | messages = [ |
| 14 | {"role": "system", "content": "You are a helpful assistant."}, |
| 15 | {"role": "user", "content": "What is 2+2?"}, |
| 16 | ] |
| 17 | prompt = encode_messages(messages, thinking_mode="thinking") |
| 18 | # => "<|begin▁of▁sentence|>You are a helpful assistant.<|User|>What is 2+2?<|Assistant|><think>" |
| 19 | |
| 20 | # Parse model output back to structured message |
| 21 | completion = "Simple arithmetic.</think>2 + 2 = 4.<|end▁of▁sentence|>" |
| 22 | parsed = parse_message_from_completion_text(completion, thinking_mode="thinking") |
| 23 | # => {"role": "assistant", "reasoning_content": "Simple arithmetic.", "content": "2 + 2 = 4.", "tool_calls": []} |
| 24 | ``` |
| 25 | |
| 26 | > **Note:** The `parse_message_from_completion_text` function is designed to handle well-formatted model output only. It does not attempt to correct or recover from malformed output that the model might occasionally generate. For production use, additional error handling is recommended. |
| 27 | |
| 28 | ## Message Format |
| 29 | |
| 30 | ### Special Tokens |
| 31 | |
| 32 | | Token | Purpose | |
| 33 | |-------|---------| |
| 34 | | `<|begin▁of▁sentence|>` | Beginning of sequence (BOS) | |
| 35 | | `<|end▁of▁sentence|>` | End of assistant turn (EOS) | |
| 36 | | `<|User|>` | User turn prefix | |
| 37 | | `<|Assistant|>` | Assistant turn prefix | |
| 38 | | `<|latest_reminder|>` | Latest reminder (date, locale, etc.) | |
| 39 | | `<think>` / `</think>` | Reasoning block delimiters | |
| 40 | | `|DSML|` | DSML markup token | |
| 41 | |
| 42 | ### Roles |
| 43 | |
| 44 | The encoding supports the following message roles: `system`, `user`, `assistant`, `tool`, `latest_reminder`, and `developer`. |
| 45 | |
| 46 | > **Note on the `developer` role:** The `developer` role is used exclusively in the internal search agent pipeline. It is not needed for general-purpose chat or tool-calling tasks, and the official API does not accept messages with this role. |
| 47 | |
| 48 | ### Basic Chat |
| 49 | |
| 50 | A simple multi-turn conversation is encoded as: |
| 51 | |
| 52 | ``` |
| 53 | <|begin▁of▁sentence|>{system_prompt} |
| 54 | <|User|>{user_message}<|Assistant|></think>{response}<|end▁of▁sentence|> |
| 55 | <|User|>{user_message_2}<|Assistant|></think>{response_2}<|end▁of▁sentence|> |
| 56 | ``` |
| 57 | |
| 58 | - The BOS token is prepended at the very beginning of the conversation. |
| 59 | - In **chat mode** (`thinking_mode="chat"`), `</think>` is placed right after `<|Assistant|>` to immediately close the thinking block, so the model generates content directly. |
| 60 | |
| 61 | ### Interleaved Thinking Mode |
| 62 | |
| 63 | In **thinking mode** (`thinking_mode="thinking"`), the model produces explicit reasoning inside `<think>...</think>` blocks before responding. |
| 64 | |
| 65 | ``` |
| 66 | <|begin▁of▁sentence|>{system_prompt} |
| 67 | <|User|>{message}<|Assistant|><think>{reasoning}</think>{response}<|end▁of▁sentence|> |
| 68 | ``` |
| 69 | |
| 70 | The `drop_thinking` parameter (default `True`) controls whether reasoning from earlier turns is preserved: |
| 71 | |
| 72 | - **Without tools**: `drop_thinking` takes effect. Reasoning content from assistant turns **before** the last user message is stripped. Only the final assistant turn retains its `<think>...</think>` block. |
| 73 | - **With tools** (on system or developer message): `drop_thinking` is automatically disabled. All turns retain their reasoning, because tool-calling conversations require full context for the model to track multi-step reasoning across tool calls. |
| 74 | |
| 75 | ### Tool Calling (DSML Format) |
| 76 | |
| 77 | Tools are defined on the `system` or `developer` message via the `tools` field (OpenAI-compatible format). When tools are present, the following schema block is injected into the system/user prompt: |
| 78 | |
| 79 | ``` |
| 80 | ## Tools |
| 81 | |
| 82 | You have access to a set of tools to help answer the user's question. You can invoke tools by writing a "<|DSML|tool_calls>" block like the following: |
| 83 | |
| 84 | <|DSML|tool_calls> |
| 85 | <|DSML|invoke name="$TOOL_NAME"> |
| 86 | <|DSML|parameter name="$PARAMETER_NAME" string="true|false">$PARAMETER_VALUE</|DSML|parameter> |
| 87 | ... |
| 88 | </|DSML|invoke> |
| 89 | <|DSML|invoke name="$TOOL_NAME2"> |
| 90 | ... |
| 91 | </|DSML|invoke> |
| 92 | </|DSML|tool_calls> |
| 93 | |
| 94 | String parameters should be specified as is and set `string="true"`. For all other types (numbers, booleans, arrays, objects), pass the value in JSON format and set `string="false"`. |
| 95 | |
| 96 | If thinking_mode is enabled (triggered by <think>), you MUST output your complete reasoning inside <think>...</think> BEFORE any tool calls or final response. |
| 97 | |
| 98 | Otherwise, output directly after </think> with tool calls or final response. |
| 99 | |
| 100 | ### Available Tool Schemas |
| 101 | |
| 102 | {tool_definitions_json} |
| 103 | |
| 104 | You MUST strictly follow the above defined tool name and parameter schemas to invoke tool calls. |
| 105 | ``` |
| 106 | |
| 107 | An actual tool call in the assistant turn looks like: |
| 108 | |
| 109 | ```xml |
| 110 | <|DSML|tool_calls> |
| 111 | <|DSML|invoke name="function_name"> |
| 112 | <|DSML|parameter name="param" string="true">string_value</|DSML|parameter> |
| 113 | <|DSML|parameter name="count" string="false">5</|DSML|parameter> |
| 114 | </|DSML|invoke> |
| 115 | </|DSML|tool_calls><|end▁of▁sentence|> |
| 116 | ``` |
| 117 | |
| 118 | - `string="true"`: the parameter value is a raw string. |
| 119 | - `string="false"`: the parameter value is JSON (number, boolean, array, object). |
| 120 | |
| 121 | Tool execution results are wrapped in `<tool_result>` tags within user messages: |
| 122 | |
| 123 | ``` |
| 124 | <|User|><tool_result>{result_json}</tool_result><|Assistant|><think>... |
| 125 | ``` |
| 126 | |
| 127 | When multiple tool results are present, they are sorted by the order of the corresponding `tool_calls` in the preceding assistant message. |
| 128 | |
| 129 | ### Reasoning Effort |
| 130 | |
| 131 | In thinking mode, the `reasoning_effort` parameter selects one of three levels, which control how much deliberation the model spends before answering. The level is realized purely as a text prefix prepended at the very beginning of the prompt (before the system message); the rest of the encoding is identical across levels. |
| 132 | |
| 133 | | `reasoning_effort` | Prompt prefix | |
| 134 | |:---|:---| |
| 135 | | `"low"` (default) | none | |
| 136 | | `"high"` | `Reasoning Effort: Absolute maximum ...` | |
| 137 | | `"max"` | `Reasoning Effort: Beyond maximum ...` | |
| 138 | |
| 139 | `reasoning_effort` has no effect in chat mode (`thinking_mode="chat"`), where the model does not produce a reasoning block at all. |
| 140 | |
| 141 | The full prefix text for `"high"`: |
| 142 | |
| 143 | ``` |
| 144 | Reasoning Effort: Absolute maximum with no shortcuts permitted. |
| 145 | You MUST be very thorough in your thinking and comprehensively decompose the problem to resolve the root cause, rigorously stress-testing your logic against all potential paths, edge cases, and adversarial scenarios. |
| 146 | Explicitly write out your entire deliberation process, documenting every intermediate step, considered alternative, and rejected hypothesis to ensure absolutely no assumption is left unchecked. |
| 147 | ``` |
| 148 | |
| 149 | And for `"max"`: |
| 150 | |
| 151 | ``` |
| 152 | Reasoning Effort: Beyond maximum — exhaustive, relentless, and uncompromising. |
| 153 | You MUST reason with the utmost depth and rigor, leaving absolutely nothing to chance: exhaustively decompose the problem into its most fundamental components, trace every causal chain to its root, and resolve the underlying cause rather than any surface symptom. |
| 154 | Do not stop reasoning until you have independently verified the solution from multiple angles and are certain that no assumption remains unchecked and no error remains undiscovered. |
| 155 | ``` |
| 156 | |
| 157 | ### Quick Instruction Special Tokens |
| 158 | |
| 159 | Quick instruction tokens are used for auxiliary classification and generation tasks. They are appended to messages via the `"task"` field to trigger specialized model behavior for a single-token or short-form output. |
| 160 | |
| 161 | | Special Token | Description | Format | |
| 162 | |:---|:---|:---| |
| 163 | | `<|action|>` | Determines whether the user prompt requires a web search or can be answered directly. | `...<|User|>{prompt}<|Assistant|><think><|action|>` | |
| 164 | | `<|title|>` | Generates a concise conversation title after the first assistant response. | `...<|Assistant|>{response}<|end▁of▁sentence|><|title|>` | |
| 165 | | `<|query|>` | Generates search queries for the user prompt. | `...<|User|>{prompt}<|query|>` | |
| 166 | | `<|authority|>` | Classifies the user prompt's demand for source authoritativeness. | `...<|User|>{prompt}<|authority|>` | |
| 167 | | `<|domain|>` | Identifies the domain of the user prompt. | `...<|User|>{prompt}<|domain|>` | |
| 168 | | `<|extracted_url|>` `<|read_url|>` | Determines whether each URL in the user prompt should be fetched and read. | `...<|User|>{prompt}<|extracted_url|>{url}<|read_url|>` | |
| 169 | |
| 170 | Usage in message format: |
| 171 | |
| 172 | - **`action`** on a user message: the `<|action|>` token is placed after the assistant prefix and thinking token, triggering a routing decision (e.g., "Search" or "Answer"). |
| 173 | - **Other tasks** (`query`, `authority`, `domain`, `read_url`) on a user message: the task token is appended directly after the user content. |
| 174 | - **`title`** on an assistant message: the `<|title|>` token is appended after the assistant's EOS. The next assistant message provides the generated title. |
| 175 | |