$ curl -s -w "\nHTTP %{http_code}\n" https://api.deepseek.com/models
Authentication Fails (governor)
HTTP 401
$ ollama pull deepseek-v4.1-flash
Error: pull model manifest: file does not exist
"deepseek-v4.1-flash:cloud" is available as a cloud model. Try:
ollama pull deepseek-v4.1-flash:cloud
$ ollama show deepseek-v4.1-flash:cloud
Model
architecture deepseek_v41
parameters 763205315794
context length 1048576
embedding length 0
quantization FP8
Capabilities
completion
thinking
tools
vision
$ ollama run deepseek-v4.1-flash:cloud "hello"
You need to be signed in to Ollama to run Cloud models.
# encode_demo.py: what DeepSeek-V4.1-Flash actually reads for one chat request with one tool.
# Uses DeepSeek's own encoding.py and tokenizer.json from the Hugging Face repo. No model is run.
import sys
sys.path.insert(0, "../raw/encoding")
from encoding import encode_messages
from tokenizers import Tokenizer
tok = Tokenizer.from_file("../raw/tokenizer.json")
count = lambda s: len(tok.encode(s, add_special_tokens=False).ids)
TOOLS = [{"type": "function", "function": {
"name": "get_order_status",
"description": "Look up the delivery status of an order by its id.",
"parameters": {"type": "object", "properties": {"order_id": {"type": "string"}},
"required": ["order_id"]}}}]
QUESTION = "Where is my order A-1042?"
def build(mode, effort=None, tools=None):
msgs = [{"role": "system", "content": "You are a support assistant for an online shop."}]
if tools:
msgs[0]["tools"] = tools
msgs.append({"role": "user", "content": QUESTION})
return encode_messages(msgs, thinking_mode=mode, reasoning_effort=effort)
print(f"the question alone: {count(QUESTION)} tokens\n")
for label, mode, effort, tools in [
("chat mode, no tool", "chat", None, None),
("chat mode, one tool", "chat", None, TOOLS),
("thinking, effort 50, one tool", "thinking", 50, TOOLS),
("thinking, effort 100, one tool", "thinking", 100, TOOLS),
]:
p = build(mode, effort, tools)
print(f"{label:32s} {count(p):4d} tokens ends with: {p[-24:]!r}")
print("\n--- the full prompt for 'thinking, effort 100, one tool' ---")
print(build("thinking", 100, TOOLS))
$ python -m pytest -q test_encoding.py # DeepSeek's own tests for encoding.py
.................................................. [100%]
50 passed in 0.02s
$ python encode_demo.py
the question alone: 9 tokens
chat mode, no tool 24 tokens ends with: '42?<|Assistant|></think>'
chat mode, one tool 290 tokens ends with: '42?<|Assistant|></think>'
thinking, effort 50, one tool 315 tokens ends with: '042?<|Assistant|><think>'
thinking, effort 100, one tool 315 tokens ends with: '042?<|Assistant|><think>'
--- the full prompt for 'thinking, effort 100, one tool' ---
<|begin▁of▁sentence|><|System|>Reasoning Effort: 100 (range 1-100, the higher the value, the more thorough the reasoning)
You are a support assistant for an online shop.
## Tools
You have access to a set of tools to help answer the user's question. You can invoke tools by writing a "<|DSML| calls>" block like the following:
<|DSML| calls>
<|DSML| invoke name="$TOOL_NAME">
<|DSML| parameter name="$PARAMETER_NAME" string="true|false">$PARAMETER_VALUE</|DSML| parameter>
...
</|DSML| invoke>
<|DSML| invoke name="$TOOL_NAME2">
...
</|DSML| invoke>
</|DSML| calls>
String parameters should be specified as is and set `string="true"`. For all other types (numbers, booleans, arrays, objects), pass the value in JSON format and set `string="false"`.
If thinking_mode is enabled (triggered by <think>), you MUST output your complete reasoning inside <think>...</think> BEFORE any tool calls or final response.
Otherwise, output directly after </think> with tool calls or final response.
### Available Tool Schemas
{"name": "get_order_status", "description": "Look up the delivery status of an order by its id.", "parameters": {"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}}
You MUST strictly follow the above defined tool name and parameter schemas to invoke tool calls.
<|User|>Where is my order A-1042?<|Assistant|><think>