DeepSeek V4 Flash

deepseek/deepseek-v4-flashDeepSeekoperational

DeepSeek V4 Flash: fast, low-cost general model with native multimodal support (retired by DeepSeek; requests route to V4.1 Flash).

Official pricing

USD per million tokens, exactly what DeepSeek publishes. ElevenRouter adds no markup on tokens.

Input
$0.440
Output
$1.32
Cache read
$0.014
Cache write
$0.440

Example: 1,000 input + 500 output tokens ≈ $0.00110. Zero-completion responses are never charged.

Specs

Context
1M
Max output
64K
Modalities
text+image->text
Tokenizer
DeepSeek
Released
2026-09-20
Family
deepseek-flash
Tool calling Structured outputs JSON mode Reasoning Vision Audio input Prompt caching Streaming

Availability on ElevenRouter

Measured by the routing engine; up when at least one credential can serve the model.

Last 3 days
100.00%
Last 30 days
100.00%
30 days agotoday
Full status page

Observed performance

Real requests over the last 7 days.

Time to first token p50
p95
Output tokens/s p50
p95

Percentiles appear after five successful requests.

Usage on ElevenRouter

Tokens per day across all customers (anonymous, aggregate).

0 tokens · 30d
0 requests in the last 7 days

Use it

Any OpenAI or Anthropic SDK works by changing the base URL. Supported parameters: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, repetition_penalty, seed, stop, tools, tool_choice, response_format, reasoning, include_reasoning.

curl https://elevenrouter.com/api/v1/chat/completions \
  -H "Authorization: Bearer $ELEVENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "deepseek/deepseek-v4-flash", "messages": [{ "role": "user", "content": "Hello" }] }'

Related models

Frequently asked questions

How much does DeepSeek V4 Flash cost on ElevenRouter?

DeepSeek V4 Flash is billed at DeepSeek's official list price: $0.440 per million input tokens and $1.32 per million output tokens, with cached input at $0.014. A request with 1,000 input and 500 output tokens costs about $0.00110.

What is the context length of DeepSeek V4 Flash?

DeepSeek V4 Flash accepts up to 1M tokens of context and can produce up to 64K output tokens.

Does DeepSeek V4 Flash support tool calling and structured outputs?

Yes, tool calling is supported and JSON mode is available; add the response-healing plugin for schema validation. Supported parameters: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, repetition_penalty, seed, stop, tools, tool_choice, response_format, reasoning, include_reasoning.

How do I call DeepSeek V4 Flash?

Send an OpenAI-compatible chat completion to https://elevenrouter.com/api/v1/chat/completions with "model": "deepseek/deepseek-v4-flash" and your ElevenRouter key, or use the Anthropic Messages endpoint. Any OpenAI or Anthropic SDK works by changing the base URL.

Is DeepSeek V4 Flash available right now?

Availability is measured continuously by the routing engine. Current status: operational, 100.00% uptime over the last 30 days. When one credential fails, requests fail over to another automatically.

DeepSeek V4 Flash by DeepSeek — pricing, context, availability · ElevenRouter