All models
DeepSeek

DeepSeek-V4-Flash

DeepSeek's lower-cost V4 model for high-volume reasoning, coding, agents, and million-token text workloads.

Plain-English overview

What DeepSeek-V4-Flash actually is

DeepSeek-V4-Flash keeps the core text, reasoning, and tool interface of the V4 family while activating a smaller model and charging substantially less than Pro. It is the practical first test for routine coding agents, extraction, long-context analysis, and repeated tool loops.

The standard Flash endpoint is text-only; DeepSeek offers a separate experimental Flash Vision model for image input. Naming that distinction explicitly prevents an application from assuming the stable text route can accept media it does not support.

Good fit for

  • Cost-efficient coding and agent workloads
  • Large-context text analysis
  • High-concurrency applications using familiar API formats

Category comparison

The facts that matter for text models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Context window
1M tokensMaximum combined prompt and working context documented by the provider.
Maximum output
Up to 384K tokensProvider-published response limit, where available.
Knowledge cutoff
Not publishedLatest reliable knowledge date explicitly published by the model provider. Search and connected tools can retrieve newer information but do not change the model's built-in cutoff.
Input price
$0.44 / 1M peak uncached; $0.22 off-peakCurrent standard list price per million input tokens unless noted.
Output price
$1.32 / 1M peak; $0.66 off-peakCurrent standard list price per million output tokens unless noted.
Inputs
Text; separate Flash-Vision-Exp accepts imagesMedia types accepted by the listed model endpoint.
Tools & agents
Tools, JSON output, Responses API, thinking effort, FIM betaSelected native tools and agent-building capabilities, not an exhaustive list.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Example: 6,000 input tokens for 10 pages, plus a 500-token summary. Page lengths vary; adjust the numbers below.

One run sends your input to the model once and receives an answer. Tokens are pieces of text: input is what you send, output is the answer you receive.

Advanced options

DeepSeek-V4-Flash

DeepSeek

Estimated total (USD)

No reviewed rate

For the usage above · USD · API pricing, not a subscription

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get DeepSeek-V4-Flash

Availability

Regions and access stage

Available in the DeepSeek API through OpenAI-, Anthropic-, and Responses-compatible formats; current API version is V4-Flash-0731.

DeepSeek's public model pages do not enumerate one global model-specific country or processing-region list; confirm account eligibility and applicable law.

Check live availability

Data and training

The route matters.

DeepSeek documents its Responses API as stateless and says responses and conversations are not stored by that endpoint. This is not a complete account-wide retention or training-use promise; review current platform terms and automatic cache behavior for sensitive data.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

DeepSeek-V4-Flash FAQ

What is DeepSeek-V4-Flash?

DeepSeek's lower-cost V4 model for high-volume reasoning, coding, agents, and million-token text workloads. DeepSeek-V4-Flash keeps the core text, reasoning, and tool interface of the V4 family while activating a smaller model and charging substantially less than Pro. It is the practical first test for routine coding agents, extraction, long-context analysis, and repeated tool loops.

When was DeepSeek-V4-Flash released?

DeepSeek-V4-Flash was released on July 31, 2026 according to the cited provider materials.

Where can I access DeepSeek-V4-Flash?

Available in the DeepSeek API through OpenAI-, Anthropic-, and Responses-compatible formats; current API version is V4-Flash-0731. The access routes listed in this guide are DeepSeek API and DeepSeek.

How much does DeepSeek-V4-Flash cost?

Peak: $0.44 input · $1.32 output / 1M tokens. Off-peak uncached input and output are $0.22 and $0.66 per million. Cached input is cheaper, and the separate experimental Vision route uses Flash token pricing.

Where is DeepSeek-V4-Flash available?

Available in the DeepSeek API through OpenAI-, Anthropic-, and Responses-compatible formats; current API version is V4-Flash-0731. DeepSeek's public model pages do not enumerate one global model-specific country or processing-region list; confirm account eligibility and applicable law.

Is my DeepSeek-V4-Flash API data used for training?

DeepSeek documents its Responses API as stateless and says responses and conversations are not stored by that endpoint. This is not a complete account-wide retention or training-use promise; review current platform terms and automatic cache behavior for sensitive data. The policy belongs to the provider route and account terms, so verify it again before production use.

Price comparison

USD per million tokens for the shown API routes at standard context length. Cache and long-context rates may differ. Models without matching reviewed prices are omitted; lower cost does not mean better quality.

Input token costsper 1M tokens · USD
  1. $0.15
    Mistral Small 4

    Mistral AI

  2. $0.75
    Gemini 3.8 Flash

    Google

  3. $4.00
    GPT-5.6 Sol

    OpenAI

  4. $10.00
    GPT-6 Astra

    OpenAI

  5. $10.00
    Claude Fable 5.1

    Anthropic

Output token costsper 1M tokens · USD
  1. $0.60
    Mistral Small 4

    Mistral AI

  2. $3.75
    Gemini 3.8 Flash

    Google

  3. $20.00
    GPT-5.6 Sol

    OpenAI

  4. $50.00
    GPT-6 Astra

    OpenAI

  5. $50.00
    Claude Fable 5.1

    Anthropic

Related comparisons

Head-to-head comparisons

  1. DeepSeek-V4-Flash vs GPT-6 Astra

    OpenAI's most capable model for difficult end-to-end work, combining frontier reasoning with a million-token context window and a broad set of agent tools.

    Open full comparison
  2. DeepSeek-V4-Flash vs GPT-5.6 Sol

    OpenAI's flagship general model for difficult coding, analysis, and professional work, with a very large context window and a broad native tool set.

    Open full comparison
  3. DeepSeek-V4-Flash vs Claude Fable 5.1

    Anthropic's frontier model for ambitious, long-running agentic and coding work, with strong vision and enterprise marketplace availability.

    Open full comparison
  4. DeepSeek-V4-Flash vs Gemini 3.8 Flash

    Google's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price.

    Open full comparison
  5. DeepSeek-V4-Flash vs Grok 4.6

    SpaceXAI's frontier text-and-image model for coding, agentic tasks, and knowledge work, with direct web, X, and code tools.

    Open full comparison
  6. DeepSeek-V4-Flash vs Grok 4.20 Multi-Agent Beta

    A specialist Grok model that sends several AI agents to investigate a difficult question in parallel, then combines their work into one researched answer.

    Open full comparison
  7. DeepSeek-V4-Flash vs Inkling

    Thinking Machines Lab's large open-weights model for customizable reasoning, coding, tools, vision, and audio workflows.

    Open full comparison
  8. DeepSeek-V4-Flash vs Mistral Medium 3.5

    Mistral's open-weight frontier model for demanding multimodal, coding, reasoning, and agent workflows, with a 256K context window.

    Open full comparison
  9. DeepSeek-V4-Flash vs Mistral Small 4

    Mistral's lower-cost open model that combines normal instruction following, reasoning, coding, vision, and agent tools in one endpoint.

    Open full comparison
  10. DeepSeek-V4-Flash vs Qwen3.8 Max

    Alibaba's frontier Qwen model for long-context reasoning, coding, tools, and understanding text, images, and video.

    Open full comparison
  11. DeepSeek-V4-Flash vs MiniMax M3

    MiniMax's million-token multimodal model for coding, agents, computer use, and long projects at a low direct API price.

    Open full comparison
  12. DeepSeek-V4-Flash vs GLM-5.3

    Z.ai's open-weight long-context reasoning model for coding, tools, and agentic work across self-hosted and managed routes.

    Open full comparison
  13. DeepSeek-V4-Flash vs Kimi K3

    Moonshot AI's frontier multimodal model for million-token coding, reasoning, knowledge work, and tool-driven agents.

    Open full comparison
  14. DeepSeek-V4-Flash vs Falcon-H1R-7B

    TII's compact open reasoning model for text tasks, long contexts, and self-hosted function-calling workflows.

    Open full comparison
  15. DeepSeek-V4-Flash vs Falcon-H1-34B-Instruct

    TII's 34B open instruction model for multilingual text, coding, and controlled self-hosted applications.

    Open full comparison
  16. DeepSeek-V4-Flash vs Muse Spark 1.3

    Meta's latest multimodal reasoning model for long-running agents, coding, tools, and complex user collaboration.

    Open full comparison
  17. DeepSeek-V4-Flash vs Celeris-1 Magnus

    Celeris' agent-focused diffusion language model for fast reasoning, tool loops, structured actions, and OpenAI-compatible integration.

    Open full comparison
  18. DeepSeek-V4-Flash vs NVIDIA Nemotron 3 Ultra

    NVIDIA's largest Nemotron 3 reasoning model for complex agents, coding, planning, tools, RAG, and million-token analysis.

    Open full comparison
  19. DeepSeek-V4-Flash vs NVIDIA Nemotron 3.5 Lightning

    NVIDIA's compact 30B mixture-of-experts model for efficient specialist agents and high-volume text workflows.

    Open full comparison
  20. DeepSeek-V4-Flash vs Qwen3.8-Flash

    Alibaba's efficient million-context multimodal model for fast agents, coding, document work, vision, video understanding, and tool use.

    Open full comparison
  21. DeepSeek-V4-Flash vs DeepSeek-V4-Pro

    DeepSeek's flagship million-context text model for difficult reasoning, coding, long-running agents, tools, and very large outputs.

    Open full comparison