All models
Google

Gemini 3.8 Flash

Google's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price.

Plain-English overview

What Gemini 3.8 Flash actually is

Gemini 3.8 Flash combines a million-token input window with text, image, video, audio, and PDF understanding. Its native API features include function calling, code execution, search grounding, file search, and structured output.

The current price is promotional. Teams comparing annual costs should model both the introductory rate and the published higher rate that begins in 2027, plus any search or grounding calls.

Good fit for

  • High-volume multimodal agents
  • Long video, audio, PDF, and code inputs
  • Teams optimizing capability per token cost

Category comparison

The facts that matter for text models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Context window
1,048,576 tokensMaximum combined prompt and working context documented by the provider.
Maximum output
65,536 tokensProvider-published response limit, where available.
Knowledge cutoff
March 2026; some domains may lag to January 2025Latest reliable knowledge date explicitly published by the model provider. Search and connected tools can retrieve newer information but do not change the model's built-in cutoff.
Input price
$0.75 / 1M promoCurrent standard list price per million input tokens unless noted.
Output price
$3.75 / 1M promoCurrent standard list price per million output tokens unless noted.
Inputs
Text, image, video, audio, PDFMedia types accepted by the listed model endpoint.
Tools & agents
Functions, search, code execution, file searchSelected native tools and agent-building capabilities, not an exhaustive list.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Example: 6,000 input tokens for 10 pages, plus a 500-token summary. Page lengths vary; adjust the numbers below.

One run sends your input to the model once and receives an answer. Tokens are pieces of text: input is what you send, output is the answer you receive.

Advanced options

Gemini 3.8 Flash

Google

Estimated total (USD)

$0.0064

For the usage above · USD · API pricing, not a subscription

Input tokens
$0.75 per 1M tokens
Output tokens
$3.75 per 1M tokens
How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get Gemini 3.8 Flash

Availability

Regions and access stage

Available through the Gemini API and Google AI Studio in supported regions, with Vertex and gateway routes also available.

Use Google's live available-regions list and verify feature restrictions for the selected route.

Check live availability

Data and training

The route matters.

Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

Gemini 3.8 Flash FAQ

What is Gemini 3.8 Flash?

Google's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price. Gemini 3.8 Flash combines a million-token input window with text, image, video, audio, and PDF understanding. Its native API features include function calling, code execution, search grounding, file search, and structured output.

When was Gemini 3.8 Flash released?

Gemini 3.8 Flash was released on September 2, 2026 according to the cited provider materials.

Where can I access Gemini 3.8 Flash?

Available through the Gemini API and Google AI Studio in supported regions, with Vertex and gateway routes also available. The access routes listed in this guide are Google and OpenRouter.

How much does Gemini 3.8 Flash cost?

$0.75 input · $3.75 output / 1M tokens. Introductory paid-tier price through December 31, 2026; Google lists $1.50 input and $7.50 output beginning January 1, 2027.

Where is Gemini 3.8 Flash available?

Available through the Gemini API and Google AI Studio in supported regions, with Vertex and gateway routes also available. Use Google's live available-regions list and verify feature restrictions for the selected route.

Is my Gemini 3.8 Flash API data used for training?

Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region. The policy belongs to the provider route and account terms, so verify it again before production use.

Price comparison

One price per model, using the same input/output mix. We weight each API rate by its share of one million total tokens and use the lowest matching reviewed route. This is a comparison rate, not a cost per task or a quality score. Tokenizers, caching, reasoning and long context can change your actual bill.

Blended token costper 1M total tokens · USD

75% input + 25% output · no cache discounts

  1. $0.20
    GPT-6 Luna

    OpenAI

  2. $0.2625
    Mistral Small 4

    Mistral AI

  3. $0.45
    GPT-5.6 Luna

    OpenAI

  4. $1.50
    Gemini 3.8 Flash

    Google

    This model

  5. $2.00
    Claude Haiku 4.5

    Anthropic

  6. $4.00
    GPT-6 Sol

    OpenAI

  7. $4.00
    Claude Sonnet 5

    Anthropic

  8. $4.50
    GPT-5.6 Terra

    OpenAI

  9. $8.00
    Daybreak Blue

    OpenAI

  10. $8.00
    Claude Opus 5.5

    Anthropic

  11. $8.00
    GPT-5.6 Sol

    OpenAI

  12. $10.00
    Claude Opus 5

    Anthropic

  13. $20.00
    Claude Fable 5

    Anthropic

  14. $20.00
    GPT-6 Astra

    OpenAI

  15. $20.00
    Claude Fable 5.1

    Anthropic

Related comparisons

Head-to-head comparisons

  1. Gemini 3.8 Flash vs GPT-6 Sol

    An OpenAI model for coding projects and assistants that carry out several steps with tools.

    Open full comparison
  2. Gemini 3.8 Flash vs GPT-6 Luna

    A low-cost OpenAI model for focused tasks repeated at scale, such as extraction, classification and short replies.

    Open full comparison
  3. Gemini 3.8 Flash vs Claude Opus 5.5

    A Claude model for large code changes, detailed research and business tasks that take many steps.

    Open full comparison
  4. Gemini 3.8 Flash vs GPT-5.6 Terra

    A lower-cost GPT-5.6 option for coding, analysis and assistants that use tools.

    Open full comparison
  5. Gemini 3.8 Flash vs GPT-5.6 Luna

    The budget GPT-5.6 tier for processing many short, repeatable tasks.

    Open full comparison
  6. Gemini 3.8 Flash vs Claude Fable 5

    The original Fable model for complex coding and multi-step work, retained for existing integrations.

    Open full comparison
  7. Gemini 3.8 Flash vs Claude Opus 5

    A Claude model for substantial code changes, reasoning and business workflows with many steps.

    Open full comparison
  8. Gemini 3.8 Flash vs Claude Sonnet 5

    A lower-priced Claude option for everyday coding, writing and tool-assisted work.

    Open full comparison
  9. Gemini 3.8 Flash vs Claude Haiku 4.5

    A compact Claude model for quick replies, classification and narrowly scoped assistants.

    Open full comparison
  10. Gemini 3.8 Flash vs GPT-6 Astra

    OpenAI's most capable model for difficult end-to-end work, combining frontier reasoning with a million-token context window and a broad set of agent tools.

    Open full comparison
  11. Gemini 3.8 Flash vs GPT-5.6 Sol

    OpenAI's flagship general model for difficult coding, analysis, and professional work, with a very large context window and a broad native tool set.

    Open full comparison
  12. Gemini 3.8 Flash vs Claude Fable 5.1

    Anthropic's frontier model for ambitious, long-running agentic and coding work, with strong vision and enterprise marketplace availability.

    Open full comparison
  13. Gemini 3.8 Flash vs Grok 4.6

    SpaceXAI's frontier text-and-image model for coding, agentic tasks, and knowledge work, with direct web, X, and code tools.

    Open full comparison
  14. Gemini 3.8 Flash vs Grok 4.20 Multi-Agent Beta

    A specialist Grok model that sends several AI agents to investigate a difficult question in parallel, then combines their work into one researched answer.

    Open full comparison
  15. Gemini 3.8 Flash vs Inkling

    Thinking Machines Lab's large open-weights model for customizable reasoning, coding, tools, vision, and audio workflows.

    Open full comparison
  16. Gemini 3.8 Flash vs Mistral Medium 3.5

    Mistral's open-weight frontier model for demanding multimodal, coding, reasoning, and agent workflows, with a 256K context window.

    Open full comparison
  17. Gemini 3.8 Flash vs Mistral Small 4

    Mistral's lower-cost open model that combines normal instruction following, reasoning, coding, vision, and agent tools in one endpoint.

    Open full comparison
  18. Gemini 3.8 Flash vs Qwen3.8 Max

    Alibaba's frontier Qwen model for long-context reasoning, coding, tools, and understanding text, images, and video.

    Open full comparison
  19. Gemini 3.8 Flash vs MiniMax M3

    MiniMax's million-token multimodal model for coding, agents, computer use, and long projects at a low direct API price.

    Open full comparison
  20. Gemini 3.8 Flash vs GLM-5.3

    Z.ai's open-weight long-context reasoning model for coding, tools, and agentic work across self-hosted and managed routes.

    Open full comparison
  21. Gemini 3.8 Flash vs Kimi K3

    Moonshot AI's frontier multimodal model for million-token coding, reasoning, knowledge work, and tool-driven agents.

    Open full comparison
  22. Gemini 3.8 Flash vs Falcon-H1R-7B

    TII's compact open reasoning model for text tasks, long contexts, and self-hosted function-calling workflows.

    Open full comparison
  23. Gemini 3.8 Flash vs Falcon-H1-34B-Instruct

    TII's 34B open instruction model for multilingual text, coding, and controlled self-hosted applications.

    Open full comparison
  24. Gemini 3.8 Flash vs Muse Spark 1.3

    Meta's latest multimodal reasoning model for long-running agents, coding, tools, and complex user collaboration.

    Open full comparison
  25. Gemini 3.8 Flash vs Celeris-1 Magnus

    Celeris' agent-focused diffusion language model for fast reasoning, tool loops, structured actions, and OpenAI-compatible integration.

    Open full comparison
  26. Gemini 3.8 Flash vs NVIDIA Nemotron 3 Ultra

    NVIDIA's largest Nemotron 3 reasoning model for complex agents, coding, planning, tools, RAG, and million-token analysis.

    Open full comparison
  27. Gemini 3.8 Flash vs NVIDIA Nemotron 3.5 Lightning

    NVIDIA's compact 30B mixture-of-experts model for efficient specialist agents and high-volume text workflows.

    Open full comparison
  28. Gemini 3.8 Flash vs Qwen3.8-Flash

    Alibaba's efficient million-context multimodal model for fast agents, coding, document work, vision, video understanding, and tool use.

    Open full comparison
  29. Gemini 3.8 Flash vs DeepSeek-V4-Pro

    DeepSeek's flagship million-context text model for difficult reasoning, coding, long-running agents, tools, and very large outputs.

    Open full comparison
  30. Gemini 3.8 Flash vs DeepSeek-V4-Flash

    DeepSeek's lower-cost V4 model for high-volume reasoning, coding, agents, and million-token text workloads.

    Open full comparison