All models
Google

Veo 3.1 Fast

Google's lower-cost Veo 3.1 variant for quickly creating short video clips with native audio and output options up to 4K.

Plain-English overview

What Veo 3.1 Fast actually is

Veo 3.1 Fast supports the same main creative inputs as the standard model—text, images, and video—but is tuned for quicker, less expensive iteration. It is a practical fit for social clips, ad variations, storyboards, and experiments where teams expect to generate several versions.

Google prices the Fast route by output resolution: 720p is the least expensive, 1080p costs slightly more, and 4K has a larger step up. Preview status and an eight-second requirement for high-resolution output still apply.

Good fit for

  • Rapid ad and social-video variations
  • Short clips that need generated audio
  • Teams seeking a lower-cost route into Veo

Category comparison

The facts that matter for video models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Generation price
$0.10 720p · $0.12 1080p · $0.30 4K / secCurrent list price per generated second or the provider's closest billing unit.
Clip duration
4, 6, or 8 secondsDocumented duration choices or maximum clip length.
Maximum output
4K at 8 secondsHighest provider-documented output resolution, with relevant constraints.
Native audio
Yes, always onWhether the endpoint generates synchronized audio with video.
Input modes
Text, image, or video to videoSupported text, image, video, reference, or motion inputs.
Generation time
11 seconds–6 minutes at peakOnly shown when a provider publishes a range; not a Cody benchmark.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Example: 8-second, 720p clips with audio. Models that do not support these settings will not receive an estimate.

Veo 3.1 Fast

Google

Estimated total (USD)

$0.80

For the usage above · USD · API pricing, not a subscription

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get Veo 3.1 Fast

Availability

Regions and access stage

Available as a preview model through the paid Gemini API in supported regions.

Use Google's live Gemini API region list. Person-generation controls are more restrictive in the EU, UK, Switzerland, and MENA.

Check live availability

Data and training

The route matters.

Google's paid Gemini API terms say prompts and responses are not used to improve its products. Generated Veo videos are stored on the server for two days and must be downloaded to keep them.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

Veo 3.1 Fast FAQ

What is Veo 3.1 Fast?

Google's lower-cost Veo 3.1 variant for quickly creating short video clips with native audio and output options up to 4K. Veo 3.1 Fast supports the same main creative inputs as the standard model—text, images, and video—but is tuned for quicker, less expensive iteration. It is a practical fit for social clips, ad variations, storyboards, and experiments where teams expect to generate several versions.

When was Veo 3.1 Fast released?

The provider does not publish a clear release date for Veo 3.1 Fast in the source material reviewed by Cody.

Where can I access Veo 3.1 Fast?

Available as a preview model through the paid Gemini API in supported regions. The access routes listed in this guide are Google.

How much does Veo 3.1 Fast cost?

$0.10–$0.30 / generated second. Google lists Veo 3.1 Fast with audio at $0.10/s for 720p, $0.12/s for 1080p, and $0.30/s for 4K on the paid Gemini API tier.

Where is Veo 3.1 Fast available?

Available as a preview model through the paid Gemini API in supported regions. Use Google's live Gemini API region list. Person-generation controls are more restrictive in the EU, UK, Switzerland, and MENA.

Is my Veo 3.1 Fast API data used for training?

Google's paid Gemini API terms say prompts and responses are not used to improve its products. Generated Veo videos are stored on the server for two days and must be downloaded to keep them. The policy belongs to the provider route and account terms, so verify it again before production use.

Related comparisons

Head-to-head comparisons

  1. Veo 3.1 Fast vs Sora 2

    OpenAI's legacy synced-audio video model for short text- or image-guided clips—still documented, but approaching a permanent API shutdown.

    Open full comparison
  2. Veo 3.1 Fast vs Veo 3.1

    Google's preview video model for text-, image-, and video-guided generation with native audio and output options up to 4K.

    Open full comparison
  3. Veo 3.1 Fast vs MiniMax H3 Max

    fal's speed-focused post-trained version of MiniMax H3 for short video with native audio, fast 768p iteration, and text, image, or multimodal reference input.

    Open full comparison
  4. Veo 3.1 Fast vs Runway Gen-4.5

    Runway's flagship developer video model for text- and image-guided clips, with professional formats and a predictable credit-per-second base rate.

    Open full comparison
  5. Veo 3.1 Fast vs Kling Video 3.0 Pro

    Kuaishou's Pro video model delivered through fal for multi-shot text/image generation, optional native audio, references, and clips up to 15 seconds.

    Open full comparison
  6. Veo 3.1 Fast vs Kling Video 2.5 Turbo Pro

    A cost-conscious Kling model on fal for fluid five- or ten-second clips generated from text or a starting image.

    Open full comparison
  7. Veo 3.1 Fast vs Grok Imagine Video 1.5

    SpaceXAI's current video model for turning text, images, or references into short clips with optional generated speech and sound.

    Open full comparison
  8. Veo 3.1 Fast vs Seedance 2.5

    ByteDance's current audiovisual model for coherent clips up to 30 seconds, with native sound, long-form storytelling, rich references, and targeted editing.

    Open full comparison
  9. Veo 3.1 Fast vs HappyHorse 1.1

    Alibaba's unified video family for creating short clips with sound from text, a starting image, or a reference video.

    Open full comparison
  10. Veo 3.1 Fast vs LTX-2.5 Pro

    Lightricks' quality-focused video model for making 720p or 1080p clips from text, images, or audio with native sound.

    Open full comparison
  11. Veo 3.1 Fast vs LTX-2.5 Fast

    LTX's speed-optimized video model for text, image, or audio input, native sound, automatic duration, and output up to 4K.

    Open full comparison
  12. Veo 3.1 Fast vs LTX-2.3 Pro

    LTX's mature full-workflow video model for generation plus Retake, Extend, Reframe, native audio, and first-to-last-frame control.

    Open full comparison