Google's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price.
$0.75 input · $3.75 output / 1M tokens
Read guideAI model company
Google develops the Gemini family alongside specialist media, document, voice, and world models from Google DeepMind and Google Cloud. Access varies between the Gemini API, Vertex AI, Google Cloud products, and research previews.
About Google
Google's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price.
$0.75 input · $3.75 output / 1M tokens
Read guideGoogle's versatile image workhorse for fast generation, conversational edits, multiple references, grounded imagery, and outputs up to 4K.
$0.067 / 1K image
Read guideGoogle's preview video model for text-, image-, and video-guided generation with native audio and output options up to 4K.
$0.40 / second at 720p or 1080p
Read guideGoogle's lower-cost Veo 3.1 variant for quickly creating short video clips with native audio and output options up to 4K.
$0.10–$0.30 / generated second
Read guideGoogle's flagship music model for complete, high-fidelity songs with richer arrangements, expressive vocals, structured lyrics, and detailed creative control.
$0.08 / full song
Read guideGoogle's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
$0.005/min audio input · $0.018/min audio output
Read guideGoogle DeepMind's research world model for generating photorealistic, controllable environments that can be explored in real time from a text description.
Not published
Read guideGoogle Cloud's Gemini-assisted parser for preserving document hierarchy and creating context-rich chunks for enterprise search and RAG.
$10 / 1,000 pages
Read guideGoogle Cloud's Gemini-powered Document AI processor for extracting the exact fields and derived values defined in a business schema.
$30 / 1,000 pages for the first 1M pages
Read guideGoogle's multimodal embedding model for placing text, images, video, audio, and PDFs in one searchable vector space.
Text $0.20 / 1M tokens
Read guideGoogle Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
$0.016/min standard · $0.003/min dynamic batch
Read guide