Gemini Embedding 2
Google's multimodal embedding model for placing text, images, video, audio, and PDFs in one searchable vector space.
Google's multimodal embedding model for placing text, images, video, audio, and PDFs in one searchable vector space.
Plain-English overview
Gemini Embedding 2 lets a search system compare meaning across several media types instead of maintaining a separate vector model for every format. A text query can retrieve a relevant image, video segment, audio clip, PDF, or passage when they are embedded into the same space.
Developers can choose a smaller output size to reduce database storage or keep the full 3,072 dimensions when retrieval quality matters more. Task types help distinguish search queries from indexed documents, classification, clustering, and other workloads.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Assumes 600 tokens per page, processed separately. Actual token counts vary. This covers embedding only, not storage, search, or generated answers.
Gemini Embedding 2
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Generally available through the Gemini Developer API for text and multimodal embeddings.
Gemini API access follows Google's supported-country list, project billing state, and age requirements.
Check live availabilityData and training
Google's Gemini API terms distinguish unpaid and paid services: content from unpaid services may be used to improve products, while paid-service prompts and responses are not used to improve products.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Google's multimodal embedding model for placing text, images, video, audio, and PDFs in one searchable vector space. Gemini Embedding 2 lets a search system compare meaning across several media types instead of maintaining a separate vector model for every format. A text query can retrieve a relevant image, video segment, audio clip, PDF, or passage when they are embedded into the same space.
Gemini Embedding 2 was released on April 22, 2026 according to the cited provider materials.
Generally available through the Gemini Developer API for text and multimodal embeddings. The access routes listed in this guide are Google and Google AI Studio.
Text $0.20 / 1M tokens. Paid standard pricing also lists image, audio, and video input separately. Batch rates are half the standard paid rates; the free tier has different data-use terms.
Generally available through the Gemini Developer API for text and multimodal embeddings. Gemini API access follows Google's supported-country list, project billing state, and age requirements.
Google's Gemini API terms distinguish unpaid and paid services: content from unpaid services may be used to improve products, while paid-service prompts and responses are not used to improve products. The policy belongs to the provider route and account terms, so verify it again before production use.
USD per million tokens for the shown API routes at standard context length. Cache and long-context rates may differ. Models without matching reviewed prices are omitted; lower cost does not mean better quality.
Mistral AI
Voyage AI
Voyage AI
Voyage AI
OpenAI
This model
Related comparisons
OpenAI's highest-capability text embedding model for semantic search, recommendations, clustering, and retrieval pipelines.
Open full comparisonCohere's enterprise embedding model for multilingual text, images, and visually rich documents with a 128K context window.
Open full comparisonVoyage AI's quality-first general embedding model for text and code retrieval, with adjustable dimensions and a shared family vector space.
Open full comparisonVoyage AI's document-aware embedding model that automatically creates vectors for chunks while preserving information from the surrounding document.
Open full comparisonVoyage AI's specialist embedding model for finding relevant code from natural-language questions or other source-code context.
Open full comparisonMistral's straightforward hosted text embedding model for semantic search, clustering, classification, and RAG.
Open full comparison