Gemini 2.0 Flash: Complete Specifications, Pricing, API Access & Use Cases (2026)
Gemini 2.0 Flash: Complete Specifications, Pricing, API Access & Use Cases (2026)
What Is Gemini 2.0 Flash?
Gemini 2.0 Flash is a Google Gemini model designed for fast, cost-efficient multimodal AI workloads. It was part of Google’s second-generation Gemini 2.0 family and was positioned as a workhorse model for developers who needed speed, long context, tool use, and multimodal input handling.
The model supported text, code, image, audio, and video inputs, with text output for the standard API model. It was especially useful for applications that needed to process large documents, visual data, long audio, video files, structured responses, tool calls, and high-volume AI requests.
As of June 2026, Gemini 2.0 Flash should be treated as a legacy model. Google’s current documentation marks Gemini 2.0 Flash as discontinued from June 1, 2026. New production systems should evaluate newer Gemini models instead of starting new deployments on Gemini 2.0 Flash.
What Are Gemini 2.0 Flash’s Key Specifications and Pricing?
The following table summarizes Gemini 2.0 Flash based on official Google documentation and pricing references available as of June 2026.
| Specification | Gemini 2.0 Flash |
|---|---|
| Model name | Gemini 2.0 Flash |
| Provider | |
| Model ID | gemini-2.0-flash; version reference: gemini-2.0-flash-001 |
| Release date | February 5, 2025 |
| Discontinuation date | June 1, 2026 |
| Model family | Gemini 2.0 |
| Model type | Multimodal large language model |
| Knowledge cutoff / data reference date | June 2024 |
| Maximum input tokens | 1,048,576 tokens |
| Maximum output tokens | 8,192 tokens |
| Supported inputs | Text, code, images, audio, video |
| Standard output | Text |
| Context window | 1M tokens |
| Input size limit | 500 MB |
| Function calling | Supported |
| Structured output | Supported |
| System instructions | Supported |
| Code execution | Supported |
| Grounding with Google Search | Supported during availability |
| Explicit context caching | Supported |
| Thinking mode | Not supported in the standard Gemini 2.0 Flash model |
| Live API | Separate preview model: gemini-2.0-flash-live-preview-04-09 |
| Current API status | Discontinued as of June 1, 2026 |
Historical Gemini Developer API pricing for Gemini 2.0 Flash listed the following paid-tier rates per 1 million tokens:
| Pricing Item | Historical Paid-Tier Price |
|---|---|
| Input: text, image, video | \$0.10 / 1M tokens |
| Input: audio | \$0.70 / 1M tokens |
| Output: text | \$0.40 / 1M tokens |
| Context caching: text/image/video | \$0.025 / 1M tokens |
| Context caching: audio | \$0.175 / 1M tokens |
| Context cache storage | \$1.00 / 1M tokens per hour |
| Batch input: text, image, video | \$0.05 / 1M tokens |
| Batch input: audio | \$0.35 / 1M tokens |
| Batch output | \$0.20 / 1M tokens |
These prices are useful for historical comparison and migration analysis, but they should not be used as live production pricing after the model’s discontinuation.
What Can Gemini 2.0 Flash Do That Makes It Useful in Production?
Gemini 2.0 Flash was useful because it combined speed, low historical token cost, long context, and multimodal input support in one model. This made it practical for high-volume applications where a slower flagship model would be too expensive or unnecessary.
Common production capabilities included:
| Pricing Item | Historical Paid-Tier Price |
|---|---|
| Input: text, image, video | \$0.10 / 1M tokens |
| Input: audio | \$0.70 / 1M tokens |
| Output: text | \$0.40 / 1M tokens |
| Context caching: text/image/video | \$0.025 / 1M tokens |
| Context caching: audio | \$0.175 / 1M tokens |
| Context cache storage | \$1.00 / 1M tokens per hour |
| Batch input: text, image, video | \$0.05 / 1M tokens |
| Batch input: audio | \$0.35 / 1M tokens |
| Batch output | \$0.20 / 1M tokens |
Gemini 2.0 Flash was not mainly a deep reasoning model. Its strongest value was fast multimodal throughput, long-context utility, and practical developer integration.
What Are Gemini 2.0 Flash’s Supported Modalities?
Gemini 2.0 Flash supported multimodal input across text, code, images, audio, and video. The standard model output was text.
| Modality | Support Status | Notes |
|---|---|---|
| Text input | Supported | Prompts, documents, instructions, knowledge-base content |
| Code input | Supported | Code review, debugging, explanation, refactoring, documentation |
| Image input | Supported | Screenshots, charts, diagrams, product images, scanned documents |
| Audio input | Supported | Audio summarization, transcription-style workflows, translation-style workflows |
| Video input | Supported | Video understanding, summarization, scene-level analysis |
| Text output | Supported | Standard generation output |
| Audio output | Not supported in the standard model | Available only through the separate Live API preview model during its availability |
| Image output | Not available after shutdown | Historical references should not be treated as live capability |
| Video output | Not supported | Use separate video-generation models where available |
The separate Gemini 2.0 Flash Live API preview model supported audio/video input and audio output, but it had different token limits and a separate model ID.
Where Does Gemini 2.0 Flash Fall Short?
Gemini 2.0 Flash had several practical limitations:
| Limitation | Explanation |
|---|---|
| Discontinued status | The main limitation in 2026 is availability. Google lists Gemini 2.0 Flash as discontinued from June 1, 2026. |
| Not ideal for new deployments | New production systems should consider newer Gemini models with active support. |
| No standard thinking mode | The standard Gemini 2.0 Flash model did not support thinking mode. |
| Text-only standard output | Although it accepted multiple input types, the standard model output was text. |
| Long-context reliability still requires design | A 1M-token window does not guarantee perfect recall across very long inputs. Chunking, retrieval, and validation are still useful. |
| Hallucination risk | Like other LLMs, it could produce inaccurate or unsupported outputs. |
| High-risk use cases require review | Legal, medical, financial, compliance, and safety-sensitive uses require human review and external verification. |
| Migration required | Teams using older model IDs need to update model selection, tests, prompts, cost assumptions, and fallback logic. |
For teams maintaining legacy workflows, the most important task is not new feature exploration but safe migration planning.
What Is Gemini 2.0 Flash Best Used For?
Before discontinuation, Gemini 2.0 Flash was best suited for fast, multimodal, high-throughput use cases.
| Use Case | Fit | Reason |
|---|---|---|
| Document summarization | High | Long context and low historical token cost made it useful for large files |
| Customer support automation | High | Fast response and structured output support helped support workflows |
| Internal knowledge-base Q&A | High | Long context and tool use supported retrieval-style systems |
| Code explanation and documentation | Medium to High | Useful for code understanding and technical writing |
| Multimodal content review | High | Could process text, screenshots, images, audio, and video inputs |
| Meeting and media summarization | High | Audio/video input support enabled transcript and recording analysis workflows |
| Data extraction | High | Structured output and function calling helped convert unstructured content into usable fields |
| Lightweight agent workflows | Medium to High | Tool use made it suitable for task automation, but not deep reasoning-heavy agents |
| Advanced reasoning | Medium | Better handled by newer reasoning-focused or thinking-enabled models |
| New production deployments in 2026 | Low | Discontinued status makes newer models more appropriate |
In 2026, Gemini 2.0 Flash is best understood as an important legacy reference point for evaluating newer Gemini models, rather than as the first choice for new builds.
How Does Gemini 2.0 Flash Compare to Gemini 2.5 Flash and GPT-4o?
Gemini 2.0 Flash is most naturally compared with Gemini 2.5 Flash as its newer Gemini-family successor and GPT-4o as a widely used multimodal general-purpose model. For a broader GPT-4o reference, teams can review this GPT-4o model profile covering specs, pricing, API access, and use cases.
| Comparison Item | Gemini 2.0 Flash | Gemini 2.5 Flash | GPT-4o |
|---|---|---|---|
| Provider | OpenAI | ||
| Primary positioning | Fast second-generation Gemini Flash model | Newer Flash model with hybrid reasoning / thinking budgets | General-purpose multimodal model |
| Context window | 1M tokens | 1M tokens | Smaller than Gemini long-context models |
| Multimodal input | Text, code, image, audio, video | Text, image, video, audio support depending on API configuration | Text, image, and audio capabilities depending on API configuration |
| Standard output | Text | Text, with model-specific multimodal options depending on product/API | Text and multimodal features depending on API |
| Tool use | Supported | Supported | Supported |
| Thinking / reasoning mode | Not supported in standard model | Supported through thinking budgets | Uses its own reasoning and response-generation behavior |
| 2026 availability | Discontinued | Active newer option | Active model family reference |
| Best fit | Legacy high-speed multimodal workflows | Newer production workloads needing speed plus reasoning control | General multimodal assistant, content, coding, and application workflows |
The practical conclusion is simple: Gemini 2.0 Flash was strong for speed and cost-efficient multimodal processing, but Gemini 2.5 Flash is a more appropriate Gemini-family choice for new production use in 2026. GPT-4o remains a useful comparison point for teams evaluating multimodal application design across providers.
How Do I Access Gemini 2.0 Flash?
As of June 2026, Gemini 2.0 Flash is listed as discontinued by Google. The historical model IDs included gemini-2.0-flash and gemini-2.0-flash-001, but these should not be used for new production builds after the shutdown date.
For teams maintaining older integrations, the recommended access path is migration rather than new setup:
- Identify whether your application still references
gemini-2.0-flashorgemini-2.0-flash-001. - Review prompt behavior, token usage, latency, and output quality under a newer Gemini model.
- Update model IDs in your application configuration.
- Re-test structured output, function calling, grounding, caching, and safety behavior.
- Monitor cost changes because newer models may have different pricing and feature behavior.
- Keep rollback and fallback logic available during migration.
Developers who need a currently supported Gemini model should review Google’s latest Gemini model documentation and choose a replacement based on context length, latency, reasoning support, modality requirements, and budget.
FAQs
What is Gemini 2.0 Flash?
Gemini 2.0 Flash is a Google multimodal AI model from the Gemini 2.0 family. It was designed for fast, cost-efficient text generation, tool use, and multimodal input processing across text, code, images, audio, and video.
Is Gemini 2.0 Flash still available?
Google’s current documentation lists Gemini 2.0 Flash as discontinued from June 1, 2026. New production deployments should use a newer supported Gemini model.
What was the Gemini 2.0 Flash context window?
Gemini 2.0 Flash supported a 1,048,576-token input limit, commonly described as a 1M-token context window. It also supported up to 8,192 output tokens.
What was Gemini 2.0 Flash pricing?
Historical Gemini Developer API pricing listed Gemini 2.0 Flash at \$0.10 per 1M text/image/video input tokens, \$0.70 per 1M audio input tokens, and \$0.40 per 1M output tokens on the paid tier.
What modalities did Gemini 2.0 Flash support?
The standard Gemini 2.0 Flash model supported text, code, image, audio, and video input, with text output. A separate Live API preview model supported audio/video input and audio output.
Is Gemini 2.0 Flash good for production?
It was previously useful for production workloads that needed speed, multimodal input, long context, and low historical token cost. In 2026, it should not be selected for new production deployments because it is discontinued.
What should developers use instead of Gemini 2.0 Flash?
Developers should review newer Gemini models, especially newer Flash-family models, and select based on context window, latency, pricing, reasoning support, modality needs, and availability.
