Gemini icon

Gemini 3.7 Flash

3.7 FlashVerified

Improve coding and agent workflows over Gemini 3.6 Flash, with higher results on FrontierCode 1.1, DeepSWE v1.1, WebDev Arena, Terminal-bench 2.1, and AutomationBench Run 1M-context multimodal inputs and 64K-token text outputs with customizable thinking configurations that balance quality, latency, and cost for production agents Use introductory Gemini API pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026 before standard 2027 rates apply

Reviewed by ToolWorthy Editors·updated today·3.7 Flash released yesterday

Pricing:Free + from $0.75/per 1M input tokens until Dec 31, 2026
Visit Site
Jump to section
Gemini 3.7 Flash overview showing coding, agent, PDF reasoning, and introductory API pricing highlights

More tools to compare

MakersClaw icon

MakersClaw

TypingMind icon

TypingMind

Doubao icon

Doubao

Z.ai icon

Z.ai

Slashy icon

Slashy

Invoko icon

Invoko

Pros & Cons

Editor-reviewed

Pros

  • Clear improvement over 3.6 Flash on Google-published coding, agent, workflow, and PDF benchmarks
  • Introductory API price is half the original 3.6 Flash cost through December 31, 2026
  • Supports 1M-token multimodal input context and 64K-token text output
  • Broad availability across AI Studio, Gemini API, Antigravity, Android Studio, Gemini Enterprise, and Gemini Spark
  • Customizable thinking behavior gives developers a cost, latency, and quality control surface

Cons

  • Introductory pricing expires on December 31, 2026, so long-term budgets should use 2027 rates
  • Vendor-published benchmark gains still need validation on real internal workloads
  • Hosted-only model with no open-weight or local deployment path
  • Occasional slowness and timeout issues are noted in the model card
  • Knowledge cutoff behavior varies by domain, so current-information workflows still need grounding

Overview

Gemini 3.7 Flash is Google's August 13, 2026 update to the Gemini 3 Flash line. Google calls it its most intelligent workhorse model for coding and agents, released three weeks after Gemini 3.6 Flash with improvements across software engineering, web development, knowledge work, PDF comprehension, and enterprise workflow automation.

The release is best understood as a targeted Flash upgrade rather than a new product. It keeps the 3.6 Flash foundation, 1M-token multimodal input context, 64K-token text output limit, and configurable thinking behavior, then improves agentic reliability and coding quality. For buyers comparing AI agent models, the key question is whether 3.7 Flash reduces retries enough to justify routing more production coding and workflow traffic through Google's stack.

What's New

Stronger coding and agent workflows

Google positions Gemini 3.7 Flash around practical agent work: debugging, issue resolution, web app generation, business workflow automation, and tool-heavy developer tasks. Compared with Gemini 3.6 Flash, Google's published results show higher scores on:

  • FrontierCode 1.1 Main: 43.6% vs 34.4% for production code quality
  • DeepSWE v1.1: 65.3% vs 49.0% for long-horizon software engineering
  • WebDev Arena: 1588 Elo vs 1538 for web development
  • Terminal-bench 2.1: 85.8% vs 78.0% for agentic terminal coding
  • AutomationBench: 30.4% vs 17.0% for enterprise workflow automation

These are Google-published benchmark results, so teams should still validate against their own repositories and tool-use traces. The practical signal is clear enough: 3.7 Flash is intended to make Flash-tier agents less brittle on multi-step coding and workflow tasks.

Better developer experience

Google says 3.7 Flash adapts better to roadblocks, clarifies intent when needed, follows instructions more reliably, and puts more effort into multi-step planning and tool calls. This matters for production agents because retries and manual supervision often cost more than the raw model call.

The model is also used in Gemini Spark, Google's personal agent for Google AI Pro and Ultra subscribers in supported countries. That gives the update a consumer workflow surface as well as developer API relevance.

Improved knowledge and document work

Gemini 3.7 Flash improves over 3.6 Flash on knowledge-dense and document-heavy tasks. Google's model card lists GDP.pdf at 34.0% vs 22.0% for 3.6 Flash, and GDPVal-AA v2 at 1525 Elo vs 1422. The model is also positioned for finance, law, biosciences, and complex PDF-to-interactive-data-story workflows.

Performance Benchmarks

Google and DeepMind published a broad model-card benchmark table for Gemini 3.7 Flash. The most selection-relevant metrics are below:

Benchmark Gemini 3.7 Flash Gemini 3.6 Flash Why it matters
FrontierCode 1.1 Main 43.6% 34.4% Production code quality
DeepSWE v1.1 65.3% 49.0% Long-horizon software engineering
WebDev Arena 1588 Elo 1538 Elo Web development quality
Terminal-bench 2.1 85.8% 78.0% Agentic terminal coding
Terminal-bench 3.0 14.9% 5.4% General agent capabilities
AutomationBench 30.4% 17.0% Enterprise workflow automation
GDP.pdf 34.0% 22.0% Expert PDF document comprehension
OSWorld-2.0 47.9% 33.8% Agentic computer use
HLE-Verified 53.6% 51.2% Multidisciplinary expert reasoning

Use these numbers as vendor-published selection evidence, not a substitute for an internal eval. For production buyers, the strongest takeaway is the repeated improvement over 3.6 Flash on coding, agent, workflow, and PDF-heavy tasks.

Availability & Access

Gemini 3.7 Flash is distributed across Google's consumer, developer, and enterprise surfaces:

Surface Access path
Gemini API / AI Studio Build with the Gemini API and Google AI Studio
Google Antigravity Agent-first developer workflows
Android Studio Developer integrations
Gemini Enterprise Agent Platform Enterprise agent deployment
Gemini Enterprise app Enterprise end-user workflows
Gemini Spark Personal agent for Google AI Pro and Ultra subscribers in supported countries

The DeepMind model card lists text, image, audio, and video inputs with up to a 1M-token context window. Outputs are text with a 64K-token output limit. There is no required local hardware because the model is distributed through Google's hosted services and APIs.

Compatibility Notes

Gemini 3.7 Flash is based on Gemini 3.6 Flash. It supports customizable thinking configurations, which let developers trade off quality, latency, and cost. If you already migrated to Gemini 3.6 Flash, this should be evaluated as a model-routing and prompt-behavior update rather than a full platform migration.

For production use, test these areas before switching traffic:

  • Tool-call reliability on your real function schemas and agent traces
  • Cost per completed task, not only cost per token
  • Multi-turn behavior with previous_interaction_id or equivalent conversation state
  • Prompt sensitivity in coding tasks that previously depended on 3.6 Flash behavior
  • Latency and timeout behavior on long documents, large repository context, and tool-heavy runs

Pricing & Plans

Gemini 3.7 Flash uses introductory Gemini API pricing through the end of 2026:

Period Input Output
Introductory pricing through Dec 31, 2026 $0.75 / 1M tokens $3.75 / 1M tokens
Standard pricing from Jan 1, 2027 $1.50 / 1M tokens $7.50 / 1M tokens

Google says the introductory price is half the original 3.6 Flash cost per million tokens. This creates a temporary window where teams can evaluate 3.7 Flash at a lower rate, but budgets should model the January 1, 2027 price increase before committing high-volume production agents.

Consumer access through Gemini Spark depends on Google AI Pro or Ultra subscriptions and supported-country availability. Enterprise usage may follow Google Cloud or Gemini Enterprise contract terms rather than the public Gemini API rate card.

Safety & Limitations

The DeepMind model card describes Gemini 3.7 Flash as having the general limitations of foundation models, including hallucinations, occasional slowness, and timeout issues. It also states the model ships with strengthened mitigations around Frontier Safety areas including CBRN and cyber offense.

Knowledge cutoff is nuanced: the model card lists March 2026 as the cutoff while noting some domains may behave closer to January 2025, in line with the Gemini 3 family. For current facts, pricing, legal analysis, or market data, teams should use retrieval, grounding, or tool access rather than relying on model memory.

Best For

  • Engineering teams running coding agents that need better issue resolution and debugging than 3.6 Flash
  • Product teams building web app generation, UI prototyping, or interactive data-story workflows
  • Enterprise automation teams using Google Cloud or Gemini Enterprise Agent Platform
  • Knowledge-work products that process PDFs, financial reports, legal documents, or bioscience materials
  • Developers already using Google Antigravity or Gemini API who can compare cost per completed task during the introductory pricing window

FAQ

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is Google's August 2026 Flash model update for coding, agents, web development, knowledge work, and enterprise workflow automation. It is based on Gemini 3.6 Flash and adds algorithmic improvements to the core reasoning foundation.

What is the model ID for Gemini 3.7 Flash?

Google's launch blog points developers to the Gemini API and developer guide, while the model card names the model as Gemini 3.7 Flash. Before production integration, confirm the exact API model string in Google AI Studio or the current Gemini API models page.

How much does Gemini 3.7 Flash cost?

Introductory Gemini API pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026. Starting January 1, 2027, Google says $1.50 per 1M input tokens and $7.50 per 1M output tokens will apply.

How is Gemini 3.7 Flash different from Gemini 3.6 Flash?

It is based on 3.6 Flash but improves coding, web development, agentic terminal tasks, workflow automation, PDF comprehension, and computer-use benchmarks in Google's published results. It also keeps configurable thinking behavior for quality, latency, and cost tradeoffs.

Where can I use Gemini 3.7 Flash?

Google lists Gemini API, Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise Agent Platform, Gemini Enterprise, and Gemini Spark as access surfaces. Consumer Spark access requires Google AI Pro or Ultra in supported countries.

Is Gemini 3.7 Flash open source?

No. Gemini 3.7 Flash is a hosted Google model distributed through Google's services and APIs. The model card does not describe an open-weight or local deployment option.

Version History

3.7 Flash

Current Version

Released on August 13, 2026

+What's new
3 updates
  • Improve coding and agent workflows over Gemini 3.6 Flash, with higher results on FrontierCode 1.1, DeepSWE v1.1, WebDev Arena, Terminal-bench 2.1, and AutomationBench
  • Run 1M-context multimodal inputs and 64K-token text outputs with customizable thinking configurations that balance quality, latency, and cost for production agents
  • Use introductory Gemini API pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026 before standard 2027 rates apply

3.5 Flash-Lite

Released on July 21, 2026

+What's new
3 updates
  • Run high-throughput agentic workloads with the fastest, lowest-cost Gemini 3.5 model at $0.30 per 1M input tokens and $2.50 per 1M output tokens
  • Improve migration quality from Gemini 3.1 Flash-Lite with stronger coding, document extraction, long-context, and multimodal benchmark performance
  • Configure thinking levels from minimal to higher reasoning modes, balancing low-latency execution against multi-step sub-agent and tool-use reliability

3.5 Flash Cyber

Released on July 21, 2026

+What's new
3 updates
  • Find, validate, and patch software vulnerabilities with a lightweight cybersecurity model built on Gemini 3.5 Flash and tuned for CodeMender workflows
  • Scale frequent code security scans by calling a faster Flash-based model across more code paths, then consolidating sub-agent findings into high-quality reports
  • Access the model through a limited pilot for governments and trusted partners, reflecting Google's staged approach to dual-use cybersecurity deployment

3.6 Flash

Released on July 21, 2026

View Update
+What's new
3 updates
  • Improve agentic coding, knowledge work, and multimodal reasoning with Gemini 3.6 Flash, Google's GA workhorse model for production agent loops
  • Reduce execution cost versus Gemini 3.5 Flash with the same $1.50 input price, lower $7.50 output pricing, and fewer turns and tool calls
  • Use built-in tools including Computer Use Preview, code execution, search grounding, file search, structured outputs, caching, and 1M-token context

3.5 Flash

Released on May 19, 2026

View Update
+What's new
3 updates
  • Tackle long-horizon agentic tasks with a GA Flash model that outperforms Gemini 3.1 Pro on Google's Terminal-Bench 2.1 and MCP Atlas benchmarks
  • Run frontier-quality workflows at $1.50 per 1M input tokens and $9.00 per 1M output tokens, with Batch, Flex, and Priority consumption options
  • Build richer interactive web UIs and graphics with improved multimodal reasoning, 1M-token context, and support across Google AI Studio and Vertex AI

3.1 Flash-Lite

Released on May 7, 2026

View Update
+What's new
3 updates
  • Promote Gemini 3.1 Flash-Lite from its March 2026 preview to the stable gemini-3.1-flash-lite model, Google's fastest and most cost-efficient Gemini 3 series model for high-volume developer workloads
  • Run high-volume work cheaply at $0.25 per 1M text, image, and video input tokens and $1.50 per 1M output tokens, with a free tier plus Batch, Flex, and Priority pricing options
  • Keep strong quality for a lightweight model with 86.9% on GPQA Diamond, 76.8% on MMMU Pro, and 2.5X faster Time to First Answer Token than Gemini 2.5 Flash

3.1 Flash Live (Preview)

Released on March 26, 2026

View Update
+What's new
3 updates
  • Build real-time voice agents that better filter background noise and complete multi-step live tasks with Gemini's highest-quality audio model, which scores 90.8% on ComplexFuncBench Audio
  • Deploy multilingual conversational experiences in 90+ languages for real-time multimodal interactions, making voice-first apps more natural across global customer and enterprise workflows
  • Control live function calling behavior and response latency with thinkingLevel settings, helping balance reasoning depth against speed in production voice applications

3.1 Flash-Lite (Preview)

Released on March 3, 2026

+What's new
3 updates
  • Build high-volume production applications with the fastest Gemini 3 series model — 2.5X faster Time to First Answer Token and 45% higher output speed than Gemini 2.5 Flash — at just $0.25/1M input and $1.50/1M output tokens
  • Handle complex reasoning and multimodal tasks with 86.9% on GPQA Diamond and 76.8% on MMMU Pro, outperforming larger Gemini models from prior generations while maintaining low latency for real-time workflows
  • Control reasoning depth with built-in thinking levels in AI Studio and Vertex AI, enabling cost-efficient scaling from high-volume translation and content moderation to UI generation and multi-step agentic tasks

3.1 Pro (Preview)

Released on February 19, 2026

View Update
+What's new
3 updates
  • Tackle advanced problem-solving with improved reasoning in Gemini 3.1 Pro Preview, which scores 77.1% on ARC-AGI-2 and more than doubles Gemini 3 Pro on that benchmark
  • Use the dedicated gemini-3.1-pro-preview-customtools endpoint to better prioritize your own tools when combining bash and custom tool workflows
  • Build with Gemini 3.1 Pro Preview in Google AI Studio, Vertex AI, Gemini CLI, Antigravity, and other supported Google developer surfaces as the model rolls out across Google products

3 Flash (Preview)

Released on December 17, 2025

View Update
+What's new
3 updates
  • Build high-volume production applications with frontier-level AI performance at lower cost per request, optimized for low-latency workloads requiring fast response times
  • Analyze complex layouts in screenshots and diagrams with upgraded visual and spatial reasoning capabilities, enabling extraction of detailed information from UI mockups and technical diagrams
  • Develop agentic coding assistants with enhanced multimodal function responses and code execution from images, supporting automated development workflows and code generation tasks

3 Pro (Preview)

Released on November 18, 2025

+What's new
3 updates
  • Tackle complex reasoning tasks with the first Gemini 3 series model combining state-of-the-art reasoning, multimodal understanding, and agentic capabilities in one unified model
  • Process long documents and conversations with 1M token context window, enabling comprehensive analysis across multiple PDFs and cross-repository code reviews
  • Control reasoning depth and output format with new behavior configurations including thinking levels and thought signatures, requiring migration validation for production systems

2.5 Pro

Released on June 17, 2025

+What's new
3 updates
  • Generate and review complex code with production-ready stable model featuring adaptive thinking that automatically adjusts reasoning depth based on task complexity
  • Handle comprehensive document analysis with long-context support (1M tokens input), supporting quarterly report synthesis and large codebase understanding
  • Migrate confidently from preview versions with clear upgrade path as experimental builds redirect to stable and lifecycle dates ensure predictable maintenance windows

2.5 Flash (Preview)

Released on April 17, 2025

+What's new
2 updates
  • Build low-latency apps with an early Gemini 2.5 Flash preview optimized for price-performance, giving teams a faster way to test stronger reasoning without moving to a larger model
  • Evaluate adaptive thinking behavior in preview, letting you balance response speed and reasoning depth for chat, summarization, and high-volume application workflows

2.0 Flash Thinking (Preview)

Released on December 19, 2024

+What's new
1 updates
  • Debug hard reasoning tasks with Thinking Mode, a test-time compute model that shows its thought process while generating answers, making evaluation and prompt iteration easier for enterprise Q&A

2.0 Flash (Experimental)

Released on December 11, 2024

+What's new
3 updates
  • Build agentic assistants that understand instructions and take actions with native tool use, enabling early workflow automation and action-oriented assistant prototypes
  • Create richer multimodal experiences with native image generation and steerable multilingual text-to-speech, reducing reliance on separate media-generation services
  • Develop real-time interactive applications with the Multimodal Live API, which supports streaming audio and video input together with multiple tools in one session

1.5 Flash (Preview)

Released on May 14, 2024

+What's new
2 updates
  • Handle large-scale API workloads with the fastest Gemini model optimized for batch summarization, real-time chat, and lightweight content generation requiring rapid response times
  • Process long legal contracts, research materials, and email threads with 1M token context window enabling comprehensive document understanding without splitting

1.5 Flash

Released on May 14, 2024

+What's new
2 updates
  • Handle large-scale API workloads with the fastest Gemini model optimized for batch summarization, real-time chat, and lightweight content generation requiring rapid response times
  • Process long legal contracts, research materials, and email threads with 1M token context window enabling comprehensive document understanding without splitting

1.5 Pro (Preview)

Released on February 15, 2024

+What's new
3 updates
  • Scale production deployments efficiently with Mixture-of-Experts architecture delivering similar quality at lower inference costs and faster iteration cycles for enterprise applications
  • Analyze entire long-form videos, multi-hour audio recordings, and 30,000+ line codebases in single context with breakthrough 1M token window reducing segmentation-induced information loss
  • Achieve reliable cross-modal understanding with 87% performance improvement over 1.0 Pro across benchmarks while maintaining high accuracy throughout long context processing

1.5 Pro

Released on February 15, 2024

+What's new
3 updates
  • Scale production deployments efficiently with Mixture-of-Experts architecture delivering similar quality at lower inference costs and faster iteration cycles for enterprise applications
  • Analyze entire long-form videos, multi-hour audio recordings, and 30,000+ line codebases in single context with breakthrough 1M token window reducing segmentation-induced information loss
  • Achieve reliable cross-modal understanding with 87% performance improvement over 1.0 Pro across benchmarks while maintaining high accuracy throughout long context processing

1.0

Released on December 6, 2023

+What's new
3 updates
  • Deploy AI across devices from mobile to data center with three optimized sizes supporting on-device summarization with Nano, general-purpose services with Pro, and complex tasks with Ultra
  • Process text, images, audio, video, and code in unified model built from ground up for multimodal understanding, eliminating external OCR and pipeline integration complexity
  • Achieve state-of-the-art results with Ultra reaching 90.0% on MMLU benchmark exceeding human expert performance and 59.4% on MMMU providing public baseline for research and enterprise selection

Top alternatives

Related categories

From the blog

View all →

Track Gemini in ToolWorthy Weekly

Important tool updates, better alternatives, and selected AI signals in one weekly brief.

Weekly only. Unsubscribe anytime.