Gemini icon

Gemini 3.8 Flash

3.8 FlashCurrent VersionVerified

Upgrade long-horizon coding and autonomous agent workflows with Gemini's newest Flash model, improving on 3.7 Flash across tool use and reasoning Keep the 3.7 Flash introductory API price of $0.75 input and $3.75 output per 1M tokens while testing higher-effort behavior Build across Gemini API, AI Studio, Antigravity, Gemini Enterprise, AI Mode, and the Gemini app with 1M-token multimodal input and 64K text output

Content updated today·3.8 Flash released yesterday

Pricing:Free + from $0.75/per 1M input tokens until Dec 31, 2026
Visit Site
Jump to section
Gemini screenshot

Overview

Gemini 3.8 Flash is Google's September 2, 2026 update to the Gemini 3 Flash line. It arrives three weeks after Gemini 3.7 Flash and keeps the same Flash positioning: a hosted, generally available model for coding agents, autonomous workflows, multimodal reasoning, and enterprise knowledge work.

The release is not just a routine endpoint refresh. Google says 3.8 Flash works harder on complex tasks, using extra reasoning steps and iterative tool calls when higher effort is useful. That creates a real routing decision for teams comparing AI agent models: use 3.8 Flash when completion quality matters more than minimum token usage, and keep 3.7 Flash for efficiency-first workloads that already perform well.

What's New

Stronger long-horizon coding and agents

Google positions Gemini 3.8 Flash as its strongest Flash model for software engineering and autonomous agent tasks. Compared with Gemini 3.7 Flash, the launch post emphasizes better performance on long-horizon coding, specialized professional reasoning, and tool-heavy workflows where the model must evaluate intermediate results and continue iterating.

The important product change is behavioral. Gemini 3.8 Flash can spend more effort on difficult tasks instead of optimizing only for short responses. For coding agents, repository maintenance, UI generation, and business process automation, that can reduce manual supervision when the extra reasoning steps produce a correct final result.

Higher-effort behavior changes cost planning

The API token price stays the same as Gemini 3.7 Flash during the introductory period: $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026. Starting January 1, 2027, Google says the standard rate will be $1.50 per 1M input tokens and $7.50 per 1M output tokens.

That does not mean every workload will cost the same. Google's model card and launch post both warn that 3.8 Flash may use more tokens on complex tasks, especially at higher effort levels. Teams should compare cost per completed task, not only cost per token, because better completion rates can still be worth higher output volume.

Flash Cyber is a restricted defensive variant

Gemini 3.8 also introduces Gemini 3.8 Flash Cyber, a cybersecurity-tuned variant built for vulnerability discovery, automated patching, and defensive research. It uses the same foundational intelligence as 3.8 Flash but ships through Google's Fairwind Program rather than normal Gemini app or API access.

The Cyber variant is relevant to release buyers even if most users cannot access it. It explains part of the 3.8 training emphasis, sets expectations for defensive security capabilities, and creates a separate eligibility question for governments, critical infrastructure operators, core technology platforms, software maintainers, and approved research groups.

Same multimodal model limits with broader tool surface

The Gemini API model page lists the stable model code as gemini-3.8-flash. It supports text, image, video, audio, and PDF inputs, with text output. The token limits remain 1,048,576 input tokens and 65,536 output tokens.

For tool use, Google lists caching, code execution, Computer Use in preview, file search, function calling, grounding with Google Maps, search grounding, structured outputs, thinking at low, medium, and high levels, and URL context. minimal thinking is not supported for this model and returns an error.

Compared With Previous Version

Gemini 3.8 Flash is based on Gemini 3.7 Flash, so the migration should be evaluated as a model-routing change rather than a new platform integration. The main delta is that 3.8 Flash is tuned to do more work on hard agentic tasks, while 3.7 Flash remains supported for teams that prioritize low token overhead.

Google-published examples and model pages highlight better results for long-horizon software engineering, finance-agent workflows, legal-agent workflows, and broad expert reasoning. The public model page gives two concrete comparison points: Vals Finance Agent V2 rises to 61.4% for 3.8 Flash versus 59.0% for 3.7 Flash, and HLE-Verified rises to 54.9% versus 53.6%.

Decision Area Gemini 3.8 Flash Gemini 3.7 Flash
Model role Higher-effort Flash model for complex agents Efficient Flash model for production agents
API model ID gemini-3.8-flash gemini-3.7-flash
Input / output limits 1,048,576 input tokens / 65,536 output tokens 1M input tokens / 64K output tokens
Intro API pricing $0.75 input / $3.75 output per 1M tokens $0.75 input / $3.75 output per 1M tokens
Cost risk May use more tokens on complex tasks Better fit for efficiency-first workloads
Best fit Long-horizon coding, agents, specialist workflows Existing Flash workloads with stable cost targets

Performance Benchmarks

Google's public materials support a performance case for 3.8 Flash, but the model card renders the full benchmark table as images. The safest interpretation is to use the named official results as vendor-published selection evidence and rerun private evals before moving production traffic.

Benchmark or Evaluation Gemini 3.8 Flash Signal Why It Matters
Vals Finance Agent V2 61.4% Specialized finance-agent task execution
HLE-Verified 54.9% Multi-step reasoning across STEM, humanities, and professional fields
DeepSWE v1.1 Google says 3.8 Flash outperforms most larger frontier models Long-horizon software engineering
Harvey's Legal Agent Benchmark Google says 3.8 Flash outperforms 3.7 Flash and other frontier models Professional legal-agent workflows

Use these results as a shortlist signal, not a migration proof. A good rollout test should include your own codebase, tool schemas, retry policy, latency budget, and cost per successful task.

Availability & Access

Gemini 3.8 Flash is broadly distributed across Google's consumer, developer, and enterprise surfaces:

Surface Access Path
Gemini API / Google AI Studio Build with model ID gemini-3.8-flash
Google Antigravity Agent-first coding and workflow experiments
Android Studio Developer tooling integrations
Gemini Enterprise Agent Platform Enterprise agent deployment
Gemini Enterprise Enterprise end-user workflows
Gemini app / AI Mode / Google Sheets Google AI Pro and Ultra consumer access where supported

Gemini 3.8 Flash Cyber is different. Fairwind access is restricted to approved trusted partners, including governments, national cyber authorities, critical infrastructure operators, core technology platforms, software maintainers, and selected defensive research groups. Organizations are vetted, must apply governance controls, and cannot share or resell access.

Compatibility Notes

For existing Gemini 3.7 Flash users, the first compatibility check is effort behavior. If your production system assumes predictable short outputs, fixed latency, or narrow tool-call loops, test 3.8 Flash before switching default traffic.

Key areas to validate:

  • Model ID change from gemini-3.7-flash to gemini-3.8-flash
  • Thinking levels, especially the lack of minimal support
  • Output-token usage at high effort levels
  • Tool-call reliability with your real function schemas
  • Computer Use behavior if your agent controls browsers or desktop-like environments
  • Prompt sensitivity for workflows tuned around 3.7 Flash response style

Pricing & Plans

Gemini 3.8 Flash uses the same introductory Gemini API pricing as 3.7 Flash through December 31, 2026:

Period Input Output
Introductory pricing through Dec 31, 2026 $0.75 / 1M tokens $3.75 / 1M tokens
Standard pricing from Jan 1, 2027 $1.50 / 1M tokens $7.50 / 1M tokens

The practical budget question is output volume. Google says 3.8 Flash may use more tokens on complex tasks to maximize performance, so high-volume teams should compare three numbers before migration: token price, tokens per completed task, and human retry or review time saved.

Flash Cyber Notes

Gemini 3.8 Flash Cyber is best treated as a restricted companion release, not a normal replacement for Gemini 3.8 Flash. Google and DeepMind describe it as a model for trusted defenders who need stronger vulnerability discovery and patching capability under controlled access.

Official Fairwind materials report several cyber-specific results:

  • CyberGym Pass@1: 86.2% for Gemini 3.8 Flash Cyber
  • Internal real-world vulnerability discovery benchmark: more than 70% success across complex codebases spanning 20 programming languages
  • CWE-Bench patching: 47.2% pass@1, close to Fable 5 at 47.8% but at lower reported cost
  • Chrome Security found 2.6x more correct patches than the best larger commercial models in Google's internal comparison
  • Wiz reported +7.5-9.7% higher recall at 2.3-5.2x lower cost than other leading frontier models in its internal benchmark

For most ToolWorthy readers, the useful decision is eligibility. Use public 3.8 Flash for general coding agents and enterprise workflows. Explore Fairwind only if your organization has an approved defensive security role and can satisfy Google's access controls.

Safety & Limitations

Gemini 3.8 Flash has the usual foundation-model risks, including hallucinations, occasional slowness, and timeout issues. The model card also notes that some domains may reflect a March 2026 knowledge cutoff while others may behave closer to January 2025, consistent with the Gemini 3 family.

The safety picture is mixed enough to test rather than assume. Google says Gemini 3.8 Flash does not materially increase Frontier Safety tracked capability levels versus 3.7 Flash, but the model card also notes a slight regression on automated multilingual safety relative to 3.7 Flash. For current facts, legal analysis, finance workflows, or security work, use retrieval, policy checks, and human review where the output carries real risk.

Who Should Upgrade / Who Should Wait

Upgrade to Gemini 3.8 Flash if:

  • You run coding agents where better task completion is worth extra reasoning and output tokens.
  • You use Gemini 3.7 Flash for agentic workflows and want to test higher-quality routing on difficult tasks.
  • Your product needs long-context multimodal input, Computer Use, code execution, grounding, and function calling in one hosted model.
  • You evaluate models by cost per completed task rather than only cost per million tokens.

Wait or stay on Gemini 3.7 Flash if:

  • Your workload is already optimized for short responses, low latency, and predictable token use.
  • You cannot absorb possible output-token growth at higher effort levels.
  • Your prompts and tool schemas are tightly tuned to Gemini 3.7 Flash behavior and have not been regression-tested.
  • You need the Cyber variant but do not qualify for the Fairwind Program.

FAQ

What is the model ID for Gemini 3.8 Flash?

Use gemini-3.8-flash. Google's Gemini API model page lists it as the stable model code for Gemini 3.8 Flash.

Is Gemini 3.8 Flash generally available?

Yes. Google DeepMind's Gemini model page lists Gemini 3.8 Flash as generally available, and the API page lists the stable model code.

How much does Gemini 3.8 Flash cost?

Introductory API pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026. Starting January 1, 2027, Google says $1.50 per 1M input tokens and $7.50 per 1M output tokens will apply.

How is Gemini 3.8 Flash different from Gemini 3.7 Flash?

Gemini 3.8 Flash builds on Gemini 3.7 Flash and is tuned to work harder on complex coding, agentic, and professional reasoning tasks. Google warns that this can mean more token use, so production teams should compare cost per successful task.

What is Gemini 3.8 Flash Cyber?

Gemini 3.8 Flash Cyber is a cybersecurity-tuned variant for vulnerability discovery, patching, and defensive research. It is available only to approved trusted defenders through Google's Fairwind Program.

Should I create a separate integration for Gemini 3.8 Flash Cyber?

Only if your organization is approved for Fairwind access. Most developers should integrate the public gemini-3.8-flash model first and treat Cyber as a restricted security program rather than a standard Gemini API model.

Sources

Release navigation

More tools to compare

MakersClaw icon

MakersClaw

Construct Computer icon

Construct Computer

TypingMind icon

TypingMind

Doubao icon

Doubao

Chert icon

Chert

Z.ai icon

Z.ai

Top alternatives

Related categories

From the blog

View all →

Track Gemini in ToolWorthy Weekly

Important tool updates, better alternatives, and selected AI signals in one weekly brief.

Weekly only. Unsubscribe anytime.