Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
7 min read·Updated September 3, 2026

Gemini Flash is Google DeepMind's workhorse agentic line, and it now carries the company's whole frontier story — Gemini 3.5 Pro is still unreleased. The newest release is 3.8 Flash, shipped September 2, 2026, the third Flash generation in six weeks. Its security sibling, 3.8 Flash Cyber, is not sold openly at all: access runs through a vetted-defender program called Fairwind. Introductory pricing runs at half rate through December 31, 2026.

Share

Listen to this overview

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand how the Gemini Flash line positions Google DeepMind's agentic stack against frontier alternatives
  • Identify the benchmarks that matter for agentic and coding workloads (DeepSWE, FrontierCode, AutomationBench, OSWorld-Verified)
  • Evaluate when Flash is the right default versus other frontier models

What Is Gemini Flash?

Gemini Flash is Google DeepMind's workhorse agentic line. Where earlier Flash generations were positioned as "speed-optimized siblings" to a flagship Pro model, Flash now leads the shipping lineup outright — Gemini 3.5 Pro was announced at Google I/O in May 2026 and still has no public release, leaving Flash as the model Google has staged as the daily driver for both consumer and developer surfaces.

The line is moving unusually fast. Gemini 3.8 Flash shipped on September 2, 2026 — the third Flash release in six weeks, following 3.7 Flash on August 13 and 3.6 Flash on July 21. That cadence is short enough that most teams will still be evaluating the previous generation when the next one lands.

Gemini 3.8 Flash — the current release

Google calls 3.8 Flash its "most intelligent workhorse model," and the framing is deliberately incremental: it claims improvements in software engineering, agentic tasks and multi-step reasoning while holding the speed and cost of 3.7 Flash. On DeepSWE v1.1, Google says it outperforms most larger frontier models — a comparative claim the company has not attached a public number to, so treat it as vendor positioning until independent results land.

The release also tightened safety rather than only capability. Google reports safeguards against chemical, biological, radiological and nuclear (CBRN) misuse and against cyber offense, plus improved robustness to prompt injection as measured on the Gray Swan benchmarks.

Availability: developers reach 3.8 Flash through Google AI Studio, Android Studio and Google Antigravity; enterprises through Gemini Enterprise; and consumers on Google AI Pro and Ultra subscriptions inside the Gemini app and Google Search.

Gemini 3.7 Flash — the previous release

Gemini 3.7 Flash went generally available on August 13, 2026. Its gains over 3.6 Flash were large for a three-week gap, and they concentrated in exactly the areas the previous generation was weakest:

BenchmarkWhat it measures3.6 Flash3.7 Flash
DeepSWE v1.1Long-horizon software engineering49.065.3
FrontierCode 1.1Production code quality34.4%43.6%
AutomationBenchEnterprise workflow automation17.0%30.4%
GDP.pdfPDF document comprehension22.0%34.0%
WebDev ArenaWeb development (Elo)15381588

Google attributes the jump to better instruction-following, improved multi-step planning, fewer wasted retries, and stronger tool use across Google Workspace applications. The AutomationBench result is the one to watch: nearly doubling enterprise workflow automation in a single generation is the difference between an agent that needs supervision and one that can be pointed at a routine process.

3.7 Flash reached Google AI Studio, Android Studio, the Gemini Enterprise Agent Platform, and Gemini Spark for AI Pro and Ultra subscribers across more than 160 countries, and remains available.

⚠️Warning

The introductory price is temporary and roughly doubles on January 1. Gemini 3.8 Flash and 3.8 Flash Cyber both run at 75 cents per million input tokens and three dollars seventy-five per million output through December 31, 2026 — half the standard rate. On January 1, 2027 they return to one dollar fifty per million input and seven dollars fifty per million output. If you are modeling unit economics off a pilot run this year, budget for the step change rather than the introductory number.

Gemini 3.6 Flash — the previous generation

Gemini 3.6 Flash shipped on July 21, 2026 as the direct successor to 3.5 Flash, and remains available.

The headline numbers from the launch — all gains over the model it replaces:

  • 49% on DeepSWE — software-engineering agent benchmark (up from 37% on 3.5 Flash)
  • 63.9% on MLE-Bench — machine-learning engineering tasks (up from 49.7%)
  • 83.0% on OSWorld-Verified — computer-use agent benchmark (up from 78.4%)
  • Roughly 17% fewer output tokens than 3.5 Flash for the same work (per the Artificial Analysis Index) — meaning it is not just more capable but less verbose and cheaper to run
  • $1.50 per million input tokens, $7.50 per million output — a workhorse price point well below flagship tiers

The framing is "frontier intelligence with action" — a model tuned specifically for agentic workflows that automate multi-week tasks in hours, with parallel subagent dispatch and long-horizon tool use as first-class behaviors rather than retrofitted capabilities.

Flash supports the same 1 million token context window as larger Gemini models and the same native multimodal capabilities (text, image, video, audio, code). It also supports over 100 simultaneous tool calls — enabling complex agentic workflows where the model orchestrates multiple external services in parallel.

Flash-Lite — the speed-and-cost floor

Gemini 3.5 Flash-Lite is the fastest model in the series at 350 output tokens per second, priced at 30 cents per million input tokens and two dollars fifty per million output. It still posts large agentic gains over the prior Lite generation (Terminal-Bench 2.1 54% up from 31%, SWE-Bench Pro 54.2%, OSWorld-Verified 74.0%), making it the right default for high-throughput, latency-sensitive workloads.

Flash Cyber — a model you cannot simply buy

Gemini 3.8 Flash Cyber is the security-tuned member of the line, and the more interesting thing about it is the distribution model rather than the benchmark. Google claims frontier-level performance on vulnerability detection and automated patching: on the CWE-Bench patching task it reports 47.2 percent on the first attempt, at what Google describes as significantly lower cost than competing models, and it posts frontier-level results on CyberGym. Paired with Google's CodeMender multi-agent system, the pitch is that a defender can go from a discovered vulnerability to a verified, deployment-ready patch in minutes.

You cannot sign up for it. Access runs through the Fairwind Program, announced alongside the model on September 2, 2026, and limited to three categories: government agencies and national cyber authorities, critical-infrastructure operators in healthcare, telecommunications, energy and financial services, and core technology platforms. Participants must accept operational conditions — restricting use to their own internal cybersecurity, incident-response or penetration-testing teams, and deploying protections such as multi-factor authentication. Google reports more than 650 partners globally, naming CrowdStrike, Palo Alto Networks, Snowflake and Wiz among them.

Why gate a defensive model at all? Because a model good enough to find and patch unknown vulnerabilities at scale is, by construction, good enough to find them for someone else. That is the same reasoning that led OpenAI to hold Astra behind a Critical cybersecurity threshold — two labs, two different mechanisms, one conclusion: frontier cyber capability now ships to vetted defenders rather than to whoever has a credit card. For anyone evaluating tools in this category, expect the access model rather than the benchmark to decide whether you can use them at all.

Above the Flash tier, the flagship Gemini 3.5 Pro is still partner-testing only as of August 2026 — target dates have slipped repeatedly and there is no public endpoint or pricing. Google has separately confirmed that pre-training has begun on Gemini 4, described as its "most ambitious pre-training run yet." The practical read is that Google has now shipped two Flash generations in three weeks while the Pro tier it announced in May has yet to appear at all.

Where 3.6 Flash Ships

  • Google Search AI Mode — Gemini 3.6 Flash is now the default model worldwide across 98 languages, available without an AI Pro or Ultra subscription
  • Google Antigravity 2.0 — Flash is the default routing target for the Manager View's parallel subagents, with 3.5 Pro slated to handle the highest reasoning effort once it ships
  • Gemini Spark — Google's new always-on personal agent, currently rolling to trusted testers, runs on 3.6 Flash for low-latency multi-step actions
  • Gemini CLI / Antigravity CLI — the terminal-side agentic surface is built to route between 3.5 Flash and 3.5 Pro automatically once Pro is available
  • Gemini API + Vertex AI — generally available for developers

🎯Tip

Try Gemini Flash: ai.google.dev — access via Google AI Studio (free tier available) or Vertex AI for production use

Pricing and Access

Access MethodPriceBest For
Google AI Studio (free tier)Free with rate limitsPrototyping, testing, and development
Gemini API (3.8 Flash, through Dec 31 2026)$0.75 in / $3.75 out per million tokensIntroductory rate — half the standard price
Gemini API (3.8 Flash, from Jan 1 2027)$1.50 in / $7.50 out per million tokensStandard rate once the introductory period ends
Vertex AIUsage-based with enterprise supportEnterprise deployments with SLAs and compliance
Gemini app (consumer)Included in Gemini subscriptionConsumer chat interface with Flash as default model

Flash's pricing advantage is its defining feature for production use. Gemini 3.8 Flash currently runs at 75 cents per million input tokens and three dollars seventy-five per million output — an introductory rate that lasts until December 31, 2026 and then doubles to one dollar fifty and seven dollars fifty. Note what did not change across the last two generations: Google held the price flat from 3.7 to 3.8 while claiming capability gains, so the practical upgrade path is simply to pin the newer version. Combined with roughly 17 percent fewer output tokens than the 3.5 generation for the same work, high-volume workloads see compound savings beyond raw per-token pricing alone. For latency-critical or ultra-high-volume work, 3.5 Flash-Lite drops the floor further to 30 cents per million input tokens.

Core Capabilities

Frontier Intelligence with Action — Agentic by Design

Gemini 3.6 Flash is tuned for long-horizon agentic tasks: planning across days or weeks of work, dispatching subagents in parallel, and self-correcting through tool-use loops. The DeepSWE (49%) and OSWorld-Verified (83.0%) scores measure this directly — both benchmarks test how reliably a model can drive a real software environment through sustained multi-step actions, not just answer a single coding question. This is the workload Google built Antigravity 2.0 and Gemini Spark around, and Flash is the model both products default to.

Multimodal Reasoning at Flash Speed

Flash keeps native processing of text, images, video, and audio in a single model, which makes it practical for real-time multimedia pipelines: ingest a research paper's figures, a meeting recording, and a chart in one prompt and get integrated analysis — at Flash's low latency rather than a slower flagship's.

Agentic Video Understanding (September 2026)

On September 1, 2026 Google shipped agentic video understanding across 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The change is architectural rather than cosmetic: instead of sampling a video at a fixed frame rate and reasoning over whatever those frames happened to capture, the model searches, scans and re-inspects specific segments across visual frames, audio and transcript — deciding where to look and going back for a closer pass. That is what makes sub-second moment retrieval, long-form search, anomaly detection and accurate counting work, all of which fixed-rate sampling handles badly.

Google reports up to 88 percent fewer tokens, up to 66 percent lower cost, and up to 7 percent better accuracy on video tasks — figures that are vendor-reported and stated as ceilings rather than typical results. The economics are the point: fixed-rate sampling makes long video expensive precisely because most sampled frames are irrelevant.

It is live now through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, with the Gemini app to follow and a YouTube question-answering feature after that.

1 million Token Context at Speed

Flash processes the same 1 million token context window as the rest of the Gemini family, but at significantly lower latency. Entire codebases, lengthy legal contracts, or multi-hour meeting transcripts fit in a single prompt with fast response times. For applications that process long documents at volume — legal discovery, financial analysis, research synthesis — Flash's speed advantage compounds with every document.

100+ Parallel Tool Calls

Flash issues over 100 simultaneous tool calls per request, which is what enables Antigravity 2.0's parallel-subagent pattern and Gemini Spark's always-on web monitoring. Where earlier agentic models had to serialize external calls into long synchronous chains, Flash dispatches them in parallel and reasons over the joint result set — a structural fit for the multi-week-to-multi-hour collapse Google highlighted at I/O.

Strengths

  • Top-tier agentic benchmarks: 3.7 Flash posted DeepSWE 65.3 and AutomationBench 30.4% — large gains over 3.6 Flash's 49.0 and 17.0% just three weeks earlier — and Google says 3.8 Flash improves further on software engineering, agentic tasks and multi-step reasoning
  • Capability gains at a flat price: 3.8 Flash holds 3.7's rate of 75 cents / three dollars seventy-five per million input/output tokens through December 31, 2026, plus roughly 17% fewer output tokens than the 3.5 generation — compounding throughput savings at scale
  • 1 million token context: Full million-token context window at faster processing speeds than alternatives in the same quality band
  • 100+ parallel tool calls: Backbone of Antigravity 2.0's parallel subagent dispatch and Gemini Spark's always-on monitoring
  • Native multimodal: Text, image, video, and audio processing in a single model without separate pipelines
  • Specialized siblings: 3.5 Flash-Lite for 350-tokens-per-second throughput and 3.8 Flash Cyber for vulnerability detection and automated patching, at 47.2 percent first-attempt on CWE-Bench
  • Default model in production-grade surfaces: Search AI Mode (98 languages), Antigravity 2.0, Gemini Spark, and Gemini CLI / Antigravity CLI

Limitations & Considerations

  • Pro reasoning ceiling: Gemini 3.5 Pro — announced at I/O in May 2026 and still partner-testing only as of August 2026, after three slipped target dates — is positioned for the deepest reasoning tasks; Google says Pro will outperform Flash on extended chain-of-thought workloads
  • Closed model: No self-hosting option — all inference runs through Google's API, which means data leaves your infrastructure
  • Google ecosystem dependency: Deepest integration is within Google Cloud (Vertex AI), Google AI Studio, and Google Workspace — less seamless outside the Google ecosystem
  • Rate limits on free tier: The free Google AI Studio tier has rate limits that are insufficient for production workloads — plan for paid API access
  • Introductory pricing expires: the half-rate pricing ends December 31, 2026 and roughly doubles on January 1 — a pilot costed this year will not reflect next year's bill
  • Release cadence outruns evaluation: three Flash generations shipped inside six weeks, which is faster than most teams can qualify a model; pinning an explicit version is safer than tracking the default
  • The Cyber variant is not purchasable: 3.8 Flash Cyber is restricted to Fairwind Program participants — government cyber authorities, critical-infrastructure operators and core platforms — so for most organizations it is a capability to be aware of rather than one to plan around
  • 3.8 Flash has no published headline number: Google's DeepSWE claim for 3.8 is comparative rather than numeric, so the concrete generational evidence on this page still stops at 3.7
  • Benchmark vintage: The headline benchmarks (DeepSWE, FrontierCode, AutomationBench) are recent, rapidly evolving, and vendor-reported; expect leaderboard movement as other labs respond

Best Use Cases

TaskWhy Gemini Flash
Long-horizon agentic workflowsDeepSWE + OSWorld-Verified leadership; tuned for multi-day task automation
Parallel-subagent development100+ simultaneous tool calls + Antigravity 2.0 integration; multi-week tasks in hours
Real-time agentic assistantsGemini Spark's underlying model — low-latency always-on monitoring and action
Search-side AI applicationsDefault for Google Search AI Mode in 98 languages — proven scale
Document processing pipelines1 million context + fast speed = process thousands of long documents efficiently
Ultra-high-throughput, low-latency3.5 Flash-Lite at 350 tokens per second — the speed-and-cost floor of the Flash tier
Vulnerability detection and repair3.8 Flash Cyber inside CodeMender — Fairwind Program participants only
Cost-sensitive production AI$1.50 / $7.50 per million tokens plus 17% fewer output tokens = compound throughput savings

When to choose alternatives:

  • Maximum reasoning depth → a closed frontier flagship from Anthropic or OpenAI today; Gemini 3.5 Pro is still not shipped, so it is not an option you can actually pick
  • Self-hosted deployment → Gemma 4, GPT-OSS, or Mistral Medium 3.5 (open-weight models you can run locally)
  • OpenAI ecosystem → OpenAI's GPT-5 series or Codex (if you are already invested in OpenAI tooling)
  • Source-cited research → Perplexity (built specifically for research with citations)

Getting Started

  1. Go to ai.google.dev and create a Google AI Studio account (free tier available)
  2. Generate an API key from the Google AI Studio console
  3. Select gemini-3-8-flash as your model for the newest generation (or gemini-3-5-flash-lite for maximum throughput)
  4. Test with an agentic task (multi-step tool use) to feel the DeepSWE-grade workflow behavior
  5. Pin an explicit model version in production rather than tracking the default — three Flash generations shipped inside six weeks, and the default moved underneath anyone who had not pinned
  6. Cost your pilot at the January 2027 standard rate, not the introductory one, so the price step does not surprise your budget
  7. For production, set up Vertex AI for enterprise-grade SLAs, monitoring, and compliance features

🎯Tip

Decision framework for Flash vs Pro: Use Gemini 3.8 Flash as your default and treat Pro as a tier that does not yet exist — Gemini 3.5 Pro has been in partner testing since May 2026 with no public release, so "wait for Pro" is not a plan you can schedule around. With Flash the default behind Search AI Mode, Antigravity 2.0, and Gemini Spark, Google has effectively staged Flash as both the everyday workhorse and, for now, the ceiling.

Key Takeaways

  • Gemini 3.8 Flash is the current release, shipped September 2, 2026 — the third Flash generation in six weeks — and reaches AI Studio, Android Studio and Google Antigravity for developers, Gemini Enterprise for organizations, and AI Pro and Ultra subscribers in the Gemini app and Google Search
  • The gains arrive at a flat price: Google claims better software engineering, agentic work and multi-step reasoning for 3.8 while holding 3.7's rate, so upgrading is mostly a matter of pinning the newer version
  • The measured generational evidence still comes from 3.7, whose jump over 3.6 was unusually large for a three-week interval: DeepSWE 49.0 to 65.3, AutomationBench 17.0% to 30.4%, FrontierCode 34.4% to 43.6%
  • Pricing halves, then unhalves: 75 cents / three dollars seventy-five per million input/output tokens through December 31, 2026, returning to one dollar fifty and seven dollars fifty on January 1 — cost pilots at the later number
  • 3.8 Flash Cyber is gated, not sold: 47.2 percent first-attempt on CWE-Bench patching, reachable only through the new Fairwind Program for government cyber authorities, critical-infrastructure operators and core platforms, with more than 650 partners including CrowdStrike, Palo Alto Networks, Snowflake and Wiz — the same vetted-defender pattern OpenAI applied to Astra
  • 3.5 Flash-Lite remains the speed-and-cost floor at 350 tokens per second
  • Flash is the default behind Google Search AI Mode (98 languages, no AI Pro subscription required), Antigravity 2.0's Manager View, Gemini Spark, and the Gemini CLI / Antigravity CLI stack
  • Gemini 3.5 Pro remains partner-testing only after repeated slipped dates, while pre-training has begun on Gemini 4 — Google has now shipped two Flash generations in three weeks without releasing the Pro tier it announced in May, so Flash is both the daily driver and the practical ceiling

Keep track of the tools you’re evaluating

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub — sample data.

📰Gemini Flash in the News

View all 6 Gemini Flash stories in Top AI Stories

Other tools in Foundation Models & Open Source (12 of 75)

Show 7 more →

Other tools from Google

Show 18 more →
🧭Recommended for you

Optional detours — these connect to what you just read, and your next lesson will be waiting.