Gemini 3.7 Flash is Google’s August 2026 update to Gemini 3.6 Flash, designed for coding, knowledge work and multi-step agent workflows. Google reports meaningful gains on several software-engineering benchmarks without changing the introductory API price, but those gains vary by task and should not be interpreted as a universal performance improvement.
The short answer: choose Gemini 3.7 Flash when tool use, repository-scale coding, web development or long multi-step workflows matter. If Gemini 3.6 Flash already meets your quality target, test 3.7 against a representative workload before changing a production system.
Gemini 3.7 Flash vs 3.6 Flash: quick verdict
- Best reason to upgrade: stronger official results on long-horizon software engineering, terminal tasks and web development.
- What stays familiar: the model is based on Gemini 3.6 Flash and retains a 1M-token input window.
- Pricing: Google lists an introductory price of $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026.
- Important caution: benchmark results measure defined test conditions. They do not guarantee the same improvement for every application, prompt or tool environment.
Gemini 3.7 Flash specifications
The model supports function calling, code execution, file search, structured output, URL context and search grounding. Computer use is available in preview, so production teams should treat browser or interface actions as supervised workflows rather than assuming perfect execution.
Official specification: Gemini 3.7 Flash in the Gemini API documentation
Official benchmark comparison
The following figures are Google’s reported results as of August 2026. Improvements shown as “points” are absolute percentage-point changes, not relative percentages.
Source and evaluation details: Google DeepMind Gemini 3.7 Flash model card
What the benchmark gains mean
Software engineering
DeepSWE evaluates long-horizon software-engineering work. The increase from 48.6% to 65.3% is the clearest reason to test 3.7 Flash for repository-level debugging, multi-file changes and tasks that require several tool calls. It is still a benchmark result, not a promise that 65.3% of your own issues will be solved.
Terminal and agent workflows
The Terminal-bench improvements suggest better performance when an agent must inspect files, run commands and recover from roadblocks. Real results will also depend on the surrounding agent, tool permissions, time limits and quality of the task description.
Web development
The 50-point Code Arena gain indicates improved web-development output under that evaluation. Teams should still test accessibility, responsive behavior, framework conventions and visual accuracy with their own design system.
What changed in Gemini 3.7 Flash?
Google describes 3.7 Flash as an algorithmic improvement to the reasoning foundation of 3.6 Flash. It is intended to follow instructions more reliably, clarify intent when needed and put more effort into multi-step planning and tool calls. The underlying relationship matters: this is an iteration on 3.6 Flash, not an unrelated architecture.
For developers, the practical change is less about a single headline score and more about cost per accepted task. A model that completes a workflow with fewer retries can be cheaper even when its per-token price is unchanged. That benefit should be measured with task completion rate, latency, review time and total tokens—not benchmark scores alone.
Multimodal inputs and long context
Gemini 3.7 Flash accepts text, images, audio, video and PDF input, with text output. Its 1,048,576-token input limit makes it suitable for large document collections and repositories, but a large window does not remove the need for retrieval, chunking or relevance filtering. Sending unnecessary context can increase latency and cost while making the prompt less focused.
The 65,536-token output limit supports long reports and larger code responses. In production, shorter structured outputs are usually easier to validate, retry and store.
Gemini Spark and product availability
Google says Gemini 3.7 Flash powers Gemini Spark for Google AI Pro and Ultra subscribers in more than 160 countries. Developers can access the model through Google AI Studio and the Gemini API, while enterprise availability includes Google’s enterprise agent products.
Google’s announcement and current availability details
Pricing and the real cost of an agent
Google’s introductory price is $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Because this is explicitly introductory pricing, verify the live pricing page before forecasting costs beyond that period.
For a realistic comparison, track these four metrics on the same task set:
- Successful task completion rate.
- Median input and output tokens per accepted result.
- Latency and number of tool retries.
- Human review and correction time.
Should you upgrade from Gemini 3.6 Flash?
Upgrade or start a controlled test if your workload involves software repositories, terminal tools, long multi-step plans, UI generation or complex document processing. Keep 3.6 Flash temporarily if it already meets a stable production target and migration risk is more important than incremental quality.
A safe migration process
- Build a test set of 20–50 real tasks, including common cases and known failures.
- Run both models with the same prompts, tools, limits and evaluation criteria.
- Measure accepted-task rate, latency, token cost and human correction time.
- Review safety-sensitive actions separately, especially computer-use workflows.
- Roll out gradually and retain a fallback until the new model is stable.
Limitations
Gemini 3.7 Flash can hallucinate, misunderstand instructions or make incorrect tool decisions. Google also notes possible slowness and timeout issues. Computer use remains a preview capability. Require confirmation before consequential actions, validate generated code and treat retrieved web content as untrusted input.
Related Premier Solutions guides
- For the wider model architecture and product family, read .
- Compare the previous release in .
- For agent-first development workflows, continue with .
- Developers working from the terminal can use .
- For a local open-model alternative, compare .
Frequently asked questions
Is Gemini 3.7 Flash generally available?
Yes. Google lists gemini-3.7-flash as a stable, generally available model.
Does it have a 1M-token context window?
Yes. The documented input limit is 1,048,576 tokens and the output limit is 65,536 tokens.
Is 3.7 Flash always better than 3.6 Flash?
No. It scores higher on the official benchmarks listed above, but performance depends on the task, prompt, tools and evaluation method. Test both models on your own workload.
Is the listed price permanent?
No. Google describes $0.75 per million input tokens and $3.75 per million output tokens as introductory pricing available through the end of 2026.
Resumen
Gemini 3.7 Flash is a focused upgrade for coding and agentic work. Its strongest official gains appear in long-horizon software engineering, terminal tasks and web development. The best migration decision should combine those results with controlled testing of quality, latency, total token usage and review time.




