Google rolled out Gemini 3.7 Flash on August 13, calling it its smartest workhorse yet for coding and AI agents. The launch comes just three weeks after Gemini 3.6 Flash, and Google’s model card describes it as an algorithmic improvement built on 3.6 rather than a new architecture, targeting developers who are building software, web apps, knowledge systems, and production agents. Pricing is a big part of Google’s strategy: Gemini 3.7 Flash costs $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, with those promotional rates running to December 31, 2026, after which they double to $1.50 and $3.75 rises to $7.50, matching what 3.6 Flash launched at.
That price tag puts Gemini 3.7 Flash in direct competition with OpenAI’s GPT-5.6 Terra and Anthropic’s Claude Sonnet 5. All three models focus on practical work like coding, reasoning through problems, using tools, and managing multi step tasks rather than just casual daily conversations. Still, they differ in strengths, pricing, and where each one leads.
Gemini 3.7 Flash: What Google’s New Model Brings
Google is pitching Gemini 3.7 Flash as a model for coding and agents, promising a step up from Gemini 3.6 Flash in software engineering, web development, knowledge work, and business automation. The company says this version is faster at getting past roadblocks, clarifies what users want, follows instructions better, and deals with complex planning and tool use. Google also claims these upgrades mean less manual supervision and fewer retries during engineering tasks.
Gemini 3.7 Flash’s main improvements include:
- Coding: FrontierCode 1.1 Main rises to 43.6% from 34.4%
- Web development: The WebDev Arena Elo score rises to 1588, a step up from Gemini 3.6 Flash’s 1538. Google says it can deliver more complete apps and layouts in fewer prompts.
- Knowledge work: On GDP.pdf, Gemini scores 34.0%, compared with 22.0% for 3.6 Flash. On AutomationBench, it jumps to 30.4% from 17.0%.
- Pricing: $0.75 per 1 million input tokens and $3.75 per 1 million output tokens through December 31, then $1.50 and $7.50 from January 1, 2027.
- Availability: Gemini Spark now runs on 3.7 Flash for Google AI Pro and Ultra subscribers in over 160 countries.
- Safety: Google says this version has improved safeguards, including protections against CBRN and cyber offense misuse.
A caveat on the numbers. FrontierCode is produced by Cognition, the company behind the Devin coding agent, and DeepSWE is run by Datacurve, an AI data company. Neither is independently audited. Sanchit Gogia of Greyhound Research put it plainly, calling these vendor benchmark claims until the model builds up independent production evidence.
Some early external testing supports the direction. Niko Grupen, head of applied research at Harvey, reported a 2.6-point all-pass improvement over 3.6 Flash on Legal Agent Bench. Arena’s preliminary leaderboard placed 3.7 Flash at 1588 Elo after 2,544 votes, though its confidence interval still overlaps nearby models.
Developers and enterprise customers can access Gemini 3.7 Flash through the Gemini API, Google AI Studio, Android Studio, Google Antigravity, and Gemini Enterprise.
Today we're introducing Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents.
— Google (@Google) August 13, 2026
This model brings substantial gains across software engineering, web development, and complex knowledge work.
Now through the end of the year, Gemini 3.7 Flash is available… pic.twitter.com/RSCBDipjKn
Where Each Model Wins on Google’s Own Benchmarks
Google’s own comparison table puts the three models on the same benchmarks, and the results are split rather than one-sided. Gemini 3.7 Flash leads on FrontierCode 1.1 Main at 43.6%, narrowly ahead of Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. It leads more clearly on AutomationBench at 30.4%, against 23.6% for Terra and 10.7% for Sonnet 5, and on web development, where its Code Arena Elo of 1588 beats Sonnet 5’s 1541 and Terra’s 1523.
Terra holds its ground elsewhere. It leads DeepSWE v1.1 at 69.6% to Gemini’s 65.3%, Terminal-bench 2.1 at 87.4% to 85.8%, and also comes out ahead on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 tops Agent’s Last Exam, the multimodal desktop and operating-system evaluation, with a 33.3% pass rate against Gemini’s 26.3%.
In other words, Google’s own numbers don’t show 3.7 Flash displacing the higher-priced models across the board. They show a cheaper model that has become genuinely competitive in coding and agent work.









