AI News

Google’s New Gemini Flash Models Chase Efficiency, Not Power

Google Gemini gets new flash models for users
Image Source - Google

Google has released three new Gemini models built around a single idea that is cheaper and leaner beats bigger. The lineup, announced on July 21 by Senior Director of Product Management Tulsee Doshi, covers a workhorse model, a low-cost speed model, and a locked-down security model. The pitch is efficiency for companies running AI agents at scale.

Google is not claiming a leap in raw intelligence here. It is claiming its models now do more work with fewer tokens. For anyone paying per token to run an AI agent, that is the number that lands on the invoice.

What Gemini 3.6 Flash Changes for Developers

Gemini 3.6 Flash is the headline release. Google positions it as the everyday model for coding, knowledge work, and multimodal tasks. The company says it consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. On one benchmark, DeepSWE by Datacurve, Google claims the reduction reaches 65%.

Price dropped alongside it. Google lists 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Fewer tokens at a lower rate compounds into a lower cost per task. That is the actual product being sold.

The quality claims sit on Google’s own testing. The company reports 3.6 Flash hitting 49% on DeepSWE against 37% for 3.5 Flash, 63.9% on MLE Bench against 49.7%, and 83.0% on OSWorld-Verified against 78.4%. It names Hebbia and Harvey as customers finding it strong at document parsing and data analysis. None of these figures has been independently confirmed.

Why Gemini 3.5 Flash-Lite Undercuts Google’s Own Models

The more revealing release is Gemini 3.5 Flash-Lite. It is the cheapest of the group at $0.30 per million input tokens and $2.50 per million output tokens. Google says it runs at 350 output tokens per second, making it the fastest model in the 3.5 series.

Here the story gets awkward for Google. The company reports that 3.5 Flash-Lite beats the older 3 Flash on several agentic tests, including SWE-Bench Pro at 54.2% against 49.6% and OSWorld-Verified at 74.0% against 65.1%. A budget model outrunning last generation’s mid-tier model is the whole thesis of this launch. It also quietly tells developers they may be paying for more model than they need.

Flash-Lite is rolling out in the Gemini app and Google Search. That reach explains the aggressive pricing. Google wants this model handling high-volume traffic where cost per query decides whether a feature ships.

How Gemini 3.5 Flash Cyber Handles Security Work

The third release is the one to read carefully. Gemini 3.5 Flash Cyber is a specialised model tuned for finding and patching security vulnerabilities, deployed inside Google’s CodeMender agent. Google says it reaches “competitive performance at the frontier” on the CyberGym benchmark.

That claim arrives without a number. For a launch this precise on pricing and token counts, the absence of a CyberGym score on its security model is worth noticing. Google frames the tool around a real problem, that AI now finds vulnerabilities faster than teams can fix them, but the frame is Google’s and the evidence is not yet public.

Google is restricting access. The model will go only to governments and trusted partners through a limited pilot, citing the dual-use risk of offensive-security tooling. That caution is reasonable. It also means outside researchers cannot test the frontier claim.

How the New Gemini Models Compare to OpenAI and Anthropic

On price, Google is fighting at the budget end. Flash-Lite’s $0.30 and $2.50 rates sit against Anthropic’s Claude Haiku 4.5 at $1 and $5, and OpenAI’s mini and budget tiers in a similar range. The 3.6 Flash rates land below Anthropic’s Sonnet tier and near OpenAI’s mid-tier models.

The catch is that every rival benchmark in this launch is missing. Google compares 3.6 Flash and Flash-Lite against its own older Gemini models, not against Claude or GPT. That is a common tactic and it leaves the real competitive question open. Cheaper than your own last model is not the same as better than the model a developer might switch to.

Google also used the announcement to tease what comes next. It says Gemini 3.5 Pro is testing with partners and that pre-training has begun on Gemini 4. Neither has a date. Both read as signals to keep developers from shopping elsewhere while the Flash models carry the near-term revenue.

Abhijay Singh Rawat
Abhijay is the News Editor at TimesofAI, who loves to follow up on the latest tech and AI trends. After office hours, you would find him either grinding competitive ranked games, or trek up his way in the hills of Uttarakhand.
You may also like
More in:AI News