Hey there! 👋
Welcome back to SavvyMonk, your one-stop for AI and tech news that actually matters.
Google put out three new Gemini models at once, and the interesting part is not that they are more capable, it is that the flagship of the batch costs less than the model it replaces. Tucked into the same announcement is the first real signal on what comes after.
Let's get into it.
TODAY'S DEEP DIVE
Gemini 3.6 Flash Cuts Output Token Use Seventeen Per Cent And Costs Less Than The Model It Replaces
On 21 July 2026, Google introduced three new Gemini models built around a single argument, that the teams shipping production AI agents care most about cost, speed and reliability rather than the top of a benchmark chart. The centrepiece is Gemini 3.6 Flash, which the company calls its workhorse model, sitting alongside a faster Gemini 3.5 Flash-Lite and a specialised Gemini 3.5 Flash Cyber.
The first two went live the day of the announcement, which came from Tulsee Doshi, a senior director of product management writing for the Gemini team. The framing throughout is agentic scale, meaning models that get run millions of times a day inside automated workflows where every wasted token and every extra second compounds into real money.
The Efficiency Gains
The core claim behind 3.6 Flash is that it does more while spending less. On the independent composite that measures model efficiency, it consumes seventeen per cent fewer output tokens than 3.5 Flash to reach the same result, and on a coding benchmark named DeepSWE that reduction climbs as high as sixty five per cent in the best case.
Fewer tokens means fewer reasoning steps and fewer tool calls to finish a multi step task, which is the figure that actually lands on a monthly invoice. And the quality did not slip to get there.

On DeepSWE the model scored forty nine against the earlier thirty seven, on MLE Bench it reached 63.9 against 49.7, on OSWorld-Verified it hit 83.0 against 78.4, and on a broader knowledge work benchmark it posted 1421 against 1349. Computer use is now a built in tool through the Gemini API rather than a separate integration, and early customers including Hebbia and Harvey said the model handled document parsing, chart analysis and report drafting more reliably than the version before it.
The Price Cut
The number that makes the rest matter is the price. Google set 3.6 Flash at one dollar fifty per million input tokens and seven dollars fifty per million output tokens, below what 3.5 Flash cost, so the efficiency gain and the lower sticker price stack on top of each other.
That combination is unusual, because the normal pattern in this industry is that a better model costs more and a cheaper model does less. Here the company is claiming both directions at once, and the reason is competitive rather than generous. When a developer picks a model to run inside an agent that fires thousands of times an hour, the decision is made on cost per completed task, and Google is trying to win that decision before a rival locks the developer in.

The model card also details stronger safeguards around chemical, biological and cyber misuse, with the company saying it hardened the model against jailbreaks while training it to refuse fewer legitimate requests.
The Faster Sibling
The second model, Gemini 3.5 Flash-Lite, is built for a different job, the high volume low latency work where throughput matters more than depth. It runs at three hundred and fifty output tokens a second, priced at thirty cents per million input tokens and two dollars fifty per million output tokens, which puts it at the cheap end of the range for developers pushing large amounts of traffic through tasks like agentic search and document processing.

The jumps over the previous generation are large, with Terminal-Bench 2.1 moving to fifty four from thirty one, a long context test reaching 72.2 from 60.1, and a real world task benchmark climbing to 1140 from 642.
More striking, Flash-Lite now beats the older and larger 3 Flash on several evaluations, including a software engineering test at 54.2 against 49.6 and computer use at 74.0 against 65.1, which means the smaller cheaper model has overtaken a bigger predecessor.

It is rolling out inside Google Search as well, and early users such as Palo Alto Networks and Ramp pointed to the mix of speed and cost as the draw.
The Cyber Model Google Is Keeping On A Leash
The third release is the one Google is deliberately holding back. Gemini 3.5 Flash Cyber is built on 3.5 Flash and fine tuned to find and fix software vulnerabilities, and it runs inside CodeMender, a code security agent where several copies of the model work together to produce one combined report.
On a well known security benchmark called CyberGym it reaches competitive performance at the frontier while costing far less per token than larger models. But the model will not be sold openly. Google is treating it as dual use, meaning the same skill that patches a hole can be turned to finding one to exploit, so access is restricted to governments and trusted partners through a limited pilot that opens later.

The reasoning is that defenders need the head start more than the open market needs the tool, and that a model this good at locating flaws is safer kept inside a controlled programme than handed to anyone with a credit card.
The Signal About Gemini 4
The most forward looking line in the whole announcement is a single sentence that is easy to miss. Google confirmed it has begun its most ambitious pre-training run yet, for Gemini 4, and that Gemini 3.5 Pro is already testing with partners ahead of a broader release.
Pre-training is the expensive foundational stage where the base model is actually built, so saying the Gemini 4 run has started is a concrete milestone rather than a vague promise, and slipping it into a Flash update is a way of setting expectations without making it the headline.
For anyone tracking the pace of the frontier, that one clause carries more weight than the benchmark tables around it, because it tells you the next generation is already consuming compute while the current generation ships.
The Bottom Line
The efficiency story is the real story here, and it is a good one, because a model that is both cheaper and better is exactly what the people building agents have been asking for. The Cyber release is worth watching for the opposite reason, since a company choosing not to sell its most capable security model tells you how seriously the risk is being taken.
And the buried Gemini 4 line is the tell that matters most, a quiet confirmation that the biggest training run Google has ever attempted is already underway while everyone reads the Flash pricing.
AI PROMPT OF THE DAY
Category: Cost Modelling
"Act as an infrastructure cost analyst. I am choosing an AI model to run inside a production agent that fires [estimate calls per day] times a day, with an average of [input token estimate] input tokens and [output token estimate] output tokens per call. Given two candidate models with these prices [paste input and output prices for each], calculate my projected monthly and annual spend for each, show the break-even point where a pricier but more efficient model becomes the cheaper choice, and list the three variables most likely to move my real bill away from this estimate so I know what to monitor once it is live."
ONE LAST THING
For most of the last two years the story of each new model was the score at the top of the chart, and the price was a footnote you found later. This release flips that order, leading with tokens saved and dollars cut and treating the benchmark wins as the supporting cast.
That shift is worth noticing, because it is what happens when a technology moves from the demo stage into the stage where people actually run it at scale, and the question stops being how clever the model is and becomes how little it costs to let it work all day.
Hit reply, I read every response.
See you in the next one.
— Vivek
P.S. Know a developer or founder deciding which model to build their agents on? Forward this to them. They can subscribe at https://savvymonk.beehiiv.com/

