Google Gemini Update Prioritizes Efficiency
Google Gemini update launched cheaper, efficient models while the flagship Pro stayed in partner testing, shifting trader focus to AI serving costs.

KEY TAKEAWAYS
- Google launched three cheaper Gemini Flash variants to cut near-term AI serving costs for agent workloads.
- Gemini 3.6 Flash used up to 17.0% fewer output tokens than prior Flash, lowering token costs.
- Reports say a Frozen v2 server chip could boost tokens-per-watt efficiency roughly 6-10× with possible deployment around 2028.
HIGH POTENTIAL TRADES SENT DIRECTLY TO YOUR INBOX
Add your email to receive our free daily newsletter. No spam, unsubscribe anytime.
Alphabet Inc. said in a Google blog post on July 21 that the Google Gemini update introduces three cheaper, efficiency-focused models—Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—while its flagship Gemini 3.5 Pro remains in partner testing, signaling a near-term push to lower AI serving costs.
New Gemini Models and Pricing
Google launched three new Gemini variants in its product blog post. The Gemini 3.6 Flash is positioned as the high-throughput “workhorse” model for coding, knowledge work, and multimodal tasks. It balances speed with intelligence and incorporates developer and customer feedback from prior Flash releases. The company said 3.6 Flash reduces output-token usage by up to 17% compared with the previous Flash tier. It is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens in the API.
The 3.5 Flash-Lite model is the fastest and lowest-cost tier in the 3.5 family, aimed at high-volume agent workloads and smaller in-agent tasks. It is priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. Google said Flash-Lite will also be extended into Search. Both 3.6 Flash and 3.5 Flash-Lite are available through the Gemini app and developer channels including the Gemini API, Google AI Studio, Antigravity, and Android Studio. The earlier 3.5 Flash has been deprecated in favor of 3.6 Flash.
The third variant, 3.5 Flash Cyber, is a lightweight model tailored to cybersecurity use cases, focusing on detecting and patching software vulnerabilities. It will be introduced via a limited-access pilot for governments and trusted partners and integrated into tools such as CodeMender. Google described Flash Cyber as a cost-efficient option for enterprises with aggressive security and vulnerability-management needs.
Flagship Pro Status and Infrastructure Plans
Google said the Gemini 3.5 Pro flagship model is currently testing with partners and will be made broadly available as soon as it is ready. This delay follows earlier guidance at Google I/O that set expectations for a June rollout. Reporting through mid-July documented slipped internal targets, tightening scrutiny of the product timetable.
Separately, multiple reports indicate Alphabet is designing a Gemini-dedicated server chip codenamed Frozen v2. The chip aims to embed elements of Gemini’s architecture directly into silicon, targeting roughly 6–10 times gains in tokens-per-watt efficiency compared with current accelerators. Deployment is expected in Google data centers around 2028. The project remains an engineering effort rather than a commercial product.
Google has also begun what it calls its most ambitious pre-training run yet for Gemini 4, positioning it as the next frontier model beyond the 3.5 family.
Together, these moves reflect a two-track cost strategy: reducing near-term operating expenses with lighter, faster Flash models for agents and enterprise workloads while pursuing longer-term hardware engineering to materially lower the economics of serving large-scale generative AI. This approach sharpens competition in frontier modeling and cybersecurity applications.
Customers and enterprise partners will watch for a broader release date for the flagship model and early evidence that the new Flash tiers and any bespoke hardware translate into lower per-token costs at scale.
"Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready," Google said in its blog post.





