Three fast models land July 21 while the flagship Gemini 3.5 Pro stays in partner testing — and Google says pre-training for Gemini 4 has begun.
2026-07-21 · by Moe Ameen
On July 21, 2026, Google DeepMind released three new models in its cost-efficient Flash line: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The launch arrived as a stopgap — the flagship Gemini 3.5 Pro is still in partner testing, and Google confirmed it has started its "most ambitious pre-training run yet" for Gemini 4.
Gemini 3.6 Flash is the updated general-purpose "workhorse" model. Google says it improves on 3.5 Flash across coding, knowledge work, and multimodal tasks while consuming about 17% fewer output tokens (per the Artificial Analysis Index) and taking fewer reasoning steps on multi-step workflows. It is priced at $1.50 per million input tokens and $7.50 per million output tokens — down from roughly $9 per million output on 3.5 Flash — with its knowledge cutoff advanced to March 2026 and its OSWorld-Verified computer-use score rising to 83% from 78.4%.
Gemini 3.5 Flash-Lite targets low-latency, high-throughput work like agentic search and document processing, running around 350 output tokens per second at $0.30 per million input and $2.50 per million output. Google says it beats the older 3.1 Flash-Lite and outperforms Gemini 3 Flash on several benchmarks. Gemini 3.5 Flash Cyber is a narrow security model fine-tuned to find and fix code vulnerabilities; it ships only through Google's CodeMender agent as a limited-access pilot for governments and trusted partners and is not publicly priced.
Gemini 3.6 Flash and 3.5 Flash-Lite are available immediately in the Gemini app and through the Gemini API (Google AI Studio, Android Studio) and Gemini Enterprise, with 3.6 Flash also in Google Antigravity and Flash-Lite rolling into Google Search.
Every Flash price cut lands the same way for creators: the cost of raw words drops, and the real work — turning words into on-brand posts that ship across every platform — stays exactly where it was. That gap is where Kompozy lives. Use Gemini 3.6 Flash or Flash-Lite to brainstorm angles and rough out a week of ideas cheaply, then bring them into Kompozy, where the copy engine (Claude and OpenAI, governed by your Persona Brief) rewrites them in your voice and generates the formats a language model can't: Persona Shorts and HeyGen avatar video, Carousel Posts and Persona Tweets rendered pixel-exact through HyperFrames, Photo Posts, Quote Graphics, blogs, and newsletters.
The point isn't to pick a model — it's to stop the win from evaporating at the copy-paste step. Kompozy schedules and fans that finished set to Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads, plus Mailchimp and your blog, from one queue with autopilot and a per-post review pipeline. Cheaper drafting only helps if the drafts become published content — that's the half Kompozy owns.
Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, 2026. Gemini 3.6 Flash and 3.5 Flash-Lite are available in the Gemini app and via the Gemini API; Flash Cyber ships only through the CodeMender pilot.
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, down from roughly $9 per million output on 3.5 Flash, and Google says it uses about 17% fewer output tokens per task — a compounding cost reduction on high-volume work.
No. Gemini 3.6 Flash and 3.5 Flash-Lite are text and multimodal reasoning models — they write and reason but do not generate images or video. To turn their drafts into finished visual posts and publish them, creators use a generation-and-publishing engine like Kompozy.
Gemini 3.5 Flash Cyber is a security-focused model fine-tuned to detect and patch code vulnerabilities at a lower price per token than larger models. It is available only through Google's CodeMender agent as a limited-access pilot for governments and trusted partners.