Google's cheaper, faster workhorse Flash tier — a text-and-reasoning model built for high-volume, low-cost drafting and agentic work.
Last verified · 2026-07-21 · by Moe Ameen
Gemini 3.6 Flash is Google DeepMind's updated "workhorse" Flash model, released July 21, 2026 alongside two siblings: Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. These are text and multimodal reasoning models — the tier you reach for when you need fast, cheap output at volume, not a frontier model for the hardest problems. Note this is the Flash reasoning line, not Gemini Omni Flash, which is Google's separate video-generation model.
Gemini 3.6 Flash's headline is efficiency. Google says it delivers better coding, knowledge work, and multimodal performance than 3.5 Flash while consuming about 17% fewer output tokens (per the Artificial Analysis Index) and taking fewer reasoning steps and tool calls on multi-step tasks. It is priced at $1.50 per million input tokens and $7.50 per million output tokens — down from roughly $9 per million output on 3.5 Flash. Its knowledge cutoff advances to March 2026, and its computer-use score on OSWorld-Verified rises to 83% from 78.4%.
Gemini 3.5 Flash-Lite is the low-latency, high-throughput tier — Google points it at agentic search and document processing, running around 350 output tokens per second at $0.30 per million input and $2.50 per million output. Google says it beats the older 3.1 Flash-Lite and even outperforms Gemini 3 Flash on some benchmarks. Gemini 3.5 Flash Cyber is a narrow, security-focused model fine-tuned to find and fix code vulnerabilities; it ships only through Google's CodeMender agent as a limited-access pilot for governments and trusted partners, and is not publicly priced.
Gemini 3.6 Flash and 3.5 Flash-Lite are available in the Gemini app and through the Gemini API (Google AI Studio, Android Studio) and Gemini Enterprise, with 3.6 Flash also in Google Antigravity and Flash-Lite rolling into Google Search. Google framed the trio as a stopgap while Gemini 3.5 Pro finishes partner testing, and confirmed it has started pre-training Gemini 4.
Gemini 3.6 Flash and 3.5 Flash-Lite are excellent at one thing that matters to creators: producing a lot of raw text, cheaply. Ask for thirty hook variations or a blog outline and you get it in seconds for pennies. But raw text out of a chat window is not a content week. It has no brand voice guardrails, no design, no video, no schedule, and no way onto your nine platforms — you still copy-paste it into every app by hand. Kompozy is the engine that closes that gap.
Concretely: take an idea or a rough draft you brainstormed with Gemini Flash and drop it into Kompozy. Kompozy's copy engine (Claude and OpenAI, governed by your Persona Brief and banned-word filters — with bring-your-own-key on the Founding tier) rewrites it in your actual voice, then generates the finished formats Flash can't touch: Persona Shorts and HeyGen avatar video, Carousel Posts and Persona Tweets rendered pixel-exact through HyperFrames, Photo Posts, Quote Graphics, blog articles, and email newsletters. From there Kompozy schedules and publishes the whole set to Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads, plus Mailchimp and your blog. Flash gives you cheap raw material; Kompozy turns it into on-brand, multi-format, scheduled output — the part a language model was never built to do.
They are three fast, cost-efficient models Google DeepMind released on July 21, 2026. Gemini 3.6 Flash is the updated general "workhorse" tier for coding, knowledge work, and multimodal tasks; 3.5 Flash-Lite is a lower-latency, high-throughput tier for agentic search and document processing; and 3.5 Flash Cyber is a narrow security model that finds and fixes code vulnerabilities, available only through Google's CodeMender pilot.
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens — cheaper than 3.5 Flash, and Google says it uses about 17% fewer output tokens per task. Gemini 3.5 Flash-Lite is cheaper still at $0.30 per million input and $2.50 per million output. 3.5 Flash Cyber is not publicly priced.
No. Gemini 3.6 Flash is a text and multimodal reasoning model — it writes and reasons. Gemini Omni Flash is a separate Google model that generates and edits video. They share the "Flash" name for the fast, cost-efficient tier but do different jobs.
The Flash reasoning models write text and reason over multimodal inputs; they are not image or video generators. For finished visual content — persona and avatar video, carousels, quote graphics, photo posts — you use a generation engine like Kompozy, which then also schedules and publishes across platforms.
Brainstorm or draft with Gemini Flash, then bring the ideas into Kompozy. Kompozy rewrites in your brand voice via the Persona Brief, generates video, image, carousel, blog, and newsletter formats, and schedules and publishes them across nine social platforms plus email and blog — the packaging and distribution a language model does not do.