// AI NEWS · MODEL RELEASE

Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — a Cheaper, More Token-Efficient Flash Tier

Three fast models land July 21 while the flagship Gemini 3.5 Pro stays in partner testing — and Google says pre-training for Gemini 4 has begun.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

2026-07-21 · by Moe Ameen

What happened

On July 21, 2026, Google DeepMind released three new models in its cost-efficient Flash line: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The launch arrived as a stopgap — the flagship Gemini 3.5 Pro is still in partner testing, and Google confirmed it has started its "most ambitious pre-training run yet" for Gemini 4.

Gemini 3.6 Flash is the updated general-purpose "workhorse" model. Google says it improves on 3.5 Flash across coding, knowledge work, and multimodal tasks while consuming about 17% fewer output tokens (per the Artificial Analysis Index) and taking fewer reasoning steps on multi-step workflows. It is priced at $1.50 per million input tokens and $7.50 per million output tokens — down from roughly $9 per million output on 3.5 Flash — with its knowledge cutoff advanced to March 2026 and its OSWorld-Verified computer-use score rising to 83% from 78.4%.

Gemini 3.5 Flash-Lite targets low-latency, high-throughput work like agentic search and document processing, running around 350 output tokens per second at $0.30 per million input and $2.50 per million output. Google says it beats the older 3.1 Flash-Lite and outperforms Gemini 3 Flash on several benchmarks. Gemini 3.5 Flash Cyber is a narrow security model fine-tuned to find and fix code vulnerabilities; it ships only through Google's CodeMender agent as a limited-access pilot for governments and trusted partners and is not publicly priced.

Gemini 3.6 Flash and 3.5 Flash-Lite are available immediately in the Gemini app and through the Gemini API (Google AI Studio, Android Studio) and Gemini Enterprise, with 3.6 Flash also in Google Antigravity and Flash-Lite rolling into Google Search.

Why it matters for creators

  • Cheap, fast text generation is now a commodity — the AI price war means anyone can draft captions, scripts, and outlines at pennies per thousand words.
  • When drafting is free, the moat moves downstream: to brand voice, finished formats, and getting content actually published across platforms.
  • The token-efficiency win (17% fewer output tokens) lowers the cost of high-volume batch generation — more variations, more A/B copy, for less.
  • These are reasoning and text models, not image or video generators — they do not make the visual content that dominates social feeds.
  • A raw model still leaves creators with the manual work: designing, formatting per platform, and copy-pasting into every app by hand.

How to act on this with Kompozy

Every Flash price cut lands the same way for creators: the cost of raw words drops, and the real work — turning words into on-brand posts that ship across every platform — stays exactly where it was. That gap is where Kompozy lives. Use Gemini 3.6 Flash or Flash-Lite to brainstorm angles and rough out a week of ideas cheaply, then bring them into Kompozy, where the copy engine (Claude and OpenAI, governed by your Persona Brief) rewrites them in your voice and generates the formats a language model can't: Persona Shorts and HeyGen avatar video, Carousel Posts and Persona Tweets rendered pixel-exact through HyperFrames, Photo Posts, Quote Graphics, blogs, and newsletters.

The point isn't to pick a model — it's to stop the win from evaporating at the copy-paste step. Kompozy schedules and fans that finished set to Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads, plus Mailchimp and your blog, from one queue with autopilot and a per-post review pipeline. Cheaper drafting only helps if the drafts become published content — that's the half Kompozy owns.

Quick takeaways

  • Gemini 3.6 Flash: $1.50/$7.50 per million in/out tokens, ~17% fewer output tokens, March 2026 knowledge cutoff.
  • Gemini 3.5 Flash-Lite: $0.30/$2.50 per million, ~350 tokens/sec, built for agentic search and document processing.
  • Gemini 3.5 Flash Cyber: security-only, CodeMender pilot for governments and trusted partners, not publicly priced.
  • Gemini 3.5 Pro still in partner testing; Gemini 4 pre-training has begun.

Frequently asked questions

When did Google release Gemini 3.6 Flash and 3.5 Flash-Lite?

Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, 2026. Gemini 3.6 Flash and 3.5 Flash-Lite are available in the Gemini app and via the Gemini API; Flash Cyber ships only through the CodeMender pilot.

How much cheaper is Gemini 3.6 Flash than 3.5 Flash?

Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, down from roughly $9 per million output on 3.5 Flash, and Google says it uses about 17% fewer output tokens per task — a compounding cost reduction on high-volume work.

Can Gemini 3.6 Flash create social media videos or images?

No. Gemini 3.6 Flash and 3.5 Flash-Lite are text and multimodal reasoning models — they write and reason but do not generate images or video. To turn their drafts into finished visual posts and publish them, creators use a generation-and-publishing engine like Kompozy.

What is Gemini 3.5 Flash Cyber for?

Gemini 3.5 Flash Cyber is a security-focused model fine-tuned to detect and patch code vulnerabilities at a lower price per token than larger models. It is available only through Google's CodeMender agent as a limited-access pilot for governments and trusted partners.

Related news

← All AI news · Get started →