Aurora Mobile's unified AI media API — one standardized call to more than 200 image, video, and audio models, including engines that render a clip and its soundtrack together.
Last verified · 2026-08-12 · by Moe Ameen
Modellix is Aurora Mobile's (NASDAQ: JG) unified AI media generation platform, launched April 9, 2026. Instead of training its own model, Aurora Mobile aggregates leading third-party image, video, and audio engines — more than 200 by its own count — behind one standardized API, a browser Playground, a command-line tool (modellix-cli), and MCP support, with public per-second pricing and unified call logs. The pitch is to remove the busywork of AI media at scale: one request format instead of a different SDK per vendor, one place to compare and swap models, and one bill.
For video and audio specifically, Modellix has become a place where the soundtrack comes with the picture. Several of its video engines now generate synchronized native audio in the same request: Alibaba's Wan series renders a roughly 10-second HD clip with matching voiceover, sound effects, and music in a single pass; Vidu Q3-Mix (added June 15, 2026) does native audio-video up to 16 seconds with multi-character dialogue and multilingual output; and PixVerse V6 (added August 12, 2026) returns up to a 15-second, 1080p clip with synchronized native audio. Alongside those, Modellix carries dedicated audio models — text-to-speech and voice engines such as Qwen Audio, CosyVoice, MiniMax Speech, and Grok Voice — plus speech-to-text and speech-to-speech categories, so a standalone voice track is one call too. It also routes to other cinematic video heavyweights like Veo, Kling, Seedance, and Hailuo, many of which now generate their own synchronized audio as well.
The honest framing: Modellix is developer and enterprise infrastructure, not a creator studio. You reach every model through an API, a CLI, an MCP connection, or a bare Playground, and what comes back is a rendered file. There is no editor, no caption burner, no per-platform reframing, no persona system to hold one on-camera identity across posts, and no scheduler or multi-platform publishing. Because Modellix rotates models, promotions, and per-second prices frequently, treat any specific model, clip length, or price as a snapshot and confirm current details on modellix.ai before you build against it.
Think of Modellix as a model supermarket: you shop the best engine for each shot — Vidu for a dialogue scene, Wan for HD-with-music, PixVerse V6 for a finished-grade cut, a TTS model for narration — and you leave with a cart of raw files. That à-la-carte breadth is real power for an engineer, and it is also the problem for anyone building a channel: those files have different looks, different voices, no captions, no formats, and no home. [Kompozy](/) is the layer that turns a cart of Modellix renders into a recurring, on-brand show — and it needs no code at all. The [Persona Brief](/glossary/persona-brief) enforces one voice across every caption and script, and [HyperFrames](/glossary/hyperframes) applies pixel-exact brand styling to every clip no matter which engine rendered it, so a week assembled from three different models reads as a single identity rather than a stock-footage mixtape.
Then Kompozy does the two things a gateway structurally cannot. First, it finishes and distributes: branded captions burned in for the muted feed, reframes to 9:16, 1:1, and 16:9 per destination, a hook overlay on the silent opening, and scheduling plus publishing across the eight social platforms plus blog and email from one queue with [Autopilot](/glossary/autopilot). Second, it generates the formats no video model makes — [carousels](/glossary/output-buckets), quote cards, face-locked persona photos, blog articles, and email newsletters — and its own [Persona Shorts](/glossary/persona-shorts) and [Persona Frames](/glossary/persona-frames) avatar video, so you are not limited to whatever Modellix rendered. Modellix is the standardized way to summon raw media; Kompozy is the no-code way to make it a brand and ship it everywhere.
Modellix is Aurora Mobile's (NASDAQ: JG) unified AI media generation platform, launched April 9, 2026. It gives developers and enterprises one standardized API — plus a browser Playground, a CLI, and MCP support — to more than 200 image, video, and audio models, with transparent per-second pricing.
Yes, for several video engines. Alibaba's Wan renders a ~10-second HD clip with matching voiceover, effects, and music in one pass; Vidu Q3-Mix does native audio-video up to 16 seconds with multi-character dialogue; and PixVerse V6 returns up to a 15-second 1080p clip with synchronized native audio. It also offers standalone text-to-speech and voice models.
It depends on the job. As developer infrastructure — one API to hundreds of models with per-second pricing — it is strong. But it returns a raw file: no captions, no per-platform formats, no brand layer, and no publishing. Creators who want finished, posted content need a content engine like Kompozy on top of it.
Modellix is an aggregator: one key and one request shape across 200+ models, so you can swap PixVerse V6 for Vidu, Wan, Veo, Kling, or Seedance without rewiring. It is infrastructure (API, CLI, MCP, Playground), not any one vendor's consumer app, and it has no editor or publishing layer of its own.
Generate the clip or voice track on Modellix, download it, and bring it into Kompozy — no code needed. Kompozy adds branded captions, reframes per platform, and unifies brand styling across engines, then fans the idea into a carousel, quote card, blog, and captions in your voice, scheduled and published across the eight social platforms plus blog and email.