Aurora Mobile's unified AI media gateway now routes to several models that return a finished clip with synchronized native audio — dialogue, sound effects, and music — in one request, with no separate voiceover or scoring step.
2026-08-12 · by Moe Ameen
Modellix, Aurora Mobile's (NASDAQ: JG) unified AI media platform, has quietly crossed a line that matters more than any single model launch: on its gateway, generating a video and the audio that goes with it is now one API call. The platform, which launched April 9, 2026, aggregates more than 200 third-party image, video, and audio engines behind one standardized request format, a browser Playground, a command-line tool (modellix-cli), and MCP support, with public per-second pricing. What is new is not the gateway — it is how many of the video models on it now produce a synchronized soundtrack in the same pass.
Three examples make the shift concrete. Alibaba's Wan series generates a roughly 10-second HD clip with a matching voiceover, sound effects, and music produced together in a single pass. Vidu Q3-Mix, which Aurora Mobile added to Modellix on June 15, 2026, does native audio-video generation up to 16 seconds with multi-character dialogue and multilingual output, aimed at advertising, e-commerce, and AI drama. And PixVerse V6, added August 12, 2026, returns up to a 15-second, 1080p clip with synchronized native audio — dialogue, effects, and ambient sound — in one request. Each arrives through the same Modellix request shape, so switching engines is a parameter change rather than a new integration.
Modellix also carries dedicated audio models — text-to-speech and voice engines alongside speech-to-text and speech-to-speech categories — so a creator can generate a standalone voice track as easily as a clip. The common thread is that the busywork of stitching sound onto silent renders is leaving the generation step. Aurora Mobile paired several of these additions with launch promotions; treat model availability, clip lengths, and pricing as a snapshot, because Modellix rotates its roster and terms often — confirm current details on modellix.ai.
The honest framing has not changed: Modellix is developer and enterprise infrastructure, not a creator studio. You reach these models through an API, a CLI, an MCP connection, or a bare Playground, and what comes back is a rendered file with a soundtrack — not a captioned, reframed, on-brand post ready for a feed.
The tempting read is that once a clip ships with its own dialogue and music, it is done. It is not — because the feeds that clip is headed for play muted by default, and because a video model's soundtrack does nothing to make a Wan render, a Vidu scene, and a PixVerse cut look like they came from the same channel. That is the specific gap [Kompozy](/) closes. Kompozy is a full AI generation and multi-platform publishing engine, not another model gateway. Bring any Modellix export in and it burns in branded, on-style captions so the audio-optional feed still lands, reframes to 9:16, 1:1, and 16:9 per destination, and stacks a hook overlay through [HyperFrames](/glossary/hyperframes) — then schedules and publishes across the eight social platforms plus blog and email from one queue behind a per-post review gate with [Autopilot](/glossary/autopilot). The [Persona Brief](/glossary/persona-brief) and HyperFrames apply the same voice and pixel-exact styling to every clip regardless of which engine rendered it, so a mixed-model week reads as one identity.
It also makes the formats no video model can. Feed the concept behind a Modellix clip into Kompozy and one render fans into a [carousel](/glossary/output-buckets) breaking down each beat, a quote card pulled from the script, a face-locked persona photo, a captioned [Persona Short](/glossary/persona-shorts), a blog draft, and platform-native captions in your voice. Modellix ends when the file — sound and all — is ready; Kompozy turns that file into a channel.
Yes, for several of its video models. Alibaba's Wan produces a ~10-second HD clip with matching voiceover, sound effects, and music in one pass; Vidu Q3-Mix does native audio-video up to 16 seconds with multi-character dialogue; and PixVerse V6 returns up to a 15-second 1080p clip with synchronized native audio. Modellix also offers standalone text-to-speech and voice models.
Alibaba's Wan series, Vidu Q3-Mix (added June 15, 2026), and PixVerse V6 (added August 12, 2026) each generate a clip and its soundtrack in the same request. Clip lengths run roughly 10–16 seconds depending on the model. Modellix rotates its roster, so confirm current models and terms on modellix.ai.
Not quite. It arrives as a rendered file with a soundtrack, but no captions for muted feeds, no per-platform aspect ratios, no brand styling, and no way to publish. Most short-form is watched on mute, so you still need captions, reframing, and a scheduler — which is a content engine's job, not a model gateway's.
Generate the clip on Modellix, then bring it into Kompozy to burn in branded captions, reframe per platform, and stack a hook overlay — and fan the same idea into a carousel, quote card, blog, and captions in your voice, scheduled and published across the eight social platforms plus blog and email.