// AI TOOLS · MODELLIX

Modellix

Aurora Mobile's unified AI media API — one standardized call to more than 200 image, video, and audio models, including engines that render a clip and its soundtrack together.

Last verified · 2026-08-12 · by Moe Ameen

What Modellix is

Modellix is Aurora Mobile's (NASDAQ: JG) unified AI media generation platform, launched April 9, 2026. Instead of training its own model, Aurora Mobile aggregates leading third-party image, video, and audio engines — more than 200 by its own count — behind one standardized API, a browser Playground, a command-line tool (modellix-cli), and MCP support, with public per-second pricing and unified call logs. The pitch is to remove the busywork of AI media at scale: one request format instead of a different SDK per vendor, one place to compare and swap models, and one bill.

For video and audio specifically, Modellix has become a place where the soundtrack comes with the picture. Several of its video engines now generate synchronized native audio in the same request: Alibaba's Wan series renders a roughly 10-second HD clip with matching voiceover, sound effects, and music in a single pass; Vidu Q3-Mix (added June 15, 2026) does native audio-video up to 16 seconds with multi-character dialogue and multilingual output; and PixVerse V6 (added August 12, 2026) returns up to a 15-second, 1080p clip with synchronized native audio. Alongside those, Modellix carries dedicated audio models — text-to-speech and voice engines such as Qwen Audio, CosyVoice, MiniMax Speech, and Grok Voice — plus speech-to-text and speech-to-speech categories, so a standalone voice track is one call too. It also routes to other cinematic video heavyweights like Veo, Kling, Seedance, and Hailuo, many of which now generate their own synchronized audio as well.

The honest framing: Modellix is developer and enterprise infrastructure, not a creator studio. You reach every model through an API, a CLI, an MCP connection, or a bare Playground, and what comes back is a rendered file. There is no editor, no caption burner, no per-platform reframing, no persona system to hold one on-camera identity across posts, and no scheduler or multi-platform publishing. Because Modellix rotates models, promotions, and per-second prices frequently, treat any specific model, clip length, or price as a snapshot and confirm current details on modellix.ai before you build against it.

What you can make with it

  • A finished clip with synchronized native audio in one request — dialogue, sound effects, and music generated in the same pass via Wan, Vidu Q3-Mix, or PixVerse V6
  • Narrative and dialogue video up to ~16 seconds with multi-character speech and multilingual output through Vidu Q3-Mix
  • Cinematic video from other engines like Veo, Kling, Seedance, or Hailuo through the same gateway
  • Standalone voiceover and narration from dedicated text-to-speech and voice models on the same API
  • Image generation from the same gateway, for thumbnails, stills, and reference frames
  • A model-agnostic pipeline that swaps one video or audio engine for another behind a single standardized request shape

How Kompozy turns Modellix output into content

Think of Modellix as a model supermarket: you shop the best engine for each shot — Vidu for a dialogue scene, Wan for HD-with-music, PixVerse V6 for a finished-grade cut, a TTS model for narration — and you leave with a cart of raw files. That à-la-carte breadth is real power for an engineer, and it is also the problem for anyone building a channel: those files have different looks, different voices, no captions, no formats, and no home. [Kompozy](/) is the layer that turns a cart of Modellix renders into a recurring, on-brand show — and it needs no code at all. The [Persona Brief](/glossary/persona-brief) enforces one voice across every caption and script, and [HyperFrames](/glossary/hyperframes) applies pixel-exact brand styling to every clip no matter which engine rendered it, so a week assembled from three different models reads as a single identity rather than a stock-footage mixtape.

Then Kompozy does the two things a gateway structurally cannot. First, it finishes and distributes: branded captions burned in for the muted feed, reframes to 9:16, 1:1, and 16:9 per destination, a hook overlay on the silent opening, and scheduling plus publishing across the eight social platforms plus blog and email from one queue with [Autopilot](/glossary/autopilot). Second, it generates the formats no video model makes — [carousels](/glossary/output-buckets), quote cards, face-locked persona photos, blog articles, and email newsletters — and its own [Persona Shorts](/glossary/persona-shorts) and [Persona Frames](/glossary/persona-frames) avatar video, so you are not limited to whatever Modellix rendered. Modellix is the standardized way to summon raw media; Kompozy is the no-code way to make it a brand and ship it everywhere.

  1. On Modellix, generate what you need — a native-audio clip from Wan, Vidu Q3-Mix, or PixVerse V6, a cut from another engine like Veo or Kling, or a voice track from a TTS model.
  2. Download the file and bring it into Kompozy — no code required.
  3. Let Kompozy burn in branded captions, reframe per platform, and stack a hook overlay through HyperFrames so every clip matches your brand regardless of engine.
  4. Fan the same idea into a carousel, quote card, persona photo, blog, and newsletter, all governed by your Persona Brief.
  5. Schedule and publish the set across the eight social platforms plus blog and email from one queue with Autopilot.

Frequently asked questions

What is Modellix?

Modellix is Aurora Mobile's (NASDAQ: JG) unified AI media generation platform, launched April 9, 2026. It gives developers and enterprises one standardized API — plus a browser Playground, a CLI, and MCP support — to more than 200 image, video, and audio models, with transparent per-second pricing.

Can Modellix generate video and audio in the same request?

Yes, for several video engines. Alibaba's Wan renders a ~10-second HD clip with matching voiceover, effects, and music in one pass; Vidu Q3-Mix does native audio-video up to 16 seconds with multi-character dialogue; and PixVerse V6 returns up to a 15-second 1080p clip with synchronized native audio. It also offers standalone text-to-speech and voice models.

Is Modellix a good tool for creators?

It depends on the job. As developer infrastructure — one API to hundreds of models with per-second pricing — it is strong. But it returns a raw file: no captions, no per-platform formats, no brand layer, and no publishing. Creators who want finished, posted content need a content engine like Kompozy on top of it.

How is Modellix different from using a model like PixVerse or Vidu directly?

Modellix is an aggregator: one key and one request shape across 200+ models, so you can swap PixVerse V6 for Vidu, Wan, Veo, Kling, or Seedance without rewiring. It is infrastructure (API, CLI, MCP, Playground), not any one vendor's consumer app, and it has no editor or publishing layer of its own.

How do I turn Modellix video and audio into finished, published posts?

Generate the clip or voice track on Modellix, download it, and bring it into Kompozy — no code needed. Kompozy adds branded captions, reframes per platform, and unifies brand styling across engines, then fans the idea into a carousel, quote card, blog, and captions in your voice, scheduled and published across the eight social platforms plus blog and email.

Related tools

  • PixVerse V6 (Modellix)PixVerse's flagship video model served through Modellix, Aurora Mobile's unified API — finished-grade 1080p clips with native audio in a single request.
  • PixVerseA consumer AI video generator — text-to-video and image-to-video with native audio, multi-character lip sync, and a library of viral one-tap effects.
  • Google Veo 3Google DeepMind's video model that was the first to generate synchronized native audio — dialogue, sound effects, and music — inside the same pass as the video, with lip sync.
  • Kling AIKuaishou's text-to-video and image-to-video model — turn a prompt or a still into a cinematic clip with camera motion, lip sync, and native audio.
  • ByteDance Seedance 2.5AI video model that generates a 30-second clip in one pass — no stitching.
  • Hailuo AI Video GeneratorMiniMax's AI video generator, known for physically believable motion and strong instruction following from text or a single image.
  • SunoThe consumer AI music generator that writes a full song — lyrics, vocals, instrumentation, and mix — from a text prompt, now under a copyright cloud over how it was trained.
  • HeyGenAI avatar video platform that turns a text script into a talking-head video — in 175+ languages.

← All AI tools · Get started →