// AI NEWS · FEATURE

OpenArt Launches Arena, Task-Specific Leaderboards Ranking AI Image and Video Models

OpenArt shipped Arena — public leaderboards that rank AI image and video models by creative job (filmmaking, motion design, video editing, lip sync, graphic design, e-commerce) instead of one all-purpose score, judged by expert-led blind comparisons.

2026-09-15 · by Moe Ameen

What happened

OpenArt announced Arena on September 15, 2026 — a set of public leaderboards that rank AI image and video generation models by the specific creative job they're doing, rather than assigning one all-purpose "best model" score. OpenArt is a roughly four-year-old San Francisco company, co-founded by former Googlers Coco Mao and John Qiao, that runs a commercial creation platform aggregating more than 100 image and video models. Arena splits video ranking into boards such as filmmaking, motion design, video editing, and lip sync, and image ranking into boards including graphic design, e-commerce, and film-oriented imagery, each with its own overall ranking.

The methodology is a preference benchmark. Evaluators are shown two model outputs for the same prompt and pick the better one — blind, side-by-side, rather than scoring each in isolation — and OpenArt aggregates those choices with the Bradley-Terry statistical model, displaying rankings with 95% confidence intervals. Judging runs in two layers: a small "Creative Expert Council" of named practitioners (including Emmy-winning director William Lau, creative technologist Willonius Hatcher, and marketing leader David Shing, plus figures from organizations such as Edelman and UCLA), and a broader planned pool OpenArt described as 800 to 1,000 "tastemakers" drawn from its users and outside creative fields. Treat that pool figure as a plan, not a verified count of completed launch evaluations.

At launch, ByteDance's Seedance 2.5 led the overall video board, ahead of Alibaba's Wan 3.0 and ByteDance's Seedance 2.0; the video-editing board was an exception where Wan 3.0 edged Seedance 2.5. On images, ByteDance's Seedream 5.0 Pro topped the overall board ahead of OpenAI's GPT Image 2, which led the graphic-design board. OpenArt did not publish final judge counts, total pairwise judgments, or complete prompt sets — it says it will release partial prompt sets and withhold others to reduce the risk of models being tuned to the benchmark. Worth stating plainly: OpenArt is a commercial creation platform, not an independent benchmark lab, and it plans to wire Arena into its own model selector, so the leaderboard sits inside the market it evaluates. Head of Growth & Operations Stella Guan framed the goal as making Arena "a new standard" the field recognizes.

Why it matters for creators

  • Task-specific ranking is the useful part: "best video model overall" is a weak signal, but "best model for lip sync" or "best for graphic design" maps directly to a job you actually have.
  • It's a model picker, not a moat. Knowing GPT Image 2 tops graphic design or Seedance 2.5 leads video tells you which generator to open — it doesn't caption, brand, reframe, or publish anything.
  • The leaderboard rewards raw output quality, but a feed rewards consistency and cadence. The winning model still hands you a naked file with no brand voice and nowhere to post it.
  • Read the results with the conflict in mind — a vendor ranking models it also serves, with judge counts and prompt sets undisclosed, is directional, not gospel. Confirm on your own output.
  • The models on top shift monthly. Building a workflow around whatever the current leader is only works if the layer downstream — production and publishing — is model-agnostic.

How to act on this with Kompozy

The practical move the day Arena launches is to stop guessing which generator to open and start picking per task — Seedance 2.5 or Wan 3.0 for a video cut, GPT Image 2 for a graphic-design frame, whatever leads lip sync when you need a talking clip. But a leaderboard optimizes the *input*, and a raw generation is where the work starts, not ends: it arrives with no captions, no brand styling, in one aspect ratio, and with nowhere to go. [Kompozy](/) is the layer that owns the output. Drop a top-ranked model's clip into Quick Ingest and it becomes vertical [Clipped Shorts](/glossary/clipped-short) with burned-in branded captions, reframed for each platform, then scheduled and fanned across the eight social platforms plus blog and email through [Autopilot](/glossary/autopilot), behind a per-post review gate.

Arena's whole premise — different models win different jobs — is also the case for a model-agnostic engine sitting downstream. Whatever wins next month, Kompozy takes its output and, under one [Persona Brief](/glossary/persona-brief) that holds your voice and banned phrases, spins the same idea into brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), a Blog Article, native Text Posts, and an Email Newsletter. And where no leaderboard model gives you a consistent on-camera presenter, Kompozy generates its own face-locked [Persona Shorts](/glossary/persona-shorts). Use Arena to pick the best raw asset; use Kompozy to make it a published, on-brand week.

Quick takeaways

  • OpenArt launched Arena on September 15, 2026 — public leaderboards ranking AI image and video models by creative task (filmmaking, motion design, video editing, lip sync, graphic design, e-commerce) rather than one overall score.
  • Rankings use blind pairwise comparisons aggregated with the Bradley-Terry model, judged by a Creative Expert Council plus a planned pool of 800–1,000 "tastemakers."
  • At launch, Seedance 2.5 led overall video and Seedream 5.0 Pro led overall images, with GPT Image 2 topping graphic design; OpenArt withheld judge counts and full prompt sets.
  • It ranks models — it doesn't produce content. The move is to pick the top model per task, then generate, caption, brand, and publish with an engine like Kompozy.

Frequently asked questions

What is OpenArt Arena?

OpenArt Arena is a set of public leaderboards, launched September 15, 2026, that rank AI image and video generation models by specific creative task — filmmaking, motion design, video editing, lip sync, graphic design, e-commerce, and more — instead of one all-purpose score. Rankings come from blind, side-by-side comparisons of model outputs aggregated with the Bradley-Terry statistical model.

How does OpenArt Arena rank models?

Evaluators are shown two model outputs for the same prompt and pick the better one blind, rather than scoring each in isolation. OpenArt aggregates those preferences with the Bradley-Terry model and shows rankings with 95% confidence intervals. Judging runs in two layers: a small Creative Expert Council of named practitioners and a broader planned pool of 800–1,000 "tastemakers." OpenArt hasn't disclosed final judge counts or complete prompt sets.

Which models topped OpenArt Arena at launch?

At launch, ByteDance's Seedance 2.5 led the overall video board (ahead of Alibaba's Wan 3.0 and Seedance 2.0), while ByteDance's Seedream 5.0 Pro led overall images ahead of OpenAI's GPT Image 2, which topped the graphic-design board. Wan 3.0 edged Seedance 2.5 on the video-editing board. Results shift as models update, so check the live boards.

Is OpenArt Arena an independent benchmark?

Not exactly. OpenArt is a commercial creation platform that aggregates 100+ models and plans to wire Arena into its own model selector, so the leaderboard sits inside the market it evaluates. That doesn't make the results wrong, but read them as directional — a vendor ranking models it also serves, with judge counts and full prompt sets undisclosed at launch.

Related news

← All AI news · Get started →