OpenArt shipped Arena — public leaderboards that rank AI image and video models by creative job (filmmaking, motion design, video editing, lip sync, graphic design, e-commerce) instead of one all-purpose score, judged by expert-led blind comparisons.
2026-09-15 · by Moe Ameen
OpenArt announced Arena on September 15, 2026 — a set of public leaderboards that rank AI image and video generation models by the specific creative job they're doing, rather than assigning one all-purpose "best model" score. OpenArt is a roughly four-year-old San Francisco company, co-founded by former Googlers Coco Mao and John Qiao, that runs a commercial creation platform aggregating more than 100 image and video models. Arena splits video ranking into boards such as filmmaking, motion design, video editing, and lip sync, and image ranking into boards including graphic design, e-commerce, and film-oriented imagery, each with its own overall ranking.
The methodology is a preference benchmark. Evaluators are shown two model outputs for the same prompt and pick the better one — blind, side-by-side, rather than scoring each in isolation — and OpenArt aggregates those choices with the Bradley-Terry statistical model, displaying rankings with 95% confidence intervals. Judging runs in two layers: a small "Creative Expert Council" of named practitioners (including Emmy-winning director William Lau, creative technologist Willonius Hatcher, and marketing leader David Shing, plus figures from organizations such as Edelman and UCLA), and a broader planned pool OpenArt described as 800 to 1,000 "tastemakers" drawn from its users and outside creative fields. Treat that pool figure as a plan, not a verified count of completed launch evaluations.
At launch, ByteDance's Seedance 2.5 led the overall video board, ahead of Alibaba's Wan 3.0 and ByteDance's Seedance 2.0; the video-editing board was an exception where Wan 3.0 edged Seedance 2.5. On images, ByteDance's Seedream 5.0 Pro topped the overall board ahead of OpenAI's GPT Image 2, which led the graphic-design board. OpenArt did not publish final judge counts, total pairwise judgments, or complete prompt sets — it says it will release partial prompt sets and withhold others to reduce the risk of models being tuned to the benchmark. Worth stating plainly: OpenArt is a commercial creation platform, not an independent benchmark lab, and it plans to wire Arena into its own model selector, so the leaderboard sits inside the market it evaluates. Head of Growth & Operations Stella Guan framed the goal as making Arena "a new standard" the field recognizes.
The practical move the day Arena launches is to stop guessing which generator to open and start picking per task — Seedance 2.5 or Wan 3.0 for a video cut, GPT Image 2 for a graphic-design frame, whatever leads lip sync when you need a talking clip. But a leaderboard optimizes the *input*, and a raw generation is where the work starts, not ends: it arrives with no captions, no brand styling, in one aspect ratio, and with nowhere to go. [Kompozy](/) is the layer that owns the output. Drop a top-ranked model's clip into Quick Ingest and it becomes vertical [Clipped Shorts](/glossary/clipped-short) with burned-in branded captions, reframed for each platform, then scheduled and fanned across the eight social platforms plus blog and email through [Autopilot](/glossary/autopilot), behind a per-post review gate.
Arena's whole premise — different models win different jobs — is also the case for a model-agnostic engine sitting downstream. Whatever wins next month, Kompozy takes its output and, under one [Persona Brief](/glossary/persona-brief) that holds your voice and banned phrases, spins the same idea into brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), a Blog Article, native Text Posts, and an Email Newsletter. And where no leaderboard model gives you a consistent on-camera presenter, Kompozy generates its own face-locked [Persona Shorts](/glossary/persona-shorts). Use Arena to pick the best raw asset; use Kompozy to make it a published, on-brand week.
OpenArt Arena is a set of public leaderboards, launched September 15, 2026, that rank AI image and video generation models by specific creative task — filmmaking, motion design, video editing, lip sync, graphic design, e-commerce, and more — instead of one all-purpose score. Rankings come from blind, side-by-side comparisons of model outputs aggregated with the Bradley-Terry statistical model.
Evaluators are shown two model outputs for the same prompt and pick the better one blind, rather than scoring each in isolation. OpenArt aggregates those preferences with the Bradley-Terry model and shows rankings with 95% confidence intervals. Judging runs in two layers: a small Creative Expert Council of named practitioners and a broader planned pool of 800–1,000 "tastemakers." OpenArt hasn't disclosed final judge counts or complete prompt sets.
At launch, ByteDance's Seedance 2.5 led the overall video board (ahead of Alibaba's Wan 3.0 and Seedance 2.0), while ByteDance's Seedream 5.0 Pro led overall images ahead of OpenAI's GPT Image 2, which topped the graphic-design board. Wan 3.0 edged Seedance 2.5 on the video-editing board. Results shift as models update, so check the live boards.
Not exactly. OpenArt is a commercial creation platform that aggregates 100+ models and plans to wire Arena into its own model selector, so the leaderboard sits inside the market it evaluates. That doesn't make the results wrong, but read them as directional — a vendor ranking models it also serves, with judge counts and full prompt sets undisclosed at launch.