This is RCTV’s living reference to the AI video software stack — models, orchestration agents, and open-source generation. Updated as products launch, pricing changes, and capabilities evolve. Last updated: August 9, 2026.
The short answer (July 2026). There is no single best AI video model — the right pick is set by the axis you optimize. For photorealism, Veo 3.1 (4K native, free on any Google account). For broadcast-grade 4K motion, Kling 3.0 (native 4K at 60fps; Kuaishou closed a record ~$3B round at an ~$18B valuation in July 2026, the clearest capital signal in AI video to date). For character consistency across shots, Seedance 2.0 Pro (ByteDance). For stylized and VFX work, Runway; for editing existing footage, Runway’s Aleph 2.0. For free consumer distribution, Gemini Omni Flash (Google; multimodal input, SynthID + C2PA provenance by default). For open-weights local generation, LTX-2.5 (true 4K on consumer GPUs; LTX-2.x Community License, free below $10M revenue, not OSI-open). For top benchmark quality, Gemini Omni Flash swept all four Artificial Analysis boards in July 2026 (#1 on audio and silent, text- and image-to-video) — though on an early vote count and behind a 10-second preview cap; HappyHorse-1.0/1.1 (Alibaba) lead the open-API tier just behind it. The cheapest audio-capable API is Grok Imagine 1.5 at $0.08/sec. The larger 2026 shift: the models are converging, and the contest has moved off raw model quality to the orchestration and assembly layer built on top of them (see The Agentic Layer).
Quick Verdict: Best AI Video Model by Use Case (2026)
Routing decisions based on what each model actually wins at — not benchmark Elo alone, but the practical question of which model handles a specific use case best. Verdicts current as of September 1, 2026. Scroll down for full model breakdowns and the detailed routing framework.
- Best for photorealism — Veo 3.1. 4K native, the cleanest photorealistic rendering available; free 10 clips/month via Google Vids on any Google account.
- Best for broadcast 4K and motion quality — Kling 3.0. Native 4K at 60fps with multi-cut storyboards; the only production-grade model meeting broadcast delivery standards without upscaling.
- Best for character consistency across shots — Seedance 2.0 Pro. Native multi-shot storytelling with frame-level character and scene control; available in the US via CapCut with real-face restrictions.
- Best for stylized and VFX work — Runway Gen-4 Turbo. The most mature professional ecosystem for non-photorealistic aesthetics; motion brushes, scene consistency tools, $315M Series C runway.
- Best for editing existing video — Aleph 2.0 in Runway Edit Studio. Multishot edit propagation up to 30 seconds at 1080p; edit one frame, model carries it across the sequence; available on all paid Runway plans.
- Best for free consumer distribution — Gemini Omni Flash. Free inside YouTube Shorts and YouTube Create; multimodal any-to-any input (image / audio / video / text → video); conversational multi-turn editing; SynthID + C2PA provenance by default.
- Best open-weights for local generation — LTX-2.5. True 4K native on consumer GPUs (12GB+ VRAM); standalone desktop editor plus ComfyUI integration. Open weights under the LTX-2.x Community License, not an OSI-open licence — free below $10M annual revenue, and it bars training a competing model on the outputs. On measured quality it trails: ranks 21–23 of 33 on Artificial Analysis’s audio text-to-video board.
- Best benchmark quality — Wan 3.0. Leads two of the three tracked Artificial Analysis boards as of August 28, 2026 — text-to-video (Elo 1,241) and video-editing (Elo 1,189). Gemini Omni Flash held all four boards through mid-July and now sits #2 T2V (1,237), #4 I2V (1,179), #3 video-editing (1,124); its July sweep is over. On image-to-video the leader is MiniMax H3 Max, a fal post-train of an open-weight model (Elo 1,202) — the first time a third-party fine-tune tops a board this page tracks.
- Cheapest commercial API — Grok Imagine Video 1.5. $0.08/sec ($4.80/min), audio-native, GA as of June 16, 2026 — undercuts every audio-capable tier except the original Grok Imagine ($0.05/sec, no audio). Native 1080p shipped July 31, 2026, closing the resolution gap tracked since Musk’s missed April commitment. Arena positions in the entry.
- Best multi-model orchestration agent — Pika Agents. Broadest model roster in the category (Kling, Veo, Seedance, MiniMax, Sora API, Pika Video); runs inside Slack, Telegram, Discord, Notion, Figma, and 12+ other surfaces; persistent memory across sessions. See The Agentic Layer.
- Best for generative video editing — Wan 3.0. #1 on Artificial Analysis’s video-editing leaderboard (Elo 1,189, August 28, 2026), taking the category from MiniMax H3 (now #2, Elo 1,129), which had held it since its August 3 debut. Wan 3.0 is invite-gated commercial beta; H3 remains the one you can actually buy today at a published price.
- Fastest generation — MiniMax H3 Max (post-trained by fal). A 5-second 768p clip in under 3 seconds — generation faster than playback, the threshold Weekly Roundup — August 31, 2026 covered. Free tier, no signup for the first tier. The sustained-load behaviour is unproven: fal’s own livestream product built on it was paused three hours after launch.
Quick Reference: All Models at a Glance
| Model | Best For | Max Resolution | Free Tier | Paid From | API |
|---|---|---|---|---|---|
| Veo 3.1 (Google DeepMind) | Photorealism, widest free access | 4K | ✓ 10 clips/mo via Google Vids | $19.99/mo (AI Pro) | ✓ (also via Adobe Firefly) |
| Gemini Omni Flash / Omni 1.1 Flash (Google DeepMind) | Multimodal input + conversational editing; scene extension, first/last-frame control, 4K upscale (1.1, Aug 27); #2 T2V / #4 I2V / #3 video-editing (Aug 28) + Design Arena #1 | 720p native generation; 4K via 1.1 upscale | ✓ YouTube Shorts + YouTube Create | $19.99/mo (Google AI Plus) | ✓ Public preview since June 30, 2026 — $0.10/sec via AI Studio + Gemini API + Enterprise Agent Platform |
| Kling 3.0 / 3.0 Omni (Kuaishou) | Broadcast-ready 4K, 60fps, multi-shot storyboards | 4K native | ✓ | ~$8/mo | ✓ (also via Adobe Firefly) |
| Wan 3.0 (Alibaba) | #1 AA text-to-video (Elo 1,242) and #1 video-editing (Elo 1,190); 30-second single-pass; document-to-video | 1080p (no 4K) | ✗ Invite-gated preview | $0.20/sec 1080p list · 30% off live | ✓ Alibaba Cloud Model Studio (6 regions), Qwen Cloud |
| Seedance 2.0 Pro (ByteDance) | Character consistency, multi-shot | 2K | Via CapCut (US, with restrictions) | Via CapCut / third-party | Via third-party |
| Luma Ray 3.2 / Ray 3.14 (Luma AI) | Production volume, frame-level control (Ray 3.2); duration/loop workflows (Ray 3.14) | 1080p native; 20s max (Ray 3.2) | ✓ | Available | ✓ (Ray 3.2 API launched June 9, 2026) |
| Runway Gen-4 Turbo / Aleph 2.0 | Stylized/VFX, real-time avatars, multishot edit propagation | 1080p | ✗ | $12/mo | ✓ (Gen-4.5 also via Adobe Firefly) |
| Pika 2.5 / Pika Agents (Pika Labs) | Budget creators; multi-model agent orchestration (Kling, Veo, Seedance, MiniMax, Sora) | 1080p | ✗ | $8/mo | ✓ |
| Grok Imagine / Grok Imagine Video 1.5 (SpaceXAI, formerly xAI) | Speed, cheapest audio API; native I2V (#2 arena) | 720p (original); 1.5 native 1080p (Jul 31, 2026) | ✗ | X Premium / SuperGrok | ✓ Original: $0.05/sec ($3.00/min, no audio); 1.5 GA: $0.08/sec ($4.80/min, audio-native) |
| LTX-2.5 (LTX) | Local / private generation | 4K | ✓ Open weights, not OSI-open | Free under $10M revenue; paid Commercial Use Agreement above | ComfyUI |
| HappyHorse-1.0 / 1.1 (Alibaba) | 1.0: #2 T2V no-audio (Elo 1,288, behind Omni Flash’s July sweep); 1.1: #3; API live, open weights pending | 1080p (joint audio, 7-language lip-sync) | 1.0 weights still pending; API live | $0.14/sec 720p · $0.28/sec 1080p | ✓ via fal.ai + Alibaba Cloud Bailian |
| MiniMax H3 (MiniMax) | Omni-modal generation + editing; #2 video-editing (Elo 1,129), #4 T2V (1,227), #3 I2V (1,184) as of Aug 28 | 2K (2560×1440) | ✗ | Hailuo AI app | ✓ $0.13/sec at 2K (model ID MiniMax-H3) |
| MiniMax H3 Max (post-trained by fal) | #1 AA image-to-video (Elo 1,202), #3 T2V (1,235); 5-sec clip in under 3 sec | 768P (no 1080p tier) | ✓ 5/day no account; 15/day at 10s signed in | $0.08/sec 768P list · $0.02/sec until Sep 7 | ✓ via fal |
| Wan 2.7 (Alibaba) | Thinking Mode, 5 unified task types | 1080p | ✓ Open source | Free | ComfyUI / Model Studio |
| MAGI-2 Preview (Sand.ai) | Open-weights 114B MoE, no territory-restricted license; #6 I2V (Elo 1,104) | 1080p (10s clips, native audio) | ✓ Open weights (Apache 2.0) | Free (self-host, 8x Hopper GPUs) | ✗ No hosted API |
| SkyReels V4 (Skywork AI) | Joint audio-video, open-source | 1080p | ✓ 70 credits/mo | Free (open source) | — |
Rankings and pricing change weekly; arena positions are Artificial Analysis unless noted, captured September 1, 2026. Scroll down for full model breakdowns.
How to read the leaderboard figures {#quick-reference}. Artificial Analysis runs three scored leaderboard pages — text-to-video, image-to-video, and video-editing — and each one carries a With Audio / No Audio toggle that swaps the entire ranking and every Elo value on the same URL. The separate text-to-video-with-audio address retired in August 2026. So “three boards” and “audio and no-audio boards” are both half-right: three pages, two datasets each. Figures on this page are the With Audio set unless stated otherwise, because that is the default view and the one vendors quote. The no-audio set ranks differently — Gemini Omni Flash still leads text-to-video there at 1,324, where it is second with audio.
The Big Ten: Commercial Models
These are the production-grade models dominating professional and creator workflows in 2026. The market has matured to the point where no single model leads across all dimensions — the professional standard is now multi-model routing, choosing the right tool for each specific shot.
Luma Ray 3.2 / Ray 3.14 — Luma AI
Best for: Professional production volume, frame-level keyframe control, HDR/EXR output, cost-efficient multi-shot workflows
- Max resolution: 1080p native
- Max duration: Up to 20 seconds at 1080p (Ray 3.2); variable (Ray 3.14)
- Key features (Ray 3.2, June 9, 2026): Up to 16 keyframes per clip for precise motion direction; performance tracking across up to 8 faces simultaneously; native HDR generation with 16-bit EXR export; Enhanced Reframe (aspect ratio, frame extension, background replacement); API launch (first API availability for Ray 3.2)
- Key features (Ray 3.14): Duration-change and loop workflows; Ray3 Modify (hybrid performance/acting control)
- Timeline editor / EDL Export (June 19, 2026): Multi-clip timeline editor with Edit Decision List export for NLE handoff; integrates with Ray 3.2 generation workflow
- Luma Connectors (June 24, 2026): Native integrations with Airtable, Dropbox, and Google Drive for asset management and workflow automation
- Speed: 4× faster generation than previous Ray model (Ray 3.14 baseline); Ray 3.2 maintains the same 1080p baseline
- Pricing: 3× cheaper per-second than previous Ray (Ray 3.14 baseline)
- Access: Luma AI subscription; free tier available
- API: Available for both variants; Ray 3.2 API launched June 9, 2026; enterprise deployments via Luma Agents
- Note: Ray 3.2 and Ray 3.14 are parallel sub-models in the Ray3 family. Ray 3.2 handles standard video-to-video transformation and multi-keyframe guidance; Ray 3.14 handles duration-change and loop workflows. Both remain available.
Luma AI’s Ray 3.14 shipped in January 2026 as the model that stepped into the commercial tier vacated by Sora’s shutdown. (Weekly Roundup — March 27, 2026) Native 1080p output, generation 4× faster than the previous Ray 3 model, per-second pricing 3× cheaper. Ray3 Modify, a companion tool for hybrid performance and acting workflows, gives studios more control over scene continuity and character consistency across shots.
Luma launched Ray 3.2 on June 9, 2026 — framing the update as a shift “from prompting to directing.” The headline addition is keyframe control: up to 16 keyframes per clip at arbitrary positions, letting operators set precise motion direction and pacing rather than prompting for results. Alongside it: performance tracking across up to 8 simultaneous faces, native HDR generation with 16-bit EXR export for post-production pipelines, Enhanced Reframe for aspect-ratio changes and background replacement, and full API availability — the first Ray 3.2 API access. Duration extends to 20 seconds at 1080p. Ray 3.14 stays available and is the better choice specifically for duration-change and loop workflows, where its fixed-length output is an asset rather than a constraint.
Luma is positioning Ray explicitly as professional infrastructure priced for production volume rather than a consumer app — a distinction that looks strategically deliberate given Sora’s failure. The company’s $900M Series C led by HUMAIN, a London office, and enterprise Luma Agents deployments at Publicis, Adidas, and Mazda all reinforce this direction. The Mazda relationship produced a concrete deliverable in April 2026: Boundless, a Johannesburg agency, used Luma Agents to deliver Mazda’s first AI-produced commercial in under two weeks — the most credible production-deployment signal for any AI video platform to date.
Kling 3.0 / Kling 3.0 Omni / Kling 3.0 Turbo — Kuaishou
Best for: Feature density, broadcast-ready output, motion quality; Turbo for rapid iteration and cost-efficiency
- Max resolution: 4K native (60fps) — Standard and O3; 480p–720p — Turbo
- Frame rate: Up to 60fps
- Audio: Native built-in audio in six languages
- Key feature: Multi-cut storyboard generation (up to 6 camera cuts, 15s); Omni/O3 adds shot/camera/character controls; Turbo adds faster generation, lower cost, improved lip-sync, stable motion
- Kling 3.0 Turbo (June 17, 2026): Speed-and-cost tier; faster generation, lower price, improved lip-sync, stable motion; targets rapid iteration workflows
- Kling 3.0 O3 (June 17, 2026): Up to 15-second clips at full 4K; stronger prompt-and-reference consistency; targets production-quality delivery
- Kling MCP & CLI (July 8, 2026): Kling’s model is now agent-callable — a Model Context Protocol endpoint plus a command-line interface, joining the MCP-exposed tier alongside Runway, Pika, and Higgsfield (Weekly Roundup — July 13, 2026)
- Platform rollout (June 17–18, 2026): fal, SeaArt, Clipfly, Fotor, GlamAI, Runware, Morphic — seven partner platforms in a coordinated same-day launch
- Pricing (Turbo/O3): Kuaishou has not published official pricing for either variant; capability claims sourced from partner-platform announcements
- Access: Free tier available; paid plans from ~$8/mo; also via Adobe Firefly (Creative Cloud subscription)
- API: Available via Kuaishou and third-party platforms
The most capability-dense model available. Kling 3.0 is the first AI video model to meet broadcast delivery standards without upscaling, offering native 4K at 60fps. The storyboard feature generates up to six camera cuts in a single generation with visual consistency — a production-first capability no other model matches. The Kling 3.0 Omni (O3) variant adds finer-grained controls for shot duration, camera angle, and character movement across multi-shot sequences, with clips up to 15 seconds at full 4K.
On June 17–18, 2026, Kuaishou rolled two new Kling variants to seven partner platforms simultaneously — fal, SeaArt, Clipfly, Fotor, GlamAI, Runware, and Morphic — with no direct Kuaishou announcement on klingai.com (which remains 446-blocked). Kling 3.0 Turbo targets speed and cost: faster generation, lower price, improved lip-sync, and more stable motion than Standard, designed for rapid iteration and volume workflows. Kling 3.0 O3 targets production quality: clips up to 15 seconds at full 4K, with stronger prompt and reference consistency. One tier for iteration; one for delivery. Pricing for both variants has not been published by Kuaishou; all capability claims are sourced from partner-platform announcements, which agree across all seven independently.
In April 2026, both Kling 3.0 and Kling 3.0 Omni joined Adobe Firefly’s multi-model video hub alongside Veo 3.1, Runway Gen-4.5, and 30+ other AI models — significantly broadening Kling’s distribution to Adobe Creative Cloud’s existing professional user base.
In July 2026 the financing closed. Kuaishou disclosed on July 2 a round of nearly $3 billion at an ~$18 billion post-money valuation — reported as the largest ever raised by an AI-video-model company. It was led by CPE Yuanfeng, Guofang Venture Capital, Tencent, Zhongguancun Science City Fund, and CITIC Securities, with Alibaba Cloud and Baidu among 38-plus investors; Kuaishou’s stake dilutes from 100% to 68.33% (outside investors plus an employee ownership plan), and the round is widely read as pre-positioning for a Hong Kong IPO within 12 months. The revenue underneath it is real: Q1 2026 revenue topped 650M yuan (~$96M), +300% YoY, an annualized run rate near $500M — quadruple a year earlier. It priced below the $20B target floated in May, but stands as the clearest commercial-scale signal in AI video to date, and the first pure-play comparable since Sora’s shutdown removed one (Weekly Roundup — July 13, 2026). The scale behind it is now disclosed: Kling has crossed 100 million cumulative users across 224 countries, ARR **$500M as of March 2026** (up from ~$240M in December 2025; ~70–75% overseas). At Cannes Lions 2026 — whose inaugural “AI Craft” Grand Prix went to Google, not Kling — two Kling-made ad films still won real Lions (Silver + Bronze for The RealReal’s “L’Ultimo Uomo Reale”) (Weekly Roundup — July 20, 2026).
Veo 3.1 — Google DeepMind
Best for: Photorealism, 4K native output, integrated workflows, broadest free access
- Max resolution: 4K native (Flow/Vertex AI); 1080p via Veo 3.1 Lite; 720p via Google Vids free tier
- Audio: Native synchronized audio
- Key features: Flow unified workspace; Google Vids integration (avatars, Lyria 3 music, YouTube export); Veo 3.1 Lite developer tier; voice-driven generation on Gemini-enabled Google TV; Performance Control (July 2026): extracts precise 3D geometry and motion from a real performer to drive generation, with independently editable player/environment layers; debuted in Google DeepMind’s Pelé “Gol da Rua Javari” reconstruction
- Access: Free — 10 clips/month via Google Vids (any Google account); Google AI Pro ($19.99/mo) and Ultra for higher limits; Flow is free; also via Adobe Firefly multi-model hub (April 2026) and Gemini-enabled TCL Google TVs in the US (April 2026)
- API: Vertex AI ($12/min); Veo 3.1 Lite via Gemini API ($0.05/sec 720p, $0.08/sec 1080p); Veo 3.1 Fast pricing reduced April 7, 2026 (check Gemini API docs for current per-second rates)
- Milestone: 1.5 billion images and videos created by Flow users
Google’s model pushes photorealistic rendering to the point where trained observers struggle to identify generated footage in blind tests. It is the engine behind Google Flow (merged creative workspace with Whisk, ImageFX, and multi-clip sequencing) and Google Vids. Veo 3.1 has been freely available to any Google account holder since April 2026 (Weekly Roundup — April 4, 2026) — 10 generations per month, 8 seconds at 720p, from text prompts or uploaded images. Google AI Pro and Ultra subscribers get more: up to 1,000 Veo clips per month, Lyria 3 custom music generation (tracks up to 3 minutes), customizable AI avatars with scene placement and wardrobe control, and direct YouTube export. This is the first time a production-grade AI video model has been made freely accessible to Google’s full account base.
Veo arrived on Gemini-enabled TCL Google TVs in the US in April 2026 (Weekly Roundup — May 11, 2026) — voice-driven generation through Gemini’s Create tab, either from scratch or by animating still images. TCL-only and US-only at launch, with no public timeline for other manufacturers or markets.
On the developer side, Google launched Veo 3.1 Lite in March 2026 via the Gemini API (Weekly Roundup — April 4, 2026) and Google AI Studio — priced at $0.05/sec for 720p and $0.08/sec for 1080p, less than half the cost of the existing Veo 3.1 Fast tier. Veo 3.1 Fast received a further price reduction on April 7, 2026 — compressing the full developer stack from the free consumer tier through production-grade API calls. Check the Gemini API pricing documentation for current per-second rates.
In July 2026, Google DeepMind showcased a new Veo technique, Performance Control, in its reconstruction of Pelé’s never-filmed 1959 goal — extracting 3D geometry and motion from a real stunt player to drive generation, with editable layers that swap the performer and the environment independently. The project pairs the technique with a live-action shoot and traditional VFX, and Google framed the output explicitly as a reconstruction, not archival footage — drawing its own AI-disclosure line ahead of the August 2 EU/California transparency deadlines (Weekly Roundup — July 27, 2026).
Gemini Omni Flash — Google DeepMind
Best for: Multimodal input across image / audio / video / text → video output; conversational multi-turn editing; broadest free consumer distribution via YouTube Shorts; SynthID + C2PA provenance by default
- Max resolution: 720p is the only resolution Omni natively renders. Omni 1.1’s
resolutionparameter offers 360p/720p/1080p/4K, but Google’s own resolution table labels the top two “1080p output (upscaled)” and “4K output (upscaled)” in those words. A 4K Omni clip is a 720p render enlarged, not a 4K generation — the distinction that matters against Veo 3.1 and Kling 3.0, which render 4K natively - Max duration: 10 seconds per generation (cap); “longer durations coming soon” per Google
- Audio: Native multimodal output — synchronized audio generated alongside video
- Watermark: SynthID watermarking plus C2PA Content Credentials embedded by default; verifiable through the Gemini app, Chrome, and Search
- Key feature: Any-to-any multimodal input (text / image / audio / video → video); conversational multi-turn editing via a stateful Interactions API that carries context across turns; reference-driven object and brand insertion
- Omni 1.1 Flash — GA August 27, 2026 as model ID
gemini-omni-1.1-flash;gemini-omni-flash-previewis deprecated on September 30, 2026. Adds scene extension, first- and last-frame specification, up to three video input references, 1080p/4K upscale, and a 360p fast mode for iteration. Scene extension works in 10-second increments to a cumulative 40-second ceiling, each extension using the last 10 seconds as context — so a 40-second Omni clip is four chained requests, not a single pass, and individual generations stay at 3–10 seconds. Read the distribution precisely: Google frames the full set as “for developers” and “available via APIs”; the Gemini app account scopes the consumer rollout to scene extension only — “available to all Google AI Plus, Pro and Ultra subscribers globally.” Five capabilities to developers; one of them to subscribers - Access: Public preview since June 30, 2026 via Google AI Studio, the Gemini API, and Gemini Enterprise Agent Platform; also live in the Gemini app and Google Flow; free via YouTube Shorts and YouTube Create
- API: $17.50 per 1M video output tokens at 5,792 tokens per second of 720p — about $0.10/sec, $6.00/min; input $1.50/1M. Same structure on the GA and preview rows
- Arena: Debuted #1 on Design Arena’s Video Arena at Elo 1,404 (July 2, 2026) — a 101-point margin over second place. Swept all four Artificial Analysis boards by mid-July 2026. That sweep is over: as of August 28, 2026 Omni Flash is #2 text-to-video (Elo 1,237, behind Wan 3.0), #4 image-to-video (1,179), #3 video-editing (1,124). It has been displaced rather than degraded — the Elo is roughly where it was; two newer models scored higher. Artificial Analysis has not scored Omni 1.1 separately: the board entry carries none of the version markers AA uses elsewhere, so these figures describe Omni Flash, not the 1.1 update
- Next: Omni Pro shipping window (still no date); Artificial Analysis benchmark placement; durations past the 10-second preview cap
Google’s any-to-any multimodal video model entered public preview June 30, 2026, via Google AI Studio and the Gemini API — not the general availability that circulated in some coverage that week. Google’s own announcement is explicit: “Gemini Omni is available in public preview starting today.” The spec that shipped: $0.10 per second, 720p native with no 1080p or 4K tier — independently corroborated by VentureBeat’s pricing breakdown — clips capped at 10 seconds (“currently, with longer durations coming soon”), prompted in natural language and edited through a stateful Interactions API that carries context across turns: relight a shot, reframe it, change the wardrobe, without regenerating from scratch. Every output carries SynthID watermarking and C2PA Content Credentials by default.
This answers the question Google’s own I/O 2026 keynote left open: what does “a first step toward a world model” actually ship as? A price, a hard limit, and an admission — in Google’s own words — that “character consistency when changing scenes or panning movements has some limitations.”
Two days after the preview shipped, the first independent scoreboard weighed in: Omni Flash debuted #1 on Design Arena’s Video Arena at Elo 1,404 (July 2, 2026) — a 101-point margin over second place, one of the board’s largest gaps on record. The caveat mattered at preview: Design Arena isn’t Artificial Analysis, the leaderboard RCTV tracks for the rest of the Big Ten. It resolved fast — by mid-July 2026, Omni Flash had swept all four Artificial Analysis boards (with-audio T2V 1,240 / I2V 1,204, no-audio T2V 1,327 / I2V 1,374), passing both Seedance and HappyHorse — weeks after R#12 reported the AA score as pending. On Design Arena’s own image-to-video board, Seedance 2.0 still leads. A clean board sweep on an early vote count and a 10-second preview cap are both true at once (Weekly Roundup — July 6, 2026).
The sweep held about six weeks. By August 28, 2026 Wan 3.0 had taken text-to-video and video-editing and a fal fine-tune had taken image-to-video, leaving Omni Flash at #2/#4/#3 on scores that barely moved — the field caught up rather than the model regressing. Google’s answer, Omni 1.1 Flash on August 27, was not a bigger number: it was scene extension, frame control, video references, and 4K upscale. Controls, not score (Weekly Roundup — August 31, 2026).
Seedance 2.0 Pro — ByteDance
Best for: Character consistency, cinematic motion, multi-shot storytelling
- Max resolution: 4K native, 10-bit color depth (confirmed at Volcano Engine FORCE, June 23, 2026); standard platform tiers vary
- Audio: Native audio with lip-sync
- Key feature: Multi-shot storytelling, quad-modal input, frame-level precision
- Access: China via Jimeng/Dreamina; Africa, South America, Middle East, SE Asia, and US via CapCut/Dreamina Seedance 2.0; global via BigMotion ($35–$95/mo), LumeFlow AI, other third-party platforms; unlimited access via Higgsfield (exclusive BytePlus-powered unlimited tier, June 17, 2026)
- API: Official global API paused; available via third-party integrations (fal.ai, Higgsfield/BytePlus, others)
- US restrictions: Real-face image-to-video disabled; unauthorized IP generation blocked; invisible watermarks on all output
- Seedance 2.5 rollout completed (August 21, 2026): Dreamina brought Seedance 2.5 to US users, closing the regional rollout that began August 14 across Southeast Asia, the Middle East, Africa, Europe and South America. Five assembly-layer platforms — Higgsfield (Aug 14), Dreamina and Runway (Aug 17), Luma and Pika (Aug 17) — were on 1080p Seedance 2.5 within four days (Weekly Roundup — August 24, 2026)
- Seedance 2.0 mini (June 16, 2026): Lightweight tier launched via Dreamina at approximately $0.02/sec for emerging markets (Southeast Asia, Middle East, Africa, Europe, South America); US-excluded; roughly 7× below the standard 720p tier (Artificial Analysis lists the standard tier at ~$0.151/sec, but fal.ai’s live 720p-with-audio rate runs higher at ~$0.30/sec — treat the live gateway figure as current for commercial pricing)
- Higgsfield “Seedance Unlimited” (June 17, 2026): 30-day unlimited video generation on “Enhanced Seedance 2.0 Fast” (a purpose-built speed-optimized model from ByteDance’s BytePlus B2B cloud); 1-day, 7-day, and 30-day add-on tiers for Higgsfield subscribers; 480p–720p; Higgsfield is the exclusive non-ByteDance unlimited-Seedance surface globally per BytePlus commercial agreement
- Note: Benchmark position (August 28, 2026): #5 text-to-video (Elo 1,221), #2 image-to-video (Elo 1,190), #6 video-editing (1,039) — the Dreamina Seedance 2.0 720p entry. It held #1 image-to-video until MiniMax H3 Max took it in late August. Artificial Analysis retired its with-audio boards on August 11, 2026, so the mid-July with-audio figures this entry used to carry no longer have a live board behind them. Seedance 2.5’s BytePlus API opened July 16 (see What’s Coming). Copyright legislative battle remains a three-way standoff (Blackburn vs. White House vs. CLEAR Act)
The leading commercial model for character consistency and cinematic motion quality. Seedance 2.0 Pro’s Dual-Branch Diffusion Transformer generates audio and video simultaneously in a single pass. Its quad-modal input system accepts text, images, video, and audio in a single prompt. Multi-shot native storytelling and frame-level control over character appearance, object placement, and scene timing remain best-in-class for narrative work.
ByteDance’s official global API rollout was paused indefinitely in late February 2026 after the Motion Picture Association and major studios (Disney, Netflix, Paramount, Sony, Warner Bros.) issued cease-and-desist letters over copyright concerns. The “Face-to-Voice” feature was suspended in February 2026 after it was shown to clone voices from a single photo. Japan opened a separate inquiry over unauthorized anime character reproductions.
ByteDance relaunched the model in March 2026 as Dreamina Seedance 2.0 (Weekly Roundup — March 27, 2026) across markets in Africa, South America, the Middle East, and Southeast Asia. As of April 2026, Dreamina Seedance 2.0 is available in the US via CapCut (Weekly Roundup — April 11, 2026) — a significant reversal of the prior exclusion. The deployment comes with content restrictions: image-to-video generation from inputs containing real faces is disabled, and generation of unauthorized intellectual property is blocked. All output carries an invisible watermark for off-platform identification.
In June 2026, ByteDance expanded Seedance’s distribution through two parallel channels. On June 16, Dreamina Seedance 2.0 mini launched for emerging markets — Southeast Asia, the Middle East, Africa, Europe, and South America — at approximately $0.02/sec, roughly 7× below the standard 720p tier. The US is excluded. On June 17, Higgsfield launched “Seedance Unlimited” — 30-day unlimited access to an “Enhanced Seedance 2.0 Fast” model, delivered through an exclusive partnership with BytePlus (ByteDance’s global B2B cloud platform). Higgsfield is the only non-ByteDance surface with unlimited Seedance globally; the commercial infrastructure runs through BytePlus while Higgsfield owns the creator-facing product. Two channels, one underlying model family, different regulatory profiles — the structure gives ByteDance revenue from Western markets without operating a direct consumer product there.
At Volcano Engine FORCE on June 23, 2026, ByteDance confirmed that Seedance 2.0 now supports native 4K with 10-bit color depth — the underlying capability specification the current three-platform distribution wave runs on. At the same event, ByteDance previewed Seedance 2.5; see What’s Coming.
The copyright landscape around Seedance is a three-way Washington standoff: the White House’s National Policy Framework for AI (March 2026) stated that AI training on copyrighted works does not constitute infringement — the opposite of the Blackburn bill’s position. Separately, the bipartisan CLEAR Act (Schiff/Curtis) would require public disclosure of training data without resolving the fair use question either way.
Grok Imagine — SpaceXAI (formerly xAI)
Best for: Speed, low-cost audio-native API, rapid iteration, social media distribution
- Max resolution: Original Grok Imagine: 720p. Grok Imagine Video 1.5: native 1080p as of July 31, 2026 — closing the resolution gap tracked since Musk’s missed April commitment (a separate, higher-end “Grok Imagine Pro” tier was never named again after the original announcement; 1.5’s own 1080p update appears to be the resolution fix that commitment pointed to)
- Arena (Grok Imagine): #5 on Artificial Analysis T2V no-audio (Elo 1,235 — effectively tied with Kling 3.0 Omni at #4)
- Arena (Grok Imagine Video 1.5): #2 on the Artificial Analysis I2V with-audio arena (Elo 1,113, as of June 2026)
- Max duration: 15 seconds in a single request (added July 2026, per user reports amplified by Musk — no SpaceXAI release note); up to 30 seconds via chained extensions
- Audio: Original Grok Imagine: synchronized audio; Video 1.5: audio-native (synchronized in the same generation pass)
- Key feature: Video extension from frame; dual generation modes (Quality + Speed); native video understanding (Grok 4.3 Beta); fastest iteration cycle in the industry; Video 1.5 Fast variant (~25s for 6-sec 720p)
- Access: X Premium / SuperGrok subscription required
- API pricing:
- Engine: Aurora autoregressive MoE model on 110,000 NVIDIA GB200 GPUs
- Next: 1080p shipped in Video 1.5 (Jul 31, 2026) — voice references and text-to-video added the same update; watch whether pricing moves off the flat $0.08/sec now that resolution is no longer the differentiator
- Caution: Faced regulatory scrutiny over content moderation (UK ICO, France, California AG); image editing now restricted to paid subscribers
xAI shipped four major updates between January and March 2026: API launch (January 28), Grok Imagine 1.0 with 720p video and audio (February 3), Grok 4.20 (February 17), and video extension (March 2). The “Extend from Frame” feature lets users chain clips by continuing from the final frame, enabling sequences up to 30 seconds while preserving lighting, motion, and character positioning.
In April 2026, xAI released Grok 4.3 Beta with native video understanding — letting Grok analyze video as a coherent temporal sequence rather than as isolated frames. The understanding capability is distinct from Grok Imagine’s generation pipeline, but the two now stack: Grok can both generate and reason about video within the same model family. No other major lab currently offers vertical integration of native generation, native understanding, and platform-scale distribution under a single subscription.
On June 3, 2026, xAI shipped Grok Imagine Video 1.5 as a public API preview (Weekly Roundup — June 8, 2026) — the company’s first native image-to-video model. Per xAI’s docs: image-to-video with audio output, $0.08 per second, model alias grok-imagine-video-1.5-2026-05-30. It opened at #2 on the Artificial Analysis I2V with-audio arena, behind Seedance 2.0. Secondary reports described a higher-resolution tier; the docs list a single rate, so treat the tiered-pricing figures circulating elsewhere as unconfirmed by xAI.
On June 16, 2026, xAI moved Grok Imagine Video 1.5 to general availability across the Imagine API, grok.com, iOS, and Android simultaneously. The GA launch added a Video 1.5 Fast variant — a 6-second 720p clip in approximately 25 seconds, down from 40-plus seconds in preview. The single published rate remains $0.08/sec; xAI’s docs list no resolution-tiered pricing. At GA, Grok Imagine Video 1.5 holds #2 on the Artificial Analysis I2V with-audio arena (Elo 1,113), behind Seedance 2.0 (1,194). xAI’s #1 claim at launch wasn’t borne out by the live board. The “$4.20/min” figure circulating refers to the original Grok Imagine’s $0.07/sec@720p rate (a different, older model with no native audio) — not to Video 1.5.
Grok Imagine’s API pricing undercuts every audio-capable commercial tier except the original Grok Imagine itself, and as of August 2026 it no longer trades that price on a resolution ceiling. The 1080p tier Elon Musk telegraphed for late April 2026 missed its April window by more than three months, but Grok Imagine Video 1.5 shipped native 1080p on July 31, 2026, alongside voice references and text-to-video generation — closing the gap this page has tracked since the missed commitment. The single published rate remains $0.08/sec with no resolution-tiered pricing on any SpaceXAI docs page, so 1080p appears to be included at the existing rate rather than a new paid tier. The launch puts the two subjects of RCTV’s Grok vs. Seedance comparison at one and two on the I2V arena.
SpaceX acquired xAI in an all-stock merger on February 2, 2026, valuing the combined company at $1.25 trillion — CNBC called it “the biggest merger of all time.” On July 6, 2026, the renamed company confirmed the shift on X: “We are now @SpaceXAI.” Nothing below changed with it — Grok, Grok Imagine, and the pricing and model names above are unchanged, and Artificial Analysis already credits the model to SpaceXAI.
Runway Gen-4 Turbo / Aleph 2.0 — Runway
Best for: Stylized content, VFX aesthetics, professional ecosystem, real-time avatars, agentic production, multishot edit propagation
- Max resolution: 1080p (Gen-4 Turbo); 1080p multishot up to 30s (Aleph 2.0 in Edit Studio); 720p real-time (Characters)
- Audio: Supported
- Key features: Motion brushes, style control, API maturity (Gen-4 Turbo); Characters real-time avatar API (GWM-1); Runway Agent (conversational end-to-end production, May 13, 2026); Aleph 2.0 multishot edit propagation in Edit Studio (May 21, 2026)
- Access: From $12/mo (runwayml.com); Gen-4.5 also via Adobe Firefly (Creative Cloud subscription); Runway Agent and Edit Studio (Aleph 2.0) at runwayml.com on all paid plans
- API: Most mature video generation API available; Characters API at dev.runwayml.com
- Next: Standalone Gen-4.5 launch on runwayml.com still pending; real-time video model research preview on Vera Rubin hardware (sub-100ms TTF); independent hands-on for Aleph 2.0 against Veo 3.1 and Gemini Omni editing capabilities
- Ruby (August 21, 2026): A separate model converting SDR video — uploaded or Runway-generated, up to 30 seconds — to 16-bit HDR, delivering ProRes and EXR sequences at 10/12-bit BT.2020 with PQ or HLG. Max plan and Enterprise tiers only. A post-production/colour-pipeline tool rather than a generation capability, which is why it changes no benchmark position
- Note: Characters is an enterprise API product built on GWM-1, separate from the Gen-4 Turbo generation pipeline
Runway leads in non-photorealistic and stylized video — VFX-oriented aesthetics, abstract content, and artistic directions where other models default to photorealism. Gen-4 Turbo has the most mature professional ecosystem with motion brushes, scene consistency tools, and a robust API. Runway closed a $315M Series C in February 2026 at a $5.3B valuation.
In March 2026, Runway launched Characters — a real-time video agent API built on its GWM-1 world model (Weekly Roundup — March 13, 2026). Characters generates fully conversational AI avatars from a single reference image with no fine-tuning required. The avatars sustain realistic lip-sync, facial expressions, eye contact, and gesture across extended multi-minute conversations, running at 24fps at 720p in real time. BBC and Silverside are early enterprise partners.
At NVIDIA GTC in March 2026, Runway demoed a research preview of a new real-time video generation model (Weekly Roundup — March 20, 2026) running on NVIDIA Vera Rubin hardware — achieving time-to-first-frame under 100ms for HD video. Gen-4.5 became accessible via Adobe Firefly’s multi-model video hub in April 2026 (Weekly Roundup — April 27, 2026) — Runway’s first major distribution beyond its own platform.
In May 2026, Runway launched Runway Agent (Weekly Roundup — May 18, 2026) — a conversational creative partner that runs ideation, generation, sound design, and editing end-to-end from a chat interface. The same month, Runway opened a Tokyo office and committed $40M to Japan, its third-largest market, with its enterprise base tripled in twelve months.
In May 2026, Runway released Aleph 2.0 (Weekly Roundup — May 25, 2026) — an upgraded video editing model that propagates a single-frame edit across the rest of a clip while preserving everything else. Multishot sequences up to 30 seconds at 1080p, edited across cuts in one pass instead of shot-by-shot. Available on all paid Runway plans on the desktop web app.
In July 2026, Runway launched Runway Dev (Weekly Roundup — July 13, 2026) — a developer and enterprise platform putting one API over Runway’s own frontier models (Gen-4.5, Aleph 2.0, Act-Two) plus third-party models including Seedance, GPT Image 2, and ElevenLabs. Recipes package prompting and workflow expertise into single API calls; Workflows chain multiple models into custom pipelines. Runway says it already runs production for Adobe, ElevenLabs, Shutterstock, Figma Weave, Gamma, and Silverside, with SOC 2 Type II compliance and IP indemnification (company-stated). It gives the aggregator strategy a monetization shape — a business product, not just a creator UI.
Later in July, Runway pushed the aggregator play from hosting to routing: Media Router (July 23, 2026) lets a team define what “best” means once — cost, quality, or latency — and automatically selects the video, image, or audio model per request, live in Runway Dev (Weekly Roundup — July 27, 2026).
Pika 2.5 / Pika Agents — Pika Labs
Best for: Budget-conscious creators, rapid iteration, social media content; multi-model agent orchestration via Pika Agents
- Max resolution: 1080p (Pika 2.5); 480p real-time (PikaStream 1.0)
- Max duration: 42 seconds (clip); persistent for live (PikaStream)
- Audio: Supported (Pika Video native; ElevenLabs / MiniMax / OpenAI Whisper via Pika Agents)
- Key feature: Pikaswaps, Pikaffects, fast batch generation (Pika 2.5); PikaStream 1.0 for live agent video; Pika Agents for multi-model orchestration over Kling, Veo, Seedance, MiniMax, and Sora
- Access: From $8/mo (lowest entry price among major models); Pika Agents available at pika.me and across 17+ platform surfaces
- API: Available
The most accessible entry point to AI video generation. Pika’s strength is speed and volume — generate 20-30 variations of a concept in minutes, then refine. Features like Pikaswaps (face/object replacement) and Pikaffects (style transfer) add creative flexibility at a price point that undercuts every competitor.
In April 2026, Pika launched PikaStream 1.0 — a real-time AI video engine for live agent meetings at 24fps/480p with ~1.5s speech-to-video latency and persistent identity across calls.
In late April 2026, Pika reintroduced its product line as Pika Agents — a multi-modal AI creative partner that orchestrates other companies’ video models from a conversational interface. The video roster includes Pika’s own model alongside ByteDance’s Seedance 2.0, Kuaishou’s Kling, MiniMax, Google’s Veo 3, and OpenAI’s Sora. On audio: ElevenLabs, MiniMax Music and Voice, OpenAI Whisper. On images: Gemini, ChatGPT Images 2, SeedDream. The agents run inside Slack, Telegram, WhatsApp, Discord, X, Notion, GitHub, Figma, and a dozen other surfaces with persistent memory and personality across sessions.
In July 2026, Pika shipped Director’s Suite (Weekly Roundup — July 13, 2026) — an experimental AI timeline editor for agent-driven video (storyboarding, clip generation, and chat/voice-guided editing in one interface), alongside a VFX Skill and a Seedance-4K-powered 4K-VFX Skill that edits an existing clip from a single prompt while preserving faces, gestures, audio, and camera moves.
MiniMax H3 — MiniMax
Best for: Omni-modal generation and editing at the value tier — 2K with native stereo audio, priced and benchmarked from day one; the first arena leader in generative video editing
- Max resolution: 2K (2560×1440)
- Max duration: 5–15 seconds
- Frame rate: 24fps
- Audio: Native stereo audio
- Arena (August 28, 2026): #4 text-to-video (Elo 1,227); #3 image-to-video (Elo 1,184); #2 video-editing (Elo 1,129) — H3 held #1 on video-editing from its August 3 debut until Wan 3.0 displaced it on August 19. Above it on both generation boards now sits MiniMax H3 Max, a third-party post-train of H3’s own open weights
- Key feature: Omni-modal input (text/image/video/audio references, up to 9/3/3 respectively); in-place character, scene, dialogue, and voice editing; the video-editing arena win
- Access: Consumer via the Hailuo AI app; API live day one
- API: $0.13/sec at 2K, model ID
MiniMax-H3 - Weights: Shipped August 3, 2026 (H3-Base, 33.1B-parameter dense omni-modal transformer, day-one ComfyUI/vLLM-Omni/SGLang support) — but the Community License excludes the US, EU, UK, and South Korea from its “Applicable Territory” entirely, and local self-hosted generation caps at 768p, not the API’s 2K — the 2K upscaling module isn’t open-sourced and still routes through MiniMax’s hosted infrastructure
- Caution (license): Not usable at all, per MiniMax’s own terms, for anyone in the US, EU, UK, or South Korea — the four markets excluded from the Community License’s territory grant
- Caution: Not a signatory to the EU’s Article 50 transparency Code of Practice as enforcement began August 2, 2026 — along with every AI-video lab tracked on this page except Google and Meta
MiniMax launched H3 — successor to its Hailuo 2.x line — on July 31, 2026, and debuted directly into the field’s top tier: top-three finishes on both major generation boards, and an outright #1 on Artificial Analysis’s newer video-editing leaderboard, a blind-preference category the T2V and I2V boards don’t test. It’s the strongest arena debut by any non-Google lab this page tracks.
The launch posture stood out on its own terms too. Where ByteDance’s Seedance 2.5 went “official” the same week without a published price (see Seedance), H3’s API was live and priced from hour one. MiniMax also kept its open-weights promise fast — weights landed August 3, roughly 60 hours after the “within days” pledge, the first credible challenger to HappyHorse’s long-unfulfilled open-weights claim in this same tier. But “open weights” comes with real fine print: MiniMax’s own license shuts out the US, EU, UK, and South Korea from using the weights at all, and the self-hosted build doesn’t reach the API’s 2K ceiling. Read the claim precisely — H3 is genuinely open-weights somewhere, for some resolution, which is still more than HappyHorse has delivered anywhere.
Then the open weights did the thing open weights are supposed to do. In late August, fal Research shipped H3 Max, a post-train of H3 that now outranks the official H3 endpoint on both generation boards. MiniMax’s own account called it “fantastic work” and “exactly why we build with open weights.” The lab that shipped the model is now #4 on a board its own base weights lead — a strange kind of win, and a real one.
Wan 3.0 — Alibaba
Best for: Long single-pass takes and generative video editing — the current #1 on two of the three Artificial Analysis boards this page tracks
- Max resolution: 1080p (480p / 720p / 1080p tiers; no 4K)
- Max duration: 2–30 seconds in a single pass at 30fps, no chaining and no stitching; a
-1duration hands the model smart duration selection (default 5s) - Key feature: Document-to-video — accepts docx/doc, xlsx/xls, pptx/ppt, pdf, txt, key, pages, numbers and md, ≤100MB and ≤50 pages, as generation source material. Reference-to-video takes up to 10 images, 5 video clips (≤15s total) and 5 audio clips (≤15s total). Two model IDs:
wan3.0-videoand the speed-tierwan3.0-video-prime - Arena (August 28, 2026): #1 text-to-video (Elo 1,241) and #1 video-editing (Elo 1,189), displacing Gemini Omni Flash and MiniMax H3 respectively. Also debuted #1 on arena.ai’s separate Video Edit Arena
- Access: Invite-gated beta opened August 6, 2026 via Alibaba Cloud Model Studio (Beijing and Singapore endpoints) and Qwen Cloud. Label conflict, logged: trade coverage on August 24–25 described an official release, while Alibaba’s own Model Studio API reference still read “Currently in preview.” This page tracks it as preview until the vendor docs say otherwise — vendor documentation governs for procurement. Re-checked 2026-09-01 against a page Alibaba itself stamped “Updated at: 2026-09-01”: it still reads “Currently in preview.” Six endpoint regions, not five — Singapore, Beijing, Hong Kong, Tokyo, Frankfurt and US Virginia — with cross-region calls failing outright. Also live on Runway: the consumer app since August 24, the Dev API since August 26, billed at 5/10/20 credits per second for 480p/720p/1080p. This is the gap that cost R#19 its discovery score: no monitoring lane covers vendor API docs
- API: List: $0.05/sec (480p) · $0.10/sec (720p) · $0.20/sec (1080p) — so a 30-second 1080p take lists at $6.00. Alibaba’s pricing page carries a live limited-time 30% discount on this tier, which puts the same take nearer $4.20 at the currently-billed rate. The separate speed tier
wan3.0-video-primeis not discounted: $0.068/$0.14/$0.28 per second. Audio is on by default and, in Alibaba’s own words, “enabling or disabling audio does not affect pricing” - Weights: None. This is a closed commercial release, distinct from the open-weights Wan 2.7 line —
github.com/Wan-Videostill tops out at Wan2.2 - Caution: Invite-gated. A model can lead a blind-preference board and still be something most readers cannot buy
Alibaba opened Wan 3.0’s beta on August 6, 2026 and it charted for the first time on August 19 — absent from every Artificial Analysis board on the August 10, 12, 14 and 17 fetches, then straight in at #1 on two of them. It has held both since.
The interesting part is what it is, not where it ranks. Alibaba already runs the most-downloaded open-weights video line in the field (Wan 2.7); Wan 3.0 is the same company shipping a closed, invite-gated, commercially-licensed model that beats it. The 30-second single-pass ceiling is the technical claim worth watching — every other model on this page reaches long-form by chaining shorter takes, which is where continuity breaks.
Distribution moved the way it now always moves: Runway on August 24, Pika and ComfyUI within two days, then ArtArch, Media.io, OiiOii AI, Picsart, PowerDirector, A2E, PixelDojo and DeepInfra inside a week — eleven surfaces in eight days, the same absorption pattern already documented for Seedance 2.5, FLUX 3 and MiniMax H3.
MiniMax H3 Max — post-trained by fal
Best for: Speed. A 5-second 768p clip in under 3 seconds — the first model this page tracks that generates faster than the clip plays
- Max resolution: 768P — the resolution enum on fal’s API schema offers 480P and 768P only, with no 1080p or 2K tier. fal’s own X post says “720p” for the same clip; the model page governs. (This entry briefly carried “1344×768” — a plausible 16:9-at-768-height inference that no source states. Removed 2026-09-01.)
- Max duration: 10 seconds on the free tier; 5-second text-to-video and image-to-video are the standard calls
- Arena (August 28, 2026): #1 image-to-video (Elo 1,202), displacing Dreamina Seedance 2.0 720p; #3 text-to-video (Elo 1,235), ahead of the official MiniMax H3 endpoint at #4
- Speed claim: fal claims 35× the throughput of the official H3 endpoint; Design Arena independently measured 50×; against a quality-matched peer group the figure is 15×. These are three different comparisons and should never be blended
- Access: Five 5-second generations a day with no account; signing in raises it to 15 a day at up to 10 seconds (raised from 5 on August 31, text-to-video and image-to-video both)
- API: $0.08/sec at 768P ($4.80/min), $0.05/sec at 480P — list rates, from fal’s own model page. A 75%-off promotional rate is live until September 7, 2026: $0.02/sec at 768P, $0.0125/sec at 480P. (Corrected 2026-09-01: this page and R#19 described the promotion as half price for 14 days. It is 75% off and it ends September 7.)
- Base model: A post-train of MiniMax’s open-weight H3 — not an independent architecture. It inherits H3’s lineage and none of H3’s Community License territory exclusions, because fal serves it hosted
- Caution: fal’s own board label reads “#1 image-to-video with audio.” That rank is real; see the leaderboard note under Quick Reference for what Artificial Analysis’s audio labelling actually means
- Note on the price we would not print: R#19 deliberately published no per-second rate, because fal’s pricing widget did not render and a summarizing fetch produced figures we refused to launder. The page rendered cleanly on 2026-09-01 and the real list rate is $0.08/sec — the same number the summarizer had guessed. Refusing to print an unverified figure that later proves correct is the process working, not a miss
This is the first time a third-party fine-tune tops a leaderboard this page tracks, and it is worth being precise about what that demonstrates: not that fal built a better model than MiniMax, but that MiniMax’s decision to publish open weights let someone else find headroom in them. MiniMax’s own account endorsed the result rather than disputing it.
The speed number is the story (Weekly Roundup — August 31, 2026). Generation faster than playback is a category threshold, not a benchmark position: it is the precondition for live and interactive video rather than rendered-then-delivered video. fal tested that directly with fal.live on August 31 — interactive AI livestreams driven by viewer prompts, running on “H3 Max Director,” an autoregressive variant holding up to two minutes of native context. fal paused it roughly three hours after launch, citing quality and experience work, without disclosing what broke. Treat the threshold as real and the sustained-load behaviour as unproven.
Sora 2 — OpenAI
Status: Discontinued March 24, 2026; consumer app shutdown executed April 26, 2026. RCTV analysis →
OpenAI announced Sora’s discontinuation on March 24, 2026 — the app, the API, and the Disney licensing deal announced with it in December 2025. The stated reason was compute reallocation toward “world simulation for robotics.” The numbers tell the fuller story: estimated $15M/day peak inference cost against $2.1M in total lifetime in-app revenue, and a 66% download decline from its November 2025 peak to February 2026. Sora is removed from active tracking. See Weekly Roundup — March 27, 2026 for the full breakdown.
Shutdown timeline: The Sora consumer app and web interface went dark on April 26, 2026 (Weekly Roundup — April 27, 2026) — the export window closed at that time. The Sora API remains accessible through September 24, 2026, giving developers time to migrate integrations before the model line fully retires.
The Agentic Layer: Orchestration on Top of the Models
The model is no longer the product. The six agents in this section don’t generate video — they decide which model generates video, and on what schedule, across what surfaces. Think of the commercial models above as engines: Veo 3.1, Kling 3.0, Seedance 2.0. The agents are steering wheels. The operator-relevant question has shifted up a layer: not “which model is best?” but “which agent puts the right model on the right shot, in the right workflow, with the right context carried forward?” That question didn’t exist twelve months ago. Six companies have already shipped an answer.
| Agent | Vendor | Orchestrates | Single/Multi-model | Surfaces | What it automates | Access |
|---|---|---|---|---|---|---|
| Luma Agents | Luma AI | Ray 3.2 / Ray 3.14, Veo 3, Nano Banana Pro, Seedream, Seedance 2.0, ElevenLabs | Multi | Luma platform, enterprise | Multi-shot composition; sustained character and style across shots; full campaign packages | Enterprise; Luma platform subscription |
| Pika Agents | Pika Labs | Pika Video, Kling, Veo 3, Seedance 2.0, MiniMax, Sora API; audio via ElevenLabs, MiniMax, Whisper; image via Gemini, ChatGPT Images 2, SeedDream | Multi | Slack, Telegram, WhatsApp, Discord, X, Notion, GitHub, Figma + 17 other surfaces | Conversational prompting; cross-model orchestration; persistent memory and personality across sessions | Pika subscription tiers (pika.me) |
| Runway Agent | Runway | Runway Gen-4 Turbo, Aleph 2.0, Edit Studio | Single-vendor | Runway web app (app.runwayml.com) | Concept ideation, multi-shot generation, sound design, editing — end to end from a single conversation; Agent Skills (slash-command campaign/ad/localization workflows) | All paid Runway plans ($12/mo+) |
| Higgsfield Supercomputer | Higgsfield AI | Higgsfield video stack + Claude Opus 4.8 / Claude Fable 5, GPT-5.5 Pro, Gemini 3.1 Pro / Gemini Omni Flash | Multi | Browser, Telegram, MCP + 30 integrations (Slack, Google Drive, Notion, Figma, Gmail) | Marketing, production, and creative-direction workflows; faceless Explainer docs up to 10 min; audio/voiceover/dubbing (50+ languages); research-to-document conversion; scheduled content tasks | Higgsfield subscriber plans; MCP free (100 credits, 3-day trial) |
| Adobe Firefly AI Assistant | Adobe | Photoshop, Premiere Pro, Lightroom, Illustrator, Express, Firefly — full Creative Cloud stack | Single-vendor | Standalone Firefly web app + embedded in each Creative Cloud app; Premiere Pro with project-metadata access | Multi-step CC workflows via Creative Skills; multi-app handoff; format conversion; Frame.io feedback integration | Creative Cloud subscription; public beta since April 27, 2026 |
| HeyGen HyperFrames | HeyGen | HeyGen’s own HyperFrames production models | Single-vendor | hyperframes.heygen.com | Context-based routing across 9 production workflows (launch video, music video, captions/overlays, +6 more) — no manual workflow selection | HyperFrames platform (hyperframes.heygen.com) |
Luma Agents
Luma AI shipped the first production-grade conversational agent in this category in March 2026. By March 10, deployments were live at Publicis Groupe and Serviceplan Group — no beta waiting period. The agent works from a Luma Uni-1 reasoning layer that plans and coordinates across video, image, audio, and text before generating anything. It calls Luma’s own Ray 3.2 / Ray 3.14 for video, Google’s Veo 3 for photorealistic shots, Nano Banana Pro, ByteDance’s Seedream, and ElevenLabs for voice. Seedance 2.0 joined the roster in May 2026. The Mazda commercial deliverable in April 2026 — Johannesburg agency Boundless produced Mazda’s first AI-generated commercial using Luma Agents in under two weeks — is the most credible production-deployment signal any AI video agent has produced.
The critical technical feature: persistent context across the full asset suite. Luma Agents remembers what was generated in earlier steps and can revise upstream elements when downstream evaluation surfaces a problem. That is what makes it an orchestrator rather than a fancy prompt box.
Pika Agents
Pika’s April 28, 2026 launch was the moment the agentic-orchestration pattern became industry news rather than one lab’s experiment. Pika Agents orchestrates a broader model roster than any competitor — Kling, Veo 3, Seedance 2.0, MiniMax, Sora’s API, and Pika’s own model on the video side; ElevenLabs, MiniMax, and Whisper on audio — all from a conversational interface that runs inside 17 surfaces where creators already work. RCTV covered this as the R#2 lede because the framing was explicit: “the prompt is no longer the product.” That line has since become the editorial spine for this entire category.
PikaStream 1.0 (April 2, 2026): the real-time video engine that runs inside Pika Agents as its live-avatar capability — 24fps at 480p, ~1.5-second speech-to-video latency, persistent identity across calls. It is a streaming runtime, not a standalone orchestrator; the agent layer above it is Pika Agents.
Runway Agent
Runway shipped its agent in May 2026 — single-vendor, full Runway stack. The positioning is end-to-end: Runway Agent handles concept ideation, generation via Gen-4 Turbo or Aleph 2.0, sound design, and editing inside the same conversation thread. Runway’s “single-vendor” constraint is a deliberate product posture, not a technical limitation — they own the generation and editing stack, so the agent never needs to leave it. Whether that narrows or focuses the use case depends on whether the operator’s workflow already lives in Runway. Available on all paid plans, starting at $12/month.
Runway added an enterprise anchor and a new capability layer in July 2026. On July 1, Runway announced a partnership with Bertelsmann, integrating its models “across Bertelsmann’s global portfolio of businesses” — RTL Group, BMG, and Bertelsmann Marketing Services. RTL’s Fremantle already runs Runway through its in-house AI studio, Imaginae; BMG uses it for artist marketing visuals — a genuine enterprise-adoption story, not a single pilot. The next day, Runway shipped Agent Skills: a slash-command layer on top of Runway Agent for building ad campaigns, commercials, and localized ads on demand. Runway is now the third vendor, after Pika and Luma, to converge on “Skills” as the name for a portable, reusable agent workflow.
In July 2026, a new third-party benchmark for AI video agents — Physion-Arc 1.0 from Physion Labs — ranked Runway Agent 2.0 #1 of six agents tested (against Luma, MiniMax, Kling, Utopai Studios, and TapNow), leading every dimension it scored: narrative coherence, cinematic language, production quality. The caveats are load-bearing: the benchmark is brand-new, its methodology page is unreachable, its two primaries disagree on the test’s own scale (Physion Labs cites 100 screenplays / 600 videos; Runway’s amplification cites “30 cinematic prompts and 16 metrics”), and Runway is the loudest amplifier of a result that crowns Runway — no independent replication yet (Weekly Roundup — July 27, 2026).
Higgsfield Supercomputer
Higgsfield’s Supercomputer, launched in mid-May 2026, is the most enterprise-positioned agent in the set. “Supercomputer” is the framing: orchestrate Claude, GPT-5.5, and Gemini models alongside Higgsfield’s video stack to plan and execute full content campaigns end to end. On May 19, 2026 — the same day as Google I/O — Higgsfield updated the orchestration layer to Gemini, describing the swap as “8× cheaper, 3× faster.” On May 29, Higgsfield upgraded its Claude backbone to Opus 4.8, and on May 30 shipped Higgsfield Reframe — an MCP-native aspect-ratio reframing tool available inside Claude. The multi-model roster is the broadest LLM coverage in this category. Distribution extends via browser, Telegram, and 30+ third-party integrations. In a 12-day sprint through early June 2026 (Weekly Roundup — June 15, 2026), Higgsfield added five external surfaces — Claude MCP (May 28), Adobe Premiere/After Effects plugins (May 29), Figma (June 4), a Minecraft mod (June 5), and a DaVinci Resolve plugin (June 8) — and folded Grok Imagine 1.5 into its own platform, the clearest expression yet of the routing bet: a competitor’s model generating video inside Higgsfield’s surface. The pipeline expanded again at the end of June and start of July: audio generation landed in the MCP on June 27 — voiceovers, voice cloning, and dubbing across 50-plus languages via Seed Audio 1.0 and ElevenLabs v3 — and on July 2 Higgsfield shipped Explainer, generating faceless documentaries up to 10 minutes from an auto-researched script, built on Claude Fable 5 plus Gemini Omni Flash. The same day, Higgsfield made its MCP free: 100 credits and a 3-day full-access trial, removing the price floor entirely two days after Google’s own Omni Flash preview shipped. On July 3, Higgsfield pushed into the professional edit suite itself — plugins bringing Gemini Omni Flash and Seed Audio 1.0 into Adobe Premiere Pro and DaVinci Resolve Studio timelines, for background cleanup, multi-shot generation, and 18-language dubbing without leaving the NLE. Higgsfield’s Supercomputer page carries the full capability description.
The round closed. August 17, 2026: $400M Series B at a $5.4B valuation — DST Global led, with Goldman Sachs Alternatives, Valor Capital and Tribe Capital participating. That is roughly 4× the ~$1.3B mark of eight months earlier. Disclosed alongside it: $700M annualized revenue, 30M users across 200 countries, and 390 Fortune 500 enterprise clients. (This page tracked the round in July as “reportedly in talks to raise $300–500M at a $5B valuation” on a ~$500M run rate; the closed terms came in above the reported range on both figures.) CEO Alex Mashrabov frames the raise as compute rather than growth capital — “Video is one of the most compute-intensive domains in AI,” putting one minute of video at the resource cost of 60,000 words of text. The revenue trajectory is still the more interesting number than the valuation: ~$200M at the end of 2025 to $700M annualized eight months later, in a category where nobody else discloses revenue at all (Weekly Roundup — July 27, 2026).
Adobe Firefly AI Assistant
Adobe is the incumbent here, and the only one that arrived via acquisition of creative-workflow context rather than from a blank slate. Firefly AI Assistant — previewed as Project Moonlight at MAX, public beta April 27, 2026 — orchestrates the Creative Cloud stack conversationally: Premiere Pro with full project metadata access, Photoshop, Lightroom, Illustrator, Express. Creative Skills are the execution layer — predefined multi-step workflows that fire from a single natural-language instruction. The operative claim is multi-app handoff without context loss; the operator verdict is still forming in the public beta.
The AI-native vs. incumbent frame is worth naming: Luma, Pika, Runway, and Higgsfield are all building agents on top of AI video first. Adobe is extending an agent over a suite it already controls. The operator question these two approaches answer is different — and which model wins the relationship depends on whether the operator’s workflow is already Premiere-centric or starting from scratch.
HeyGen HyperFrames
HeyGen shipped HyperFrames “Next Generation Skills” on June 26, 2026 — nine production workflows (launch video, music video, captions and overlays on video, and six more) routed by context rather than by menu. HeyGen’s own framing: “it knows which one you mean from context and routes there on its own.” There’s no workflow picker to click through; the agent infers intent from the prompt and dispatches to the right pipeline. At 1.53 million impressions, it’s the highest-engagement AI-video product launch RCTV tracked in the June coverage window — a genuine adoption signal, even accounting for the platform’s own promotional push. HyperFrames is a single-vendor agent, orchestrating HeyGen’s own production models rather than a third-party roster — closer to Runway Agent’s posture than to Pika Agents’ multi-model bet. Surfaces at hyperframes.heygen.com. In July 2026, HyperFrames product engineer James Russo detailed the architecture in an X-Article, “HTML Is All Agents Need” — HyperFrames renders agent-authored video from plain HTML/CSS/JS rather than a proprietary timeline format, on the reasoning that the model already knows HTML fluently. Cited traction: 1.3M videos for 267K creators in 90 days, 32K GitHub stars, open source, running inside Claude Code, Codex, and Cursor (Weekly Roundup — July 13, 2026).
MCP server inside Claude Design (August 7, 2026). HyperFrames shipped an MCP server that puts the whole assembly step behind a share action: animate a design, hit share, and a HyperFrames agent scores the result, adds sound effects, tightens the motion, and returns a finished MP4 — “Just send it to HyperFrames. We’ll take it from there,” in the vendor’s own words. No human touches the step between design and shippable video. This is the agentic-assembly pattern reaching a design surface rather than a video platform (Weekly Roundup — August 17, 2026).
LiveAvatar concurrency limits removed (August 25, 2026). HeyGen dropped the cap on concurrent LiveAvatar sessions entirely — “run 1 avatar or 10,000 at once, same API, full-body 1080p, down to $0.01/min at scale.” A capacity and pricing change rather than a new model, but the direction is worth noting: the constraint on real-time avatar deployment moves from what the vendor will allow to what the customer will pay.
What to watch
Two August developments sit directly on these axes. ComfyUI shipped Comfy MCP local and open-source on August 18, 2026 — “the #1 ask after Comfy Cloud MCP shipped in June.” An agent now sees a user’s actual local install (every custom node, every model on disk), can fetch files, start the instance, and get a workflow running end to end. The self-hosted layer now has the same agent-driven entry point the hosted platforms have been building all year, which is the first time this category extends below the API line. Separately, Tencent Hunyuan previewed HyCreator (August 20–21, 2026), an agent harness claiming an “End2end Auto Mode” that generates 10-minute-scale films with “zero human intervention,” alongside a real-time interactive-editing mode; early access is application-gated. The claim is far ahead of anything demonstrated in this section and is carried here as a claim, not a capability — no independent hands-on exists yet.
Three axes will decide how this layer develops. First: single-vendor vs. multi-vendor convergence. Runway’s Runway-only posture contrasts with Pika’s eight-model roster; whether Runway opens to third-party models is the specific question. Second: AI-native startup agents vs. legacy-incumbent agents. Adobe’s Creative Skills framework is embedded in the world’s most-used professional NLE; Pika Agents runs in Slack. Neither is clearly winning the operator relationship yet. Third: agent-to-agent interoperability. None of the six currently calls another vendor’s agent — they call models. The day Pika Agents calls Runway Agent, the competitive dynamics of this category change entirely.
Open-Source & Local Generation
The open-source AI video ecosystem has matured significantly, making local generation on consumer hardware a viable option for privacy-conscious creators and developers.
LTX-2.5 — LTX
Best for: Local/desktop generation, consumer GPU workflows, high-frame-rate output
- Vendor: LTX — an independent company since July 8, 2026, spun out of Lightricks (which keeps its consumer apps, e.g. Facetune); CEO Zeev Farbman
- Current release: LTX-2.5 (August 11, 2026). The spec bullets below are the verified LTX-2.3 card; a full 2.5 spec card is pending primary verification of resolution, duration, and frame-rate figures — the vendor’s launch post does not state them.
- Max resolution: 4K native (true 4K, not upscaled)
- Max duration: 20 seconds
- Frame rate: Up to 50fps (24/48fps options also available)
- Audio: Native synchronized audio (improved HiFi-GAN vocoder)
- Portrait mode: Yes (9:16, up to 1080×1920)
- Hardware: Runs on GPUs with 12GB+ VRAM; optimized for RTX 50 Series (2.5× faster via NVFP4)
- Integration: ComfyUI native; standalone desktop video editor (shipped March 2026)
- License: the LTX-2.x Community License Agreement (dated August 11, 2026) — open weights, but not an OSI-open licence. (Corrected 2026-09-01: this page said “Apache 2.0” in three places while its own licence bullet described a $10M revenue threshold Apache 2.0 cannot contain. Verified against the licence text itself.) Free commercial use below $10,000,000 annual revenue, aggregated across affiliates under common control; at or above, a separate paid Commercial Use Agreement is required. Four restrictions worth knowing before you build on it: derivatives are defined broadly — fine-tunes, adapted weights, distillation via intermediate representations, and training on LTX-2.x-generated synthetic data all count; commercial use may not train a competing model on the weights or their outputs; military, weapons and nuclear applications are barred; and disabling or circumventing watermarking or provenance functionality is barred in two separate clauses. LTX states it intends the model to qualify as free and open-source under EU AI Act Article 53(2), with an explicit carve-out that the derogation does not extend to Article 53(1)(c) and (d)
A comprehensive rebuild released in March 2026 (Weekly Roundup — March 20, 2026): a new VAE for sharper detail, a 4× larger text connector for better prompt understanding, and an improved HiFi-GAN vocoder for cleaner native audio. The model ships alongside a dedicated desktop video editor, making the entire local pipeline accessible without a ComfyUI node graph.
Key capabilities: native portrait mode (9:16 up to 1080×1920), last-frame interpolation for seamless clip chaining, and 24/48fps output options. At GDC 2026, NVIDIA announced 2.5× performance gains on RTX 50 Series via NVFP4 quantization, 60% lower VRAM usage, and RTX Video Super Resolution for ComfyUI delivering 4K upscaling 30× faster than competing local alternatives. The ComfyUI App View strips the node-graph interface into a simplified prompt-in/video-out UI for non-technical users.
LTX became its own company, then shipped LTX-2.5 (August 2026). On July 8, 2026 LTX spun out of Lightricks as an independent open-world-models company, taking the video line and CEO Zeev Farbman with it; Lightricks retained its consumer app business. LTX-2.5 followed on August 11 — the vendor’s own blog describes Diffusion Fidelity Rendering, native multishot generation that holds continuity across cuts, a custom Gemma 4 text-encoder backbone, and full HDR ACES support. Within about ten hours it had open weights on Hugging Face, Day-0 native ComfyUI support, and a Runway integration. Licensing is unchanged in shape: free for organizations under $10M in annual revenue measured across the whole entity, paid above. By LTX’s own account the open-weights line has passed 22 million cumulative Hugging Face downloads since the base LTX-2 architecture launched in January 2026 — a vendor claim, not an independently audited figure.
The launch claim and the board disagree, and the gap is large. LTX’s commissioned blind tests put LTX-2.5 at a 67% win rate, ahead of Seedance 2.5 (65%) and Gemini Omni Flash (55%) — figures reported by VentureBeat, which flags them as vendor-commissioned rather than independent. Artificial Analysis, blind-preference and independent, puts both LTX-2.5 tiers at Elo 1,060 on the audio text-to-video board, ranks 21–23 of 33 — roughly 180 Elo behind Omni Flash. On the no-audio board they read 1,213 (Fast) and 1,204 (Pro). Two weeks in, the Pro tier has not measurably beaten the Fast tier on any board — they are tied within their error intervals on text-to-video and Fast leads on audio image-to-video (1,044 to 1,013), while Pro lists higher. Treat the 67% as a marketing figure with a named methodology gap, not as a result (Weekly Roundup — August 17, 2026).
Wan 2.7 — Alibaba (Tongyi Lab)
Best for: Multi-task video generation with Thinking Mode, open-source flexibility
Not to be confused with Wan 3.0, which Alibaba opened in invite-gated commercial beta on August 6, 2026. Wan 3.0 is closed — no published weights, no GitHub listing — and is tracked in the commercial section. This entry covers the open-weights 2.x line.
- Max resolution: 1080p
- Max duration: 2–15 seconds
- Task types: T2V, I2V, video continuation, reference-to-video (up to 5 persons), video editing
- Key feature: Thinking Mode (chain-of-thought reasoning before generation)
- Integration: ComfyUI 0.18.5+, Alibaba Cloud Model Studio, wan.video
- License: Open source
- Benchmark (July 2026): A dated checkpoint, Wan2.7-260612, debuted #2 on the Artificial Analysis text-to-video-with-audio board (Elo ~1,160) — the window’s only fresh top-5 leaderboard movement (Weekly Roundup — July 13, 2026)
Alibaba’s Wan 2.7, released April 3, 2026 (Weekly Roundup — April 17, 2026), is a major upgrade from the 2.2 line. The headline feature is Thinking Mode — a chain-of-thought reasoning approach where the model analyzes the prompt, plans composition, then generates. This produces noticeably more coherent output with fewer artifacts than single-pass generation.
Wan 2.7 Video unifies five task types in a single model: text-to-video, image-to-video (first-frame, first-and-last-frame, audio-driven), video continuation with text guidance, reference-to-video with up to five real-person inputs, and video editing via text, reference images, or style transfer. ComfyUI added support the same day in version 0.18.5 with workflow templates for all five task types.
HappyHorse-1.0 — Alibaba ATH AI Innovation Unit
Best for: Top-ranked benchmark quality (T2V + I2V); commercial API access with joint audio-video and seven-language native lip-sync
- Max resolution: 1080p
- Audio: Joint audio-video generation in a single forward pass; native synced output
- Lip-sync languages: 7 — English, Mandarin, Cantonese, Japanese, Korean, German, French
- Architecture: 15B-parameter unified 40-layer self-attention Transformer
- Inference speed: ~38 seconds for 1080p on a single NVIDIA H100
- Benchmark position (mid-July 2026): #2 on Artificial Analysis T2V no-audio (Elo 1,288), behind Gemini Omni Flash’s #1 (1,327); HappyHorse-1.1 #3 (Elo 1,273); Seedance 2.0 #4 (1,272). Omni Flash swept #1 on all four AA boards this month, ending HappyHorse-1.0’s no-audio reign. On the with-audio boards HappyHorse-1.1 sits #4 (T2V ~1,150) / HappyHorse-1.0 #5, both behind Omni Flash (#1) and Seedance 2.0 (#2)
- Access: API live via fal.ai ($0.14/sec 720p, $0.28/sec 1080p) and Alibaba Cloud Bailian (enterprise from April 27, 2026); open weights still pending despite ATH’s marketing claim
- API: ✓ via fal.ai (4 endpoints) and Alibaba Cloud Bailian (enterprise tier)
HappyHorse-1.0 debuted anonymously on Artificial Analysis on April 7, 2026 (Weekly Roundup — April 11, 2026), immediately ranked #1 in both text-to-video and image-to-video blind testing, surpassing Seedance 2.0. Alibaba revealed its ATH AI Innovation Unit ownership on April 10. The 15-billion-parameter model uses a unified 40-layer self-attention Transformer that generates audio and video jointly in a single forward pass — no cross-attention modules, no separate audio post-processing.
In April 2026, fal launched HappyHorse-1.0 as official API partner with four endpoints at $0.14 per second for 720p output and $0.28 per second for 1080p — pay-per-second, no minimums. Alibaba Cloud Bailian opened enterprise-grade access the same day.
The open-weights story is messier. ATH’s happyhorse.me/open-source landing page describes HappyHorse-1.0 as “fully open-sourced,” but independent verification finds a public GitHub repo with no model weights, no inference code, and no license file; the Hugging Face profile remains auth-gated. Alibaba has effectively separated commercial API access (live) from open-weight distribution (still unscheduled). Until weights ship, treat HappyHorse-1.0 as a commercial model with an open-source promise — the API is the actual access surface.
HappyHorse-1.1 appeared on the Artificial Analysis leaderboard in late June 2026 — the first model update from ATH’s AI Innovation Unit since the 1.0 debut in April — and has since settled into #2 on both Artificial Analysis with-audio boards: T2V with-audio (Elo 1,151, behind Seedance 2.0’s 1,222) and I2V with-audio (Elo 1,116, behind Seedance 2.0’s 1,194). HappyHorse-1.0 sits third and fifth respectively on those same boards (Elo 1,125 T2V; Elo 1,089 I2V) — Grok Imagine Video 1.5 and Wan 2.7 have both climbed above it on I2V since its debut. On the no-audio T2V board the ranking is unchanged: HappyHorse-1.0 #1 (1,290), HappyHorse-1.1 #2 (1,285). No separate pricing or distinct API endpoint has been announced for 1.1; access is through the same fal.ai and Alibaba Cloud Bailian surfaces. Open-weights delivery for 1.0 remains unshipped — 69 days past the April 27, 2026 commercial API launch and the “fully open-sourced” marketing claim.
MAGI-2 Preview — Sand.ai
Best for: Open-weights unified audio-video generation at the 100B+-parameter tier; the only license in this set with no territory exclusion
- Architecture: 114B-parameter MoE (MagiMoE), ~6B parameters active per token; two-stage pipeline —
magi2_preview(low-resolution denoising) +magi2_refiner(upscale to 1080p) - Max resolution: 1080p (via refiner stage)
- Max duration: 10 seconds — the only duration currently supported
- Audio: Native — generated jointly with video in the same pass, muxed into the output file
- Checkpoint size: ~307GB total — 228GB preview-stage transformer, 56GB Qwen3.5-27B text encoder, 14GB refiner, 5GB Stable Audio Open 1.0 audio VAE, 3GB Wan2.2 video VAE, 2GB turbo VAE decoder
- Hardware: Eight NVIDIA Hopper-class GPUs (~80GB VRAM each) documented for inference; no smaller-footprint path published
- License: Apache 2.0 — unmodified boilerplate plus a copyright notice; no territory carve-out, community-license rider, or acceptable-use policy anywhere in the repository
- Access: Weights-only. The Hugging Face model card confirms it “isn’t deployed by any Inference Provider” — self-host or wait for third-party inference
- Arena: #6 on Artificial Analysis’s image-to-video leaderboard (Elo 1,104)
- Vendor: Sand.ai (Beijing) — founded by Cao Yue, a Swin Transformer co-author and former Microsoft Research Asia researcher; prior release MAGI-1 (24B-parameter autoregressive video model, open-sourced April 2025)
Sand.ai released MAGI-2 Preview’s weights and inference code on August 5, 2026 — a 114B-parameter mixture-of-experts model that activates roughly 6B parameters per token to generate 10-second clips with synchronized audio in a single pipeline. The model card confirms the full spec: two-stage generation (a preview pass followed by a 1080p refiner), a ~307GB checkpoint across six weight sets, and a documented inference requirement of eight Hopper-class GPUs, with no smaller footprint published. It debuted at #6 on Artificial Analysis’s image-to-video leaderboard.
The license is the story. MAGI-2 Preview ships under a standard, unmodified Apache 2.0 grant — no separate community license, no territory clause, no acceptable-use policy layered on top. That’s a direct contrast with MiniMax H3, whose Community License excludes the US, EU, UK, and South Korea from using its open weights at all. Sand.ai’s own about-us page calls MAGI-2 Preview “the world’s first 100-billion-parameter MoE video generation model” released as fully open source — a vendor claim, not independently audited, but the license terms behind it check out unmodified against the repository.
Sand.ai is a Beijing-based lab founded by Cao Yue, a co-author of the Swin Transformer who previously led a vision-model research center at the Beijing Academy of Artificial Intelligence. MAGI-1, the company’s first open-source video model, shipped in April 2025; MAGI-2 Preview is Sand.ai’s first appearance on this page.
Other Notable Open-Source Models
- SkyReels V4 (Skywork AI) — Released April 3, 2026 (Weekly Roundup — April 17, 2026). First open-source model to co-generate video and synchronized audio in a single forward pass. Dual-stream Multimodal Diffusion Transformer (MMDiT) architecture; 1080p at 32 FPS, clips up to 15 seconds. Accepts text, images, video clips, masks, and audio references. Ranked among the top models on Artificial Analysis T2V with audio leaderboard (Elo ~1,135). Free tier: 70 monthly credits on skyreels.dev; open-source weights available for local deployment
- Mochi 1 — High-fidelity short video with strong prompt alignment
- HunyuanVideo / HY-World 2.0 (Tencent) — HunyuanVideo offers solid image-to-video with coherent motion. In April 2026, Tencent’s Hunyuan team released HY-World 2.0 — a multi-modal world model that generates editable 3D scenes (meshes plus Gaussian Splattings) from text prompts or single reference images, with WorldMirror 2.0 inference code and weights open-sourced (github.com/Tencent-Hunyuan/HY-World-2.0). The combination of editable 3D geometry and open weights makes HY-World 2.0 the more pipeline-friendly counterpart to Alibaba’s still-gated Happy Oyster
- Happy Oyster (Alibaba ATH) — Released April 16, 2026 (Weekly Roundup — April 17, 2026). World model that generates interactive, physics-aware 3D environments from text prompts; targets gaming, film, and VR. Directing and Wandering modes are designed for real-time exploration but don’t expose the underlying 3D representation in a standards-friendly way (unlike Tencent’s HY-World 2.0 above). Live demo accessible via Artificial Analysis arena; weights gated
- Helios (Peking University / ByteDance / Canva) — 14B autoregressive diffusion model; 19.5fps real-time generation on a single NVIDIA H100; capable of minute-scale video; Apache 2.0 license. Released March 2026. Notable for real-time throughput on a single accelerator
- NVIDIA Cosmos 3 (NVIDIA) — Released June 1, 2026 at Computex (Weekly Roundup — June 8, 2026). An open-weights omnimodel for physical AI that generates text, images, video, ambient sound, and actions in a unified architecture; shipped as Cosmos 3 Super and Cosmos 3 Nano on Hugging Face under the OpenMDW-1.1 license, corroborated by a 291-author arXiv paper. NVIDIA claimed top open-source rank on Artificial Analysis for text-to-image and image-to-video at launch; as of June 2026 the live board confirms Cosmos3-Super leading open-weight image-to-video at Elo 1,251 (Weekly Roundup — June 15, 2026) — the claim now borne out on the I2V side. Positioned as physical-AI / world-model infrastructure rather than a creator generation endpoint, but the open weights and unified video generation make it a tracked open-source entrant
How to Choose: A Routing Framework
The right model depends on the shot, not the project. Here’s a practical decision framework:
Which AI video model is best for broadcast-ready 4K? Kling 3.0 or Veo 3.1. Kling hits 4K at 60fps with multi-cut storyboards. Veo 3.1 leads on photorealism.
Which AI video model wins on benchmark quality with commercial API access? HappyHorse-1.0 via fal.ai ($0.14/sec 720p, $0.28/sec 1080p) or Alibaba Cloud Bailian — #1 on Artificial Analysis T2V and I2V no-audio leaderboards; joint audio-video; seven-language native lip-sync.
What’s the best free AI video model to start with? Veo 3.1 via Google Vids (10 free clips/month, any Google account).
Which AI video model is free inside an app you already use? Gemini Omni Flash via YouTube Shorts and YouTube Create (available since May 2026, any Google account).
Which AI video model accepts multimodal input (image + audio + video → video)? Gemini Omni.
Which AI video model wins character consistency across shots? Seedance 2.0 Pro via CapCut (US available since April 2026, with real-face restrictions) or Luma Ray 3.2.
Which AI video model is best for stylized and VFX work? Runway Gen-4 Turbo.
Which AI video model propagates a single-frame edit across a 30-second multishot? Runway Aleph 2.0 in Edit Studio.
Which AI video model handles professional production volume at scale? Luma Ray 3.2 (4× faster, 3× cheaper than previous Ray; adds 16-keyframe control and HDR/EXR export).
What’s the best low-cost AI video model for volume work? Pika 2.5.
What’s the cheapest AI video API in 2026? Grok Imagine ($4.20/min generated).
Which AI video model is best for local generation and privacy? LTX-2.5 via ComfyUI or desktop editor.
Which AI video API is best for real-time interactive avatars? Runway Characters (GWM-1).
What’s the best real-time AI video for live agent meetings? PikaStream 1.0 (24fps/480p, ~1.5s latency).
Which AI video model wins multi-shot narrative? Seedance 2.0 Pro via CapCut (US, with restrictions), Luma Ray 3.2, or Kling 3.0 Omni.
Which AI video models work inside Adobe Creative Cloud? Adobe Firefly multi-model hub (Veo 3.1, Kling 3.0/Omni, Runway Gen-4.5, Luma, plus 30+ others).
Which AI video orchestration agent runs Kling + Veo + Seedance + MiniMax from one chat? Pika Agents (April 28, 2026; Slack/Telegram/Discord/X/Notion/Figma, persistent memory) or Higgsfield Supercomputer (mid-May 2026; orchestrates Seedance 2.0, Gemini, GPT-5.5 on web and Telegram). Runway Agent (May 13, 2026) covers the single-vendor end-to-end case. Luma Agents for enterprise campaign production (Publicis, Adidas, Mazda deployments). Adobe Firefly AI Assistant for teams already in Creative Cloud (public beta, CC subscription). See The Agentic Layer for the full comparison table.
Which AI model is best for editable 3D world generation? Tencent HY-World 2.0 (open weights) or Alibaba Happy Oyster (gated early access).
Which AI video model has the largest built-in distribution? Grok Imagine (500M+ X users).
Most professional workflows use 2-3 models per project, routing different shots to different engines based on the specific requirements of each scene.
What’s Coming
- Reactor — real-time world models as infrastructure — Reactor emerged from stealth in May 2026 with $59M (Lightspeed led; Jeffrey Katzenberg’s WndrCo, Amplify Partners, Sky9 Capital, FPV Ventures also in). Co-founders are former Apple Vision Pro technical leads. The pitch: a unified SDK and API that makes real-time world models available to developers “with a few lines of code,” targeting media and entertainment, physical AI, and robotics. This is a different category from the commercial models and agents tracked above — it is infrastructure for interactive AI worlds, not a generation endpoint. No Stack row warranted yet; worth watching as the definition of “AI video” expands toward interactive, real-time, and physics-driven output.
- MCP as a distribution layer for AI video — In May 2026, Runway launched Runway MCP, connecting its model roster (Gen-4.5, Seedance 2.0, GPT Images 2.0, Kling) to Claude, ChatGPT, Cursor, Replit, and any MCP-compatible client. The following day, Pika and Higgsfield both shipped their own MCP skills — Pika’s Founder Starter Kit (four Claude skills: Build-a-Brand, App Screens, Product Sizzle, Founder Video) and Higgsfield Supercomputer as a Claude skill. Three Tier-1 AI video labs shipped MCP integrations inside 24 hours; Higgsfield also launched five Adobe Premiere Pro and After Effects plugins in the same window. MCP is becoming the standard distribution channel for AI video capability into developer and coding-agent workflows — a second distribution layer running alongside the model APIs.
- Seedance 2.5 — three aggregators shipped it the same Friday morning, still no price on any of them — ByteDance previewed Seedance 2.5 at Volcano Engine FORCE (June 23, 2026): single clips up to 30 seconds with no post-stitching, joint audio-video generation, up to 50 multimodal reference inputs, targeted post-generation editing. ByteDance’s own Dreamina app pushed the global consumer launch July 31; the BytePlus ModelArk developer API opened July 16. On August 7, 2026, Runway, Higgsfield, and Pika all went live with Seedance 2.5 within about two hours of each other (06:11–07:59 UTC) — the aggregator layer racing a model none of them controlled, resolving within a single morning. No independent benchmarks exist yet — every spec is ByteDance’s own claim. If durability holds, 30-second native generation from a top-tier T2V model resets the category’s duration ceiling. August 2026 additions: ComfyUI shipped native Seedance 2.5 partner-node support on August 7, putting the model on the self-hosted local-inference layer with a documented per-run cap of 50 reference assets (30 images, 10 videos, 10 audio), second-level shot control via timeline prompting, and native dialogue and lip sync in 10+ languages. On August 14 Higgsfield opened free 1080p generation as a “limited time” early-access tier — the first resolution figure attached to Seedance 2.5 on any surface, since the August 7 go-live announcements specified none. Pricing landed — corrected 2026-09-01. This page said “no published consumer or developer price from any aggregator” through August. That was wrong: ByteDance publishes first-party rates on BytePlus ModelArk for
dreamina-seedance-2-5-260628— $0.514 per 5-second 480p clip ($0.103/sec), $1.156 at 720p ($0.231/sec), $2.843 at 1080p ($0.569/sec), verified against the vendor’s own table. A dated 28% discount applies to 1080p only (480p and 720p explicitly excluded), running 14:00 UTC+8 August 14 through 14:00 UTC+8 September 17, 2026 — which puts the currently-billed 1080p rate near $0.41/sec, not $0.569. The list figure is computed off the pre-discount token rate. Anyone budgeting before mid-September should use the effective rate (Weekly Roundup — August 17, 2026). - Meta Muse Video — MSL enters generative video — Meta Superintelligence Labs previewed Muse Video on July 7, 2026, alongside Muse Image — its first in-house image and video models after years of licensing Midjourney and Black Forest Labs. Built on the Muse Image pretraining base with native audio; Meta claims competitive prompt adherence, visual fidelity, and temporal consistency, and it debuted #3 on Design Arena’s text-to-video board (Elo 1,459, July 5). But it’s a preview, not a product — no published resolution, clip length, or pricing (“coming soon to creators and in Meta AI”), and Meta flags its own gaps in audio-video sync and fast motion. A watch entrant, not a counted Big Ten row until it ships with specs (Weekly Roundup — July 13, 2026)
- PixVerse — funded to scale, world model incoming — the Singapore lab (founded 2023 by ex-ByteDance vision lead Wang Changhu) closed a $439M Series C extension on July 13, 2026, pushing its valuation past $2B with Alibaba among the new backers (and already a deployment partner). PixVerse claims 150M+ users across 177 countries; the round funds its R1 real-time world model and a push into interactive entertainment. PixVerse V6 currently sits mid-table on the Artificial Analysis audio board (Elo ~1,070) — a watch entrant with money and distribution, not a counted Big Ten row until a category-leading model ships (Weekly Roundup — July 20, 2026)
- EU AI Act transparency Code of Practice — initial signatory list published July 31, 2026 — The EU AI Office published its marking-and-labelling Code of Practice on June 10, 2026, ahead of the binding Article 50 disclosure date. Google formally signed July 24; Meta followed July 28, framing the move as keeping disclosure “practical, interoperable and genuinely useful” while warning against “a growing array of different labels and disclosures that end up overwhelming people.” The Commission’s own published list (~190 organizations) confirms it: Google and Meta are the only AI-video-relevant signatories. ByteDance, Kuaishou, Alibaba, TikTok, SpaceXAI, and the newly-launched MiniMax all remain unsigned as enforcement begins — the code stays open for signature, so absence isn’t necessarily a declined invitation (Weekly Roundup — August 3, 2026)
- Dreamina Octo — ByteDance’s “next chapter” beyond Seedance 2.0. Revealed at AI on the Lot (Culver City, May 27, 2026) under the framing “From Generation to Emergence” and “when the prompt isn’t the point.” Early access survey live; product confirmed as “arriving soon,” not yet shipped. If Octo ships as a conversational orchestrator rather than a generation model, it belongs in the agentic section; if it ships as a new Seedance-tier model, it belongs in the Big Ten. Watch @dreamina_ai for the actual launch.
- Gemini Omni Pro shipping window — Omni Flash’s June 30, 2026 public-preview API launch answered the access half of this item; Omni Pro itself remains unshipped with no date attached — “soon” is still the operative word. Omni Flash landed #1 on both Artificial Analysis with-audio boards July 14, 2026 (T2V 1,240 / I2V 1,203), passing Seedance, on top of its July 2 #1 debut on Design Arena’s Video Arena (Elo 1,404) — two boards, two methodologies, both now led by Omni Flash (RCTV flagship analysis →)
- Agentic-orchestration layer — Six conversational agents now sit on top of the model layer: Luma Agents (March 5, 2026), Pika Agents (April 28, 2026), Runway Agent (May 13, 2026), Higgsfield Supercomputer (mid-May 2026), Adobe Firefly AI Assistant (public beta April 27, 2026), and HeyGen HyperFrames (June 26, 2026). See The Agentic Layer for the full comparison
- Grok Imagine Pro (1080p) — Slipped past Musk’s late-April commitment with no new public timeline as of June 2026. Grok Imagine Video 1.5 reached GA on June 16 — audio-native at $0.08/sec across the Imagine API, grok.com, iOS, and Android; July 2026 added single-request 15-second takes (per user reports, no SpaceXAI release note) — but both models still top out at 720p. Until the Pro tier ships, Grok stays out of broadcast and large-format work. Track SpaceXAI release notes for the actual ship date
- HappyHorse-1.0 open-source weights — Now 69 days past the April 27, 2026 commercial API launch on fal, and the weights still haven’t shipped: the public GitHub repo remains empty (no weights, no inference code, no license file) and the Hugging Face profile still reads “coming soon.” Artificial Analysis now lists HappyHorse as an API-only product. ATH’s happyhorse.me/open-source “fully open-sourced” claim hasn’t softened, hasn’t acknowledged the gap, and hasn’t put a date on the artifact. Independent verification by WaveSpeedAI remains the canonical source on the marketing-vs-artifact gap
- TAKE IT DOWN Act enforcement first action — The FTC opened formal enforcement after the May 19, 2026 deadline (FTC press release, business-guidance framework). The agency stood up TakeItDown.ftc.gov as the victim-facing intake surface and on the same day sent a second-wave warning letter to twelve “nudify” tool sites — on top of the fifteen platforms named May 13, 2026. First enforcement target — platform, tool, or generator — sets the operative precedent for the rest of the year. Civil penalties up to $53,088 per violation; 48-hour removal SLA; Section 5 enforcement (Section 230 is no shield)
- AI executive order — signed June 2, 2026, light-touch — Trump signed Promoting Advanced Artificial Intelligence Innovation and Security on June 2, 2026. The version that landed is voluntary, not the FDA-style framework early drafts floated: developers of “covered frontier models” may give federal agencies a 30-day pre-release look, and the text explicitly bars any “mandatory governmental licensing, preclearance, or permitting requirement.” OpenAI, Anthropic, and Google all welcomed it. For AI video the exposure runs through the frontier-model layer, and the DOJ criminal-AI priority stacks onto TAKE IT DOWN. The month-long regulatory overhang resolved in industry’s favor (Weekly Roundup — June 8, 2026)
- Runway Cosmos Coalition — world models as an industry alliance — In June 2026, Runway, NVIDIA, and “leading AI labs” announced the Cosmos Coalition to build and open-source frontier world models; Runway joins as a founding member. As announced it’s a mission statement — no specs, no timeline, no architecture. It landed the same day as NVIDIA’s actual Cosmos 3 release (now in Open-Source above) and Luma’s Open Physical AI Lab, making “world model” three different things in 24 hours (Weekly Roundup — June 8, 2026)
- Agnes-Video — a new API price floor — Singapore’s Sapiens AI entered the Artificial Analysis arena with Agnes-Video-V2.0 at $0.30/min — the cheapest price on the board, roughly a tenth of the next commercial tier. The catch is quality: it debuts near the bottom (Elo 905, #24 on audio T2V). A price-floor entrant, not a quality threat — yet (Weekly Roundup — June 8, 2026)
- Varya — India’s government-backed open-weight entrant — Bangalore’s Avataar.ai launched Varya (June 11, 2026), an open-weight video model distilled from Alibaba’s Wan 2.2 (4 steps vs. 50; 5-second 720p in 45s on an H200), weights on India’s AI Kosh repository under the $1.2B IndiaAI Mission. The hook is price: a vendor-reported ₹0.48/sec (~$0.005), roughly 20× below the global leaders. Treated as a watch entrant, not a tracked open-source row: the 14B parameter figure is Indian-outlet-only (not confirmed by TechCrunch), Artificial Analysis hasn’t benchmarked it, and Avataar’s own model page was unreachable at writing. Real as a price-floor and sovereign-AI signal; specs unverified (Weekly Roundup — June 15, 2026)
- Black Forest Labs ships the video model — FLUX 3, July 23, 2026 — the pivot Martin Scorsese’s June 2 advisory role implied was coming: a multimodal model that “jointly learns from images, videos, and audio within a unified architecture,” generating up to 20 seconds per clip with native synced audio (text-to-video, image-to-video, video-to-video, multi-shot chaining). Early Access only via API/private weights; an open-weight “FLUX 3 Dev” is named with no ship date beyond “the next few weeks and months.” Runway integrated it August 4, Pika followed August 5, and Luma made it three on August 12 — three aggregators in eight days, the same fast-absorption pattern already established for Seedance 2.5, Grok Imagine, and MiniMax H3 (Weekly Roundup — August 17, 2026). FLUX 3 Video also debuted at #2 on Design Arena’s text-to-video board (Elo 1,496, marked preliminary by Design Arena itself), sixteen points behind Gemini Omni Flash. Still primarily an image-model company by revenue and history; still a watch entrant, not a counted Big Ten row. The three-aggregator threshold was reviewed on August 18, 2026 and read as a pattern rather than noise; whether that makes BFL a commercial frontier lab or the infrastructure underneath the platforms absorbing it is the subject of a scheduled analysis, and Big Ten membership stays open until that piece answers it. First capability update since launch: FLUX Upscale (August 20, 2026), a dedicated native-resolution upscaler taking FLUX 3 Video output to 2K/4K, via API and a public demo — shipped after
@bfl_aihad gone four consecutive pulse runs silent. A lab shipping pipeline utilities for its own model is evidence on the very question the scheduled analysis is asking - State deepfake legislation map — Federal preemption of state AI law isn’t happening in 2026. Connecticut HB 5312 (May 2026) establishes a private right of action against creators of non-consensual AI-generated intimate imagery. Vermont’s election-deepfake bill, Iowa’s chatbot-safety law, and Utah’s nine AI bills moved in the same period. Together with Tennessee’s ELVIS Act, California’s AB 2655, and New York’s election-deepfake law, the state map is denser than the federal one. AI video labs allowing image-to-video from real-face inputs need to model state civil exposure alongside federal compliance
- State synthetic-media provenance mandates — A second state-law track, distinct from the NCII-liability bills above: provenance-embedding requirements that reach the generators directly rather than punishing misuse. Connecticut SB 5 passed both chambers in May 2026 and awaits Gov. Lamont’s signature; it requires providers with >1M monthly users to embed C2PA-aligned, tamper-resistant provenance data in generated audio, image, and video (provenance obligation effective Oct 1, 2026; detectability standard Oct 1, 2027). Arizona SB 1786 died — recalled to the House for reconsideration (the Senate returned it May 4, 2026) and never enacted; the azleg.gov Bill Status Inquiry lists it “Engrossed — Dead,” confirming RCTV’s R#12 correction that it died in the legislature and was not vetoed. California’s AI Transparency Act — SB 942 (2023–2024, Chapter 291), as amended by AB 853 — is the operative California statute, in force since August 2, 2026: covered providers with more than 1,000,000 monthly California users must offer a free public AI-detection tool plus manifest and latent disclosures, at $5,000 per violation per day (Business and Professions Code § 22757.4). Two pending bills would amend it, and both carry “California AI Transparency Act” in their titles — a naming collision worth keeping straight: SB 1000 is Enrolled as of August 30, 2026 and on Governor Newsom’s desk — the Assembly concurred in Senate amendments 39–0 on August 27 and the urgency clause was adopted, so if signed it takes effect immediately rather than on the standard January 1 cycle. Not signed, therefore not yet law. Per the Assembly Privacy and Consumer Protection analysis, it would delete the user-count threshold from the definition of “covered provider” (broadening who is regulated), rename the “AI detection tool” a “disclosure verification tool,” drop the requirement that covered providers offer a manifest-disclosure option, and require the latent disclosure to state whether content was AI-generated or AI-modified — directly amending the regime the August 24 roundup covered when an audit found most SB 942-covered companies without a working detector; AB 2713 (“California AI Transparency Act: system provenance data,” amending § 22757.3.1) has cleared the Assembly and is in the Senate. Hawaii HB 2137 was signed — it became Act 247 on July 14, 2026 (Gov. Msg. No. 1349), “Relating to Artificial Intelligence: Realistic Digital Imitations; Protections for Individuals,” after passing both chambers on final reading May 6 and reaching the Governor May 7. Verified against the Hawaii Legislature’s own measure-status record. Unlike removal mandates, provenance embedding is a model-build requirement — it changes what the model ships, not just what a platform takes down (Weekly Roundup — May 25, 2026)
- The training-data cases — live, and deliberately carrying no trial date — Andersen v. Stability AI (N.D. Cal., 3:23-cv-00201, Judge Orrick, filed January 2023) is the generative-media fair-use case whose outcome would govern training-data law for video models as much as image ones, which is why it sits on a page about video. RCTV is not publishing a trial date for it. Trade coverage has circulated one; the court’s own public case page carries no trial-date field, govinfo’s case record carries no scheduling data, and the docket’s own recent-filings list ran through August 31, 2026 on contested discovery — evidence-preservation disputes, deposition fights, sealing motions. That is not the shape of a docket about to open a trial, and we would rather print nothing than a date sourced to repetition. Separately and firmly on the record: Stability AI closed a $76M Series B on August 25, 2026 — $232M total to date — with Electronic Arts and all three major music groups (Sony Music, Universal Music, Warner Music) taking equity, alongside AMD Ventures and Pacific Alliance Ventures. Rightsholders buying into the company they are litigating against is the more interesting fact than any scheduling order
- Hollywood and ByteDance sign — the first AI-safeguards accord between the MPA and a Chinese tech company (August 17, 2026) — ByteDance and the Motion Picture Association signed a global memorandum of understanding on AI intellectual-property protections covering Seedance (video) and Seedream (image) across TikTok, the TikTok USDS joint venture, CapCut and Dreamina, per the MPA’s own announcement. Read what it is: a voluntary truce, not a licensing deal — studios receive no payment for catalogue use, and neither party published technical enforcement details, thresholds, or detection mechanisms alongside it. It resolves the enforcement chain running from the February 12 cease-and-desist over Seedream 5.0 Lite and Seedance 2.0 through the March 9 legal escalation (Weekly Roundup — August 24, 2026)
- New York Synthetic Performer Disclosure Law — live June 9, 2026 — First-in-the-nation requirement that advertisers conspicuously disclose AI-generated “synthetic performers” (fabricated human likenesses) in visual/audiovisual ads reaching New York consumers; civil penalties $1,000 first violation / $5,000 subsequent; audio-only ads and expressive works (film, TV, games) exempt. A distinct mechanism from content-labeling — ad-specific disclosure — adding a jurisdiction to the compliance staircase (Weekly Roundup — July 13, 2026)
- Runway Gen-4.5 — Accessible via Adobe Firefly’s multi-model hub since April 2026; standalone Gen-4.5 launch on runwayml.com still pending. Previewed on NVIDIA Vera Rubin hardware at GTC (March 2026)
- NVIDIA Vera Rubin cloud deployment — AWS, Google Cloud, Microsoft Azure, and OCI all confirmed H2 2026 availability. Vera Rubin delivers 10× lower inference token cost versus Blackwell — the number that will reshape per-second AI video pricing across all major cloud platforms
- DLSS 5 — NVIDIA’s neural rendering technology, launching Fall 2026. Explicitly positioned for filmmaking and VFX beyond gaming; uses generative AI to infuse photoreal lighting and materials anchored to source 3D geometry
- Blackburn draft AI bill — GOP Senate draft (March 19, 2026) declares AI training on copyrighted works not fair use; targets deepfakes and Section 230. Not yet introduced as legislation; path to passage uncertain
- White House AI framework vs. CLEAR Act — White House (March 2026) takes the opposite position from Blackburn: AI training is not infringement; courts should decide. Bipartisan CLEAR Act (Schiff/Curtis) proposes mandatory training data disclosure without resolving fair use. Three irreconcilable positions now active in Washington simultaneously
- Seedance 2.0 copyright litigation — US CapCut access available since April 2026 with real-face and IP restrictions, but the underlying copyright dispute with Disney, Paramount, Warner Bros., and Netflix remains unresolved. The restrictions are a negotiating posture, not a settlement
- OpenAI robotics / world simulation — OpenAI redirected Sora’s compute toward “world simulation for robotics” after shutting the product down. The consumer app went dark on April 26, 2026 as scheduled; the Sora API remains accessible until September 24, 2026
- Adobe Firefly multi-model expansion — Firefly’s video hub now hosts 30+ third-party AI models including Kling 3.0/Omni, Veo 3.1, Runway Gen-4.5, ElevenLabs Multilingual v2, Luma AI, Black Forest Labs, and Topaz Labs. Firefly AI Assistant orchestrates multi-step workflows across Photoshop, Premiere, Lightroom, Express, and Illustrator
- Tencent vs. Alibaba 3D world model race — Two of China’s largest AI labs shipped 3D world models on the same day, April 16, 2026 (Alibaba’s Happy Oyster, gated; Tencent’s HY-World 2.0, open weights). Western labs have nothing comparable in production; the 6-to-12 month head start is real if world simulation matters as much as OpenAI’s Sora-shutdown framing implied
- Google Vids / Workspace expansion — YouTube export is live; paid creative tiers (Pro/Ultra) include Lyria 3 music generation and AI avatars. Further Workspace AI integration expected throughout 2026
- EU AI Act Article 50 — now in effect. The disclosure and deepfake-labeling obligations took effect August 2, 2026 as scheduled. Article 50 carries no penalty schedule of its own — violations fall under the AI Act’s general enforcement framework (Article 99), capped at €15 million or 3% of global annual turnover. A two-step deadline, not one cliff: the harder machine-readable watermarking requirement (Art. 50(2)) still carries a grace period to December 2, 2026 for generative systems already on the market before August 2 (Weekly Roundup — August 3, 2026)
- Unlimited-length AI video — EPFL’s drift elimination breakthrough (presenting at ICLR 2026) could remove the duration ceiling entirely
- SpaceXAI targeting 30-minute video — Announced goal for late 2026, with full-length films targeted for 2027
This page is maintained by RCTV as a public reference. For weekly updates on model releases and industry shifts, see our Weekly Roundup.
Have a correction or update? Contact us at rctv.oxncw@simplelogin.com
Changelog
Showing the four most recent updates. Full changelog archive →
September 1, 2026
- Catch-up refresh — this page missed the R#18 and R#19 paired updates. The standing cadence is that the Stack refreshes alongside every published roundup. It did not happen on August 23 or August 30, and nothing caught it: the changelog invariant check asserts structure (four inline blocks, each mirrored in the archive) and holds perfectly on a page nobody touched. Two roundups shipped against a
lastmodof August 18. What that cost readers: this page claimed Gemini Omni Flash had swept all four Artificial Analysis boards and that MiniMax H3 led video-editing. Both were true in July and neither was true by August 19. A freshness check now runs inrctv-gatesand in the daily task sweep, comparinglastmodagainst the newest published roundup. - Wan 3.0 — Alibaba (NEW ENTRY, commercial) and the section is now the Big Ten: Alibaba opened an invite-gated commercial beta August 6, 2026; it first charted August 19 and went straight to #1 on Artificial Analysis text-to-video (Elo 1,241) and #1 video-editing (Elo 1,189), displacing Gemini Omni Flash and MiniMax H3 respectively, and has held both. 30-second single-pass generation, document-to-video input (doc/xls/ppt/pdf/md up to 100MB), 480p/720p/1080p, no 4K. Closed — no published weights, no GitHub listing, distinct from the open-weights Wan 2.7 line, which now carries a disambiguation note. Promoted on the criterion this page already states for watch entrants (“not a counted row until a category-leading model ships”); the section header moves from Big Nine to Big Ten and eight live references follow it. Historical changelog entries keep “Big Nine” as written, per the audit-trail contract. Distribution followed the established pattern: eleven surfaces in eight days (Runway Aug 24, Pika and ComfyUI within two days, then ArtArch, Media.io, OiiOii AI, Picsart, PowerDirector, A2E, PixelDojo, DeepInfra).
- MiniMax H3 Max — post-trained by fal (NEW ENTRY, commercial): the first third-party fine-tune to top a leaderboard this page tracks — #1 image-to-video (Elo 1,202), #3 text-to-video (1,235), ahead of the official MiniMax H3 endpoint at #4. A 5-second 720p clip in under 3 seconds: generation faster than playback. The three speed multipliers are three different comparisons and are labelled as such (35× vs the official H3 endpoint per fal; 50× per Design Arena independently; 15× against a quality-matched group) — never blended. No per-second rate is published on fal’s own pricing surface, so none is carried. fal’s board label reads “#1 image-to-video with audio”; Artificial Analysis retired its with-audio boards on August 11, so the live #1 is on the plain board — noted in the entry. fal.live (August 31) and its “H3 Max Director” variant are recorded including the pause roughly three hours after launch.
- Gemini Omni Flash — the July sweep is over, and Omni 1.1 Flash added: now #2 T2V (1,237), #4 I2V (1,179), #3 video-editing (1,124) as of August 28. Displaced rather than degraded — the Elo barely moved; two newer models scored higher. Omni 1.1 Flash (August 27) adds scene extension, first/last-frame specification, up to three video references, 4K upscale, and a 360p fast mode. The distribution is split and the entry says so: Google frames the full set as “available via APIs” for developers, while the Gemini app account scopes the consumer rollout to scene extension only. Artificial Analysis has not scored Omni 1.1 separately — the board entry carries none of AA’s version markers — so the figures above describe Omni Flash, not the update.
- Higgsfield — the round closed, above the reported range: $400M Series B at $5.4B on August 17, DST Global leading with Goldman Sachs Alternatives, Valor Capital and Tribe Capital. Disclosed: $700M annualized revenue, 30M users across 200 countries, 390 Fortune 500 clients. This page carried the July report of “$300–500M at a $5B valuation” on a ~$500M run rate; the closed terms came in above the reported range on both, and the entry now says so rather than silently replacing the earlier figure.
- Regulatory — California SB 1000 is Enrolled, not law: the Assembly concurred in Senate amendments 39–0 on August 27 and the bill was Enrolled August 30, 2026, now on Governor Newsom’s desk with an urgency clause, meaning immediate effect if signed. Verified against the leginfo primary record. What it would change to SB 942 is now stated (deletes the user-count threshold, renames the detection tool a “disclosure verification tool,” drops the manifest-disclosure option requirement, requires the latent disclosure to distinguish AI-generated from AI-modified). Still unsigned; the operative statute remains SB 942. Given the ten-week SB 1000 misattribution corrected on August 11, the status wording here is deliberately literal.
- Regulatory — MPA and ByteDance sign (NEW What’s Coming item): the first AI-safeguards accord between Hollywood and a Chinese tech company, August 17, covering Seedance and Seedream across TikTok, the USDS joint venture, CapCut and Dreamina. Carried with the two qualifications that matter: it is a voluntary truce, not a licensing deal — studios receive no payment — and neither party published enforcement thresholds or detection mechanisms.
- Routine freshness across five entries: Runway Ruby (August 21) — SDR-to-16-bit-HDR conversion up to 30s, ProRes/EXR, Max and Enterprise only, a colour-pipeline tool that moves no benchmark. HeyGen LiveAvatar (August 25) — concurrency cap removed entirely, “1 avatar or 10,000,” down to $0.01/min at scale. Seedance 2.5 — US rollout completed August 21, closing the regional sequence begun August 14; five platforms on 1080p within four days. Seedance 2.0 Pro — benchmark note rewritten to the live boards (#5 T2V, #2 I2V, #6 video-editing); the mid-July with-audio figures it carried no longer have a live board behind them. Black Forest Labs — FLUX Upscale (August 20), the first capability update since the July 23 launch, folded into the What’s Coming item as evidence for the scheduled analysis rather than as a new story.
- Agentic layer — two additions to What to watch: Comfy MCP shipped local and open-source August 18, giving the self-hosted layer the same agent entry point the hosted platforms have — the first time this category extends below the API line. Tencent Hunyuan previewed HyCreator (August 20–21), claiming 10-minute films with “zero human intervention” plus a real-time interactive-editing mode, application-gated; carried as a claim, not a capability — no independent hands-on exists.
- First Cowork exit test ever run against this page — three live errors found, all confirmed against RCTV’s own records rather than adopted on the comparison’s word. (1) MiniMax H3 Max resolution: this refresh first wrote 720p; R#19 had already verified 768p (1344×768) against fal’s model page seven times over, with fal’s own X post carrying the conflicting 720p. Corrected to 768p across four surfaces. (Later the same day: the parenthetical “1344×768” carried here and in the entry was itself unsourced — an inference, not a figure any surface states — and is removed. fal’s API schema offers 480P and 768P as labels only.) (2) LTX-2.5 licence: the page called it Apache 2.0 in three places while the entry’s own licence bullet described a $10M-revenue threshold — a condition Apache 2.0 does not contain. Both cannot be true; the Apache label is removed, the revenue threshold retained, and the precise licence name marked pending primary verification. Same failure class as the MiniMax territory-exclusion catch: the release event was tracked, the licence text was not read. (3) Seedance 2.5 pricing: “no published consumer or developer price from any aggregator” is no longer safe — a first-party BytePlus rate card is reported to exist. Flagged, not adopted, pending verification against ByteDance’s own surface.
- Wan 3.0 — the vendor-docs label conflict logged: trade coverage August 24–25 described an official release; Alibaba’s own Model Studio API reference still read “Currently in preview.” Tracked as preview until the vendor docs change. This is RCTV’s own R#19 discovery gap — no monitoring lane covers vendor API documentation — recorded on the page rather than left in the ledger.
- Queued for primary verification, deliberately not carried: Gemini Omni 1.1’s chained-extension ceiling and whether its 1080p/4K are upscaled rather than natively rendered; the
gemini-omni-flash-previewdeprecation date; Wan 3.0’s per-second rate card and API-documented specs; H3 Max’s per-second rate; LTX-2.5’s licence text and arena scores against its 67% launch claim; Seedance 2.5’s first-party rates; FLUX Upscale pricing. Each would improve this page and none is verified to primary yet. - WR verification pass returned the same day — seven page-facing facts resolved to primary, three of them corrections to what this page carried hours earlier. Nothing below was adopted from the comparison document; every figure is from the vendor’s own surface, fetched 2026-09-01.
- Gemini Omni’s 1080p and 4K are UPSCALED, not natively rendered. Google’s own resolution table uses that word: “1080p output (upscaled),” “4K output (upscaled).” 720p is the only resolution Omni renders natively, which is the distinction that matters against Veo 3.1 and Kling 3.0. Also pinned: GA on 2026-08-27 as
gemini-omni-1.1-flash,gemini-omni-flash-previewdeprecated 2026-09-30, scene extension in 10-second increments to a 40-second cumulative ceiling (four chained requests, not one pass), and $17.50/1M video output tokens ≈ $0.10/sec. - MiniMax H3 Max pricing is publishable after all, and our discount figure was wrong. fal’s model page rendered cleanly today when it would not for R#19 three days ago: $0.08/sec at 768P list, $0.05/sec at 480P, with a 75% promotional discount ending September 7 — not the “half price for 14 days” this page and the roundup both carried. R#19 declined to print an unverified per-second figure that has now proved correct; that is the discipline working, not a miss.
- CORRECTION to this page’s own entry, made hours after writing it: the “1344×768” pixel figure was unsourced. No surface states it; fal’s API schema offers 480P and 768P as labels only. It was an inference that read like a specification. Removed here and in the entry.
- LTX-2.5’s licence, verified against the licence text itself (LTX-2.x Community License Agreement, August 11, 2026): $10M revenue threshold aggregated across affiliates; derivatives defined to include fine-tunes, distillation and training on LTX-2.x-generated synthetic data; a competing-model training ban on commercial use; military/weapons/nuclear ban; two separate clauses barring removal of watermarking or provenance; and a stated intent to qualify under EU AI Act Article 53(2), carved out from 53(1)(c) and (d). Added alongside the measured board position — both tiers at Elo 1,060, ranks 21–23 of 33 — against the vendor-commissioned 67% win-rate launch claim.
- Seedance 2.5 first-party pricing exists and this page was wrong to say otherwise. BytePlus ModelArk publishes $0.103/$0.231/$0.569 per second for 480p/720p/1080p, with a dated 28% discount on 1080p only, August 14 through September 17, putting the currently-billed rate near $0.41/sec.
- Wan 3.0 list pricing confirmed at $0.05/$0.10/$0.20 per second with a live 30% discount the page now names, six endpoint regions rather than five, and the full API-documented spec set. Its “Currently in preview” label was re-checked against a page Alibaba itself stamped 2026-09-01 and still reads preview — trade coverage describing an official release remains uncorroborated by the vendor.
- Gemini Omni’s 1080p and 4K are UPSCALED, not natively rendered. Google’s own resolution table uses that word: “1080p output (upscaled),” “4K output (upscaled).” 720p is the only resolution Omni renders natively, which is the distinction that matters against Veo 3.1 and Kling 3.0. Also pinned: GA on 2026-08-27 as
- Leaderboard structure clarified, and a standing ambiguity closed. Artificial Analysis runs three scored leaderboard pages, each carrying a With Audio / No Audio toggle that swaps the whole ranking on the same URL. RCTV read that as “three boards”; the comparison document read it as separate audio and no-audio boards. Both were half-right, and the disagreement was really about whether different Elo per audio state implies a different board. It does not. The Quick Reference now states the structure and says which set this page quotes.
- Training-data litigation added to What’s Coming — with no trial date, on purpose. Carlos ruled generative-media fair-use cases on-beat for the regulatory lane at the 2026-09-01 meeting, reversing a pulse exclusion. Andersen v. Stability AI is now carried. The circulating September trial date is not published: the court’s own case page has no trial-date field, govinfo’s record carries no scheduling data, and the docket’s recent filings ran through August 31 on contested discovery — evidence-preservation disputes, deposition fights, sealing motions — which is not the shape of a docket about to open. CourtListener/RECAP, which would carry the scheduling order itself, is bot-blocked even under headed rendering and is now recorded as a known-blocked surface. Verified and carried instead: Stability AI’s $76M Series B (August 25, 2026), with Electronic Arts and all three major music groups taking equity in the company they are litigating against.
- Changelog trim: August 9, 2026 block removed from the main page inline (retained in the archive). Main page inline now shows exactly four blocks: September 1 / August 18 / August 16 / August 11.
- Considered and excluded: Wan 3.0’s wider aggregator sweep beyond the eleven named — the pattern is recorded once in the entry rather than enumerated per platform. Pika Audio Models (Soundtrack, SFX, Speech, Music) — an audio model line, excluded per this page’s AI-video-first scope, same treatment as MiniMax Music 3. Higgsfield KÖK BÖRÜ at SIGGRAPH and the Global Film Festival thread — business and culture rather than spec or capability, the treatment already applied to Hell Grind and The Cully Hill Boys. “Cinema Studio 4” (Higgsfield) — still flagged, still not added: a second credit on an August 17 short strengthens the case it is a real named tool, but what changed from prior versions remains unconfirmed. EU Article 50(2) watermarking primary pin — still outstanding; the lane entry cites a page that carries no December date, flagged in place and not silently repaired here.
August 18, 2026
- CORRECTION — Grok Imagine Video 1.5’s 1080p ship date was July 31, not August 1: This page stated August 1, 2026 in five places (Quick Verdict, Quick Reference row, spec card, What’s Next, and the closing prose). The correct date is July 31, 2026, per SpaceXAI’s own release page, whose published metadata reads July 31 in UTC. The error came from anchoring the date to a same-week amplification post on the @imagine sub-account (August 1, 20:32 UTC) rather than to the vendor’s own release page or its primary @grok announcement (August 1, 00:46 UTC — July 31 in both US Eastern and the company’s home Pacific time). All five occurrences corrected and the closing prose re-anchored to the release page. The substance is unchanged: 1080p is generally available, the feature set (native 1080p, text-to-video, voice references) is as reported, and the $0.08/sec rate carries no resolution tier. Earlier changelog entries are left exactly as written per this page’s audit-trail contract; this entry is the correction of record. The affected roundup carries its own correction note — see Weekly Roundup — August 10, 2026.
- MAGI-2 Preview — Sand.ai (NEW ENTRY, Open-Source & Local Generation): Verified and added a full entry for the August 5 open-weights release — 114B-parameter MagiMoE, ~6B active/token, 10-second clips with native audio at up to 1080p via a two-stage preview+refiner pipeline; ~307GB checkpoint; eight Hopper-class GPUs documented for inference; weights-only, no hosted API. License is unmodified Apache 2.0 with no territory carve-out — the direct contrast to MiniMax H3’s US/EU/UK/South-Korea-excluded Community License, checked three independent ways (repository LICENSE, repo root listing, model-card metadata). Debuted #6 on Artificial Analysis’s image-to-video leaderboard (Elo 1,104, fetched live 2026-08-18). Vendor: Sand.ai (Beijing), founded by Cao Yue (Swin Transformer co-author); MAGI-1 was its prior release (April 2025), and its bare unattributed bullet in Other Notable Open-Source Models is superseded by this entry and removed. Funding figures circulating in secondary coverage are not carried — no primary confirmation located. Closes the item queued in the August 16 “Considered and excluded” note.
- Black Forest Labs — three-aggregator threshold reviewed, watch-entrant status held: The August 18, 2026 editorial review read three aggregator integrations in eight days (Runway August 4, Pika August 5, Luma August 12) plus a #2 Design Arena text-to-video debut as a pattern rather than distribution noise. What that pattern means — a commercial frontier lab, or the infrastructure underneath the platforms absorbing it — is the subject of a scheduled analysis. Big Ten membership stays open until that piece answers it; the What’s Coming entry now says so explicitly instead of carrying the question as an unresolved review note.
- Sourcing rule applied going forward: a vendor ship date anchors to the vendor’s own dated release page where one exists, or to the primary announcing account — never to a downstream amplification post whose timestamp happens to be what got read.
- Changelog trim: August 2, 2026 block removed from the main page inline (already in the archive). Main page inline now shows exactly four blocks: August 18 / August 16 / August 11 / August 9.
August 16, 2026
- Weekly Roundup R#17 paired refresh:
lastmod→ 2026-08-16 (lead-by-one before the August 17 Roundup);params.tomlogImage→roundup-2026-08-17.png; Mode B infographic regenerated. Ships Sunday per the standard Stack-paired-with-Roundup cadence. - LTX — vendor identity corrected, entry moved to LTX-2.5: LTX has been an independent company since July 8, 2026, spun out of Lightricks (which kept its consumer apps, e.g. Facetune); CEO Zeev Farbman came with the video line. The entry header, the Quick Reference row, both Quick Verdict references, and the FAQ all read “LTX-2.3 (Lightricks)” and are corrected. LTX-2.5 shipped August 11 — Diffusion Fidelity Rendering, native multishot generation holding continuity across cuts, a custom Gemma 4 text-encoder backbone, HDR ACES support, open weights on Hugging Face, Day-0 ComfyUI support, and a same-window Runway integration. The 22M+ cumulative Hugging Face downloads figure since base LTX-2 in January is LTX’s own claim and is labeled as such. The spec bullets deliberately remain the verified LTX-2.3 card: the 2.5 launch post publishes no resolution, duration, or frame-rate numbers, so the card carries an explicit pending-verification note rather than letting 2.3’s figures inherit a 2.5 header.
- HeyGen HyperFrames — MCP server inside Claude Design (August 7): animate a design, hit share, and a HyperFrames agent scores it, adds sound effects, tightens the motion, and returns a finished MP4 — no human in the step between design and shippable video. Added to the HyperFrames entry as the agentic-assembly pattern reaching a design surface rather than a video platform.
- Black Forest Labs — third aggregator, plus a leaderboard debut: Luma integrated FLUX 3 Video on August 12, making three aggregators in eight days after Runway (Aug 4) and Pika (Aug 5). FLUX 3 Video also debuted #2 on Design Arena’s text-to-video board at Elo 1,496 — marked preliminary by Design Arena itself, a qualifier carried into the text. What’s Coming item updated; watch-entrant status explicitly held, with the three-aggregator threshold named as a live editorial question rather than silently treated as settled.
- Seedance 2.5 — ComfyUI partner nodes, and the first resolution figure: ComfyUI shipped native Seedance 2.5 partner-node support August 7 (50 reference assets per run — 30 images, 10 videos, 10 audio; second-level shot control; dialogue and lip sync in 10+ languages), putting the model on the self-hosted layer. On August 14 Higgsfield opened free 1080p as a “limited time” early-access tier — the first resolution figure attached to Seedance 2.5 on any surface, since the August 7 go-live announcements specified none. Still no published consumer or developer price from any aggregator.
- Changelog trim: July 26, 2026 block removed from the main page inline (already in the archive). Main page inline now shows exactly four blocks: August 16 / August 11 / August 9 / August 2.
- Considered and excluded: Sand.ai’s MAGI-2 Preview — a 114B mixture-of-experts open-weights model under Apache 2.0 with no territory carve-outs (released August 5, ~307GB checkpoint, eight NVIDIA Hopper GPUs documented, Artificial Analysis image-to-video debut at #6), surfaced by the R#17 Cowork exit test. Real, on-beat, and a genuine gap against MiniMax H3’s excluded-territory license — but out of R#17’s coverage window and not yet re-verified to primary by the Writer-Researcher. Queued in
next-agenda.mdfor a verified Stack row rather than added on a comparison document’s word. “Cinema Studio 4” (Higgsfield) — a production-tool version credited on the open-sourced ONEIRIC short and never logged here; flagged, not added, pending confirmation of what actually changed. ONEIRIC itself — a promotional open-sourcing, business/culture rather than spec or capability, same treatment already applied to Hell Grind and The Cully Hill Boys.
August 11, 2026
- CORRECTION — California AI Transparency Act misidentified as SB 1000: The “State synthetic-media provenance mandates” item and the July 26 regulatory-timing note both attributed the August 2, 2026 effective date to California SB 1000. That was wrong. The operative statute is the California AI Transparency Act, SB 942 (2023–2024, Chapter 291), as amended by AB 853, which moved the operative date from January 1 to August 2. SB 1000 (2025–2026) is a separate, still-pending bill sharing the same popular name — passed the Senate 33–1 on May 19 with an urgency clause, still on the Assembly Third Reading File, never enrolled or signed. The provenance-mandates item now states the operative statute, its actual obligation and penalty (§ 22757.4), and lists SB 1000 and AB 2713 correctly as pending amending bills. AB 2713 was also clarified: its title is “California AI Transparency Act: system provenance data” and it amends § 22757.3.1 — it is a bill amending the Act, not the Act itself, which the prior wording blurred. All statuses verified against leginfo.ca.gov primary records. Affected roundups (June 15 through August 3) corrected separately; see the correction note on Weekly Roundup — August 3, 2026.
- Changelog trim: July 19, 2026 block removed from the main page inline (already in the archive). Maintains the four-most-recent inline window.