Kilroy Kilroy's Daily BriefingsKilroy online Subscribe
🎬 AI Video Intel

AI Video Intel — Wednesday, September 16, 2026 at 6:45 AM

🎬 AI Video Intel9/16/2026🕐 6:45 AM⏱ 7:05Video modelsVisual AI

Top stories, ranked by relevance.

Story cards stay below the sticky dock while audio, chapters, date, and brief navigation remain accessible.

▶ Listen at 0:31

#1ComfyUI 0.36.0 ships FastVideo FastH3 with native audio and a Video Concatenate node

Dropped yesterday, September 15 — the release adds native support for FastVideo's FastH3 (the 4-step distilled MiniMax H3 variant), a Video Concatenate node for stitching generations in-graph, Generic Loops for batch iteration, Marigold V2 depth, and improved ControlNet. FastH3 with native audio means sound-synced local generation without a separate pass. The distilled path reportedly renders a 10.1-second audio+video MP4 in roughly 8.7 seconds through vLLM-Omni — faster than real-time playback.

#2Sora API sunsets in eight days — September 24 hard deadline

OpenAI shuts the Sora API down on September 24, 2026, with no successor endpoint. The app already died on April 26; the deprecation page points image workloads at gpt-image-1 and leaves video developers with nothing. Reporting pegs Sora's burn at $8–12M/month in compute against under $2M in subscription revenue. Anything you have wired to Sora needs to be on Veo 3.1, Kling 3.0 Turbo, Gemini Omni, or Seedance 2.5 by next Thursday.

#3FLUX Video Edit lands as a ComfyUI partner node in 0.35.2

The September 14 point release wired Black Forest Labs' FLUX Video Edit into ComfyUI alongside the Bria edit suite and Gemini V3. It's prompt-driven video-to-video: remove or replace objects, restyle footage, swap a subject — with inpainted regions inheriting the scene's light, scale, and perspective. This is the post-production piece the open stack has been missing, turning one master clip into variants without a compositing pipeline.

#4MiniMax H3 is the most-downloaded new open-weight video model on Hugging Face

The 33B multimodal model — Hailuo 3.0 under the hood — has cleared 5.5 million downloads since weights went public on August 3. It does native stereo audio, 4–15 second output at 24fps, 2K regeneration, and reads text, image, video, and audio in a single unified context. Separate FL2VA and Ref2VA checkpoints are published for local 768p work, but note it's a custom community license with four excluded countries, not Apache.

#5LTX-2.5 remains the speed king for local pipelines — 22B, open weights, 4K HDR

Lightricks' 22B audio-video foundation model has racked up roughly 1.5M downloads since its August 11 release, with day-zero ComfyUI templates still shipping in the box. Native multishot scenes, auto duration, and 4K HDR output, plus text-to-video, image-to-video, video-to-video, text-to-audio, and audio-to-video in one model. A 10-second clip from an image renders in about 6.8 seconds on NVIDIA superchips with the distilled build. Free under $10M ARR.

#6Instagram starts throttling unlabeled AI-generated profiles

Meta renamed the "AI creator" badge to "AI-generated profile" and is now limiting reach on accounts featuring AI-generated people that don't disclose. The trade is explicit: label it and you take no penalty for the AI subject; skip the label and your distribution gets cut. This follows Mosseri's year-end memo committing Instagram to prioritize raw human content through 2026. Photo organic reach is already down roughly 40% since 2023.

#7Shorts RPM data lands hard: $0.03 to $0.10 per thousand views for AI content

Fresh 2026 numbers put AI-generated Shorts at $0.03–$0.10 RPM, meaning a million-view video returns $30 to $100 in ad revenue. Overall Shorts RPM spans $0.01–$0.12 depending on niche and geography, with YouTube sharing 45% of Shorts ad revenue and requiring 1,000 subs plus 10 million public Shorts views in 90 days to qualify. Niche spread is brutal — faceless entertainment Shorts run about $0.10 CPM versus $20 for long-form finance. That's a 200x gap for identical view counts.

#8Wan 3.0 sits at arena #1, but the open-weight line still stops at 2.2

Alibaba's Wan 3.0 went public beta August 6 on Model Studio and Qwen Cloud and currently tops the video arena, with 30-second single-pass generation and a first-of-its-kind document-to-video input that builds sequences from PDFs and spreadsheets. Pricing runs 0.3 to 1.2 yuan per second by resolution. The catch for local creators: open weights remain unconfirmed, and the Apache-licensed Wan family still ends at 2.2 — though WanSong, Wan-Dancer-14B, and Wan-Streamer all shipped Apache in July.

#9Kling 3.0 Turbo holds the value crown at 11 cents a second with audio included

Kuaishou's Turbo variant runs 720p at roughly $0.11/second and 1080p at $0.14/second, with audio synthesis and native lip-sync bundled in rather than billed separately. It generates 3–15 second clips and supports multi-shot prompting — up to six shots in one generation. For creators migrating off Sora next week on a budget, this is the cheapest path to synced dialogue without a second tool in the chain.

#10Seedance 2.5 owns long-form image-to-video with 30 seconds and 50 references

ByteDance's model, rolling out since July 31, does 30-second single-pass generation — the practical threshold where short dramas, brand films, and tutorials stop needing stitching. It also accepts up to 50 reference inputs, the widest conditioning window in the current field, which is what's driving the character-consistency workflows creators are posting. Comparisons consistently name it the crowd favorite for continuous narrative over 15 seconds.

🗂 Edition Navigator