Kilroy Kilroy's Daily BriefingsKilroy online Subscribe
🎬 AI Video Intel

AI Video Intel — Saturday, October 3, 2026 at 6:45 AM

🎬 AI Video Intel10/3/2026🕐 6:45 AM⏱ 7:11Video modelsVisual AI

Top stories, ranked by relevance.

Story cards stay below the sticky dock while audio, chapters, date, and brief navigation remain accessible.

▶ Listen at 0:25

#1Comfy Agent goes live — the canvas builds itself now

Comfy Org shipped Comfy Agent on October 1: an agent that plans, builds, runs, and debugs ComfyUI workflows directly on the canvas, dropping nodes and wiring them for you. It's running on frontier Anthropic models, supports reusable saved "skills," and handles up to 5 parallel agent chats each with full history. Live on Comfy Cloud in beta today with free starter chat credits; Comfy Desktop with local GPU support is promised "in a few weeks."

#2DMAD cuts MiniMax H3 from 50 steps to 4

ByteDance and Texas A&M released DMAD yesterday — rank-128 LoRAs that distill the 33B MiniMax H3 audio-video teacher into a 4-step generator at 1344x768, 124 frames (about 5.2 seconds at 24fps) with native 32kHz stereo. Kijai already published a rank-reduced ComfyUI conversion, so it loads through the standard H3 LoRA path with no custom node. Caveat: four steps cuts sampling cost, not memory — the full H3 stack is still roughly 170GB with the 32B text encoder.

#3NVIDIA's LongLive-Plug: distill once, reuse everywhere

NVIDIA's SANA team dropped LongLive-Plug, a distillation framework that learns few-step, CFG-folding, and long-context capability as separate reusable LoRAs instead of re-distilling every downstream model. It delivers 5x to 12.5x step reduction across three backbone families — MiniMax H3, Wan2.1-T2V-14B, and Wan2.2-TI2V-5B — and the adapters survive fine-tunes and control modules. Recommended stacking ratio is few-step to CFG at 1 to 0.5.

#4LTX-2.5 Native Resolution brings seamless 4K and 8K video edits

Lightricks shipped Native Resolution for LTX-2.5 — a "Tiled Fusion" approach that holds one latent canvas and noise field while stepping overlapping tiles at every denoise step, then Gaussian-blends them back. No seams, and peak VRAM tracks tile size rather than full canvas. Ships with two workflows plus two new IC-LoRAs: Refine Details (1024x576 tile) and Restore (960x544, for archive cleanup). Keep overlap at 0.5 or higher, use discrete samplers, and disable tiled encode on the guide node.

#5ComfyUI v0.38.0 adds Ming-Image, Wan ID-V2V, and a leaner SeedVR2

Core now natively supports Ming-Image 0.1 Design with a Text Encode Ming Image Edit node taking up to eight reference images, built on a 30-layer DiT with the 256-expert Ling-Mini-2.0 encoder. Wan ID-V2V lands for identity-preserving video restyling on a Wan 2.1 i2v base with VACE control plus an anti-drift ref_pad_image input. SeedVR2's VAE was rewritten for frame-by-frame processing with int8-quantized causal conv caches in pinned host memory — a big VRAM win on long upscales. Also new: w6a8 quantization and LogC3 / ACEScct HDR color spaces.

#6MiniMax H3 X2 Detail VAE: 2x decode plus reference sharpening in one file

Community dev speach1sdef178 released a 5.2GB checkpoint that does double duty — a 2X video VAE decoding H3 latents at twice spatial resolution via a packed 12-channel output plus PixelShuffle, and an optional node that rebuilds fine structure in reference images before generation. Drop it in models/vae/, swap the normal VAE Decode for the fast decode node from ComfyUI-MiniMaxH3_LatentUpscaler, and start at detail_strength 1.0. A tested 2-in-1 workflow ships with it.

#7Qwen-Image 2.1 Consistency LoRA kills edit drift

Creator ausboss trained a fix for the single most annoying Qwen-Image 2.1 edit failure — results coming back a few percent taller, nudged sideways, or repainted outside the mask. Rank 32, alpha 32, 950 edit pairs from 257 images, 3,000 steps on one H100 at about 1.9 seconds per step via ostris/ai-toolkit. Grab the step-1500 checkpoint (152MB) to preserve Qwen's native look; run LoraLoaderModelOnly at strength 1.0, resolution 0, 25 steps, CFG 1, euler.

#8H3 Prompt Builder 1.8.0 adds masked clip editing with SAM 3.1

Version 1.8.0 of the MiniMax H3 Prompt Builder brings masked clip editing driven by SAM 3.1 segmentation, so you can target prompt changes to a specific region of a clip rather than re-rolling the whole generation. Same-day, ComfyUI-RMBG hit v3.2.0 with a SAM 3.1 Multiplex node, Apple Silicon support, and startup node scanning cut from 15 seconds to under 0.1 — SAM 3.1 is quietly becoming the segmentation backbone for the whole video-edit stack.

#9Transformation LoRA turns H3 into a morph engine

Community creator Ashmotv's Transformation LoRA for MiniMax H3 deconstructs the first frame element-by-element and reassembles it into the last — subjects dissolving into each other, materials changing mid-travel. Trigger word is Tr@nsf0rmation_style with the connector Tr@nsf0rmation into, and you prompt element mappings explicitly. Two checkpoints at 36MB: st2500 fully converged (strength 1.0 to 1.2) and st1600 softer (1.0 to 1.4). ResMultistep sampler, 12 to 30 steps, 852x480 or 1024x576.

#10The distribution reality check: LTX-2.5 leads downloads while TikTok down-ranks AI

Hugging Face trending as of October 2 puts Lightricks' LTX-2.5 image-to-video at the top of video gen with 5.8k likes and 1.5M downloads, with Qwen-Image 2.1 and its LoRA ecosystem dominating the creative category. Meanwhile the demand side is tightening: TikTok's March 2026 algorithm change actively down-ranks AI-generated content, catching an estimated 35 to 45 percent automatically via invisible watermarks and C2PA Content Credentials from 47 platforms — and Kapwing's June study found 59 percent of videos served to new TikTok accounts qualify as AI slop, three times YouTube Shorts' rate.

🗂 Edition Navigator