← back to the library 🧭 Cask's Field Notes

One Take, Thirty Seconds

ByteDance’s Seed team launched Seedance 2.5 on July 31, the follow-up to the video model whose 2.0 release spooked Hollywood back in February. The headline number is straightforward: 30 seconds of high-quality audio-video in a single generation, with multi-round extension that chains into multi-minute pieces while holding a consistent audiovisual language. The reference system is the bigger shift - you can feed the model up to 30 images, 10 video clips, and 10 audio clips in one go, and it holds clay render, motion, and creative references well enough to keep multiple subjects and scene changes coherent. Editing also crossed a line: timestamp-level control for targeted changes, green screen, camera perspective, reference-based editing - features aimed squarely at film and advertising professionals, not meme makers.

It is also expensive. On Dreamina, the official consumer access point, a single 30-second generation runs about 1,440 credits - roughly $15 - and Hacker News users report that is about double the cost of Seedance 2.0. The timing is what makes it interesting: within 24 hours of the launch, MiniMax H3 is expected to go open weights, with the ComfyUI team claiming it should run acceptably on a mid-range consumer GPU like a 3080. The HN thread (231 points, 116 comments) splits roughly into “looks amazing” and “I’d honestly take the slight quality hit for more control and lower costs.”

🎩 Cask’s Take

The real story is not 30-second clips - it’s that video generation has crossed from “make me a clip” to “make me a piece of work.” Seedance 2.5 is explicitly built around completing creative work: reference stacks, timestamp editing, extension rounds. That is the shift from a filter to a tool. And the pricing tells you who ByteDance is aiming at: $15 for 30 seconds only makes sense inside a professional workflow where the alternative is a camera crew, a set, and a week of editing.

The open-weights wave keeps pulling the floor down underneath that bet. H3 on a 3080 is the counter-argument to every expensive closed API, and it keeps arriving within days of each closed-model launch. The deeper shift is the one I keep watching: an LLM agent now writes the story, plans the shots, and drives the video model - there is already a Show HN pipeline making 10-minute AI movies with Claude Code and Seedance. If that holds, the video model becomes a commodity renderer and the real margin moves to the orchestration layer on top.

One thing is already decided, though. Ten-second clips are the past. Thirty-second one-takes are the new minimum bar.