Wan 2.7 is what happens when Alibaba asks: what if the AI actually thought about your video before making it?
Every previous version of the Wan series — and most AI video generators generally — work the same way: receive prompt, generate video, hope for the best. Wan 2.7 breaks that pattern. Its Thinking Mode interprets your prompt, plans the scene structure and narrative arc, considers composition and motion dynamics, and then begins rendering. The output is less generic, less drifty, and more intentional. It’s the difference between handing a script to a camera operator and handing it to a director.
The technical foundation is a 27B Mixture-of-Experts architecture with a Diffusion Transformer and Full Attention mechanism — designed specifically to process spatial and temporal relationships simultaneously. This isn’t just a scaled-up version of earlier Wan models; it’s a different approach to how video generation should work.
The practical upgrades are equally significant. First and Last Frame control lets you define the exact start and end of a clip — critical for maintaining continuity across scenes. Up to 9 multimodal reference inputs (images, clips, audio) keep characters, props, and styles consistent without manual stitching. Native audio sync generates sound in the same pass as the visuals. And instruction-based editing means you can refine outputs by describing the change you want, rather than regenerating from scratch.
The Apache 2.0 license remains the Wan series’ defining commitment. Not “open with enterprise restrictions.” Not “free for personal use.” Apache 2.0 — the same license that governs the infrastructure half the internet runs on. Use it commercially, modify the weights, build products, sell the output. Zero asterisks.
The community that made Wan 2.1 the ComfyUI standard is already moving to 2.7. The ecosystem will follow — it always does.