ECCV Tutorial
September 9, 2026
Malmö, Sweden
Building Large Video Generation Models: Data Processing, Architectural Insights, Optimization Methods and Evaluation Strategies
Abstract
This tutorial provides a comprehensive and systematic review of the development process for modern large-scale video generation models. While visual generative AI research relies on pre-trained models, many practical challenges – data scaling, model architecture, training and inference optimization, and meaningful evaluation – remain uncovered in publications from the low-resource research groups yet are critical for real-world deployment and research practice. At the same time, many aspects of the development of modern large video models are shrouded in mystery, as this is often non-open-source proprietary information. This tutorial describes the full development pipeline, from petabyte-scale data curation and efficient Diffusion Transformer designs to multi-stage training and evaluation. Drawing on recent advances from open-source projects like Kandinsky 5.0, Wan, and HunyuanVideo, we balance theoretical insight with practical strategies for improving model efficiency and enabling low-resource usage. Attendees will gain actionable knowledge how to efficiently use the large-scale modern video generation systems.
To be updated
Tutorial details and materials will be posted here