ECCV Tutorial

September 9, 2026

Malmö, Sweden


Building Large Video Generation Models: Data Processing, Architectural Insights, Optimization Methods and Evaluation Strategies

Viacheslav Vasilev, Maria Kovaleva, Vladimir Korviakov, Denis Dimitrov

Kandinsky Lab

Contact: viacheslav.vasilev@kandinskylab.ai

Abstract

This tutorial provides a comprehensive and systematic review of the development process for modern large-scale video generation models. While visual generative AI research relies on pre-trained models, many practical challenges – data scaling, model architecture, training and inference optimization, and meaningful evaluation – remain uncovered in publications from the low-resource research groups yet are critical for real-world deployment and research practice. At the same time, many aspects of the development of modern large video models are shrouded in mystery, as this is often non-open-source proprietary information. This tutorial describes the full development pipeline, from petabyte-scale data curation and efficient Diffusion Transformer designs to multi-stage training and evaluation. Drawing on recent advances from open-source projects like Kandinsky 5.0, Wan, and HunyuanVideo, we balance theoretical insight with practical strategies for improving model efficiency and enabling low-resource usage. Attendees will gain actionable knowledge how to efficiently use the large-scale modern video generation systems.

To be updated

Tutorial details and materials will be posted here