← all repositories

test-time-training/ttt-video-dit

A text-to-video generation model based on CogVideoX that extends video context to 63 seconds using Test-Time Training layers.

ttt-video-dit
Not currently ranked — collecting fresh signals.
star history

This repository provides training and inference code for a diffusion transformer that generates videos up to 63 seconds long. The architecture adapts the CogVideoX 5B model by incorporating Test-Time Training layers to handle long-range relationships across the global context while retaining original attention layers for local 3-second segment processing. The model is fine-tuned in stages at increasing video lengths (3s, 9s, 18s, 30s, 63s) for context extension.

Frequently asked

What is test-time-training/ttt-video-dit?
A text-to-video generation model based on CogVideoX that extends video context to 63 seconds using Test-Time Training layers.
Is ttt-video-dit open source?
Yes — test-time-training/ttt-video-dit is open source, released under the MIT license.
What language is ttt-video-dit written in?
test-time-training/ttt-video-dit is primarily written in Python.
How popular is ttt-video-dit?
test-time-training/ttt-video-dit has 2.4k stars on GitHub.
Where can I find ttt-video-dit?
test-time-training/ttt-video-dit is on GitHub at https://github.com/test-time-training/ttt-video-dit.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.