Phantom-video/Phantom
A text-to-video generation model that maintains subject consistency across generated videos using cross-modal alignment.

Phantom is a video generation model developed by ByteDance that produces videos from text prompts while preserving subject consistency across frames. It employs cross-modal alignment techniques to ensure visual coherence between text descriptions and generated video content. The project includes model weights, training code, and inference tools, along with a companion dataset (Phantom-Data) for improving subject consistency in generative video models.
Frequently asked
- What is Phantom-video/Phantom?
- A text-to-video generation model that maintains subject consistency across generated videos using cross-modal alignment.
- Is Phantom open source?
- Yes — Phantom-video/Phantom is open source, released under the Apache-2.0 license.
- What language is Phantom written in?
- Phantom-video/Phantom is primarily written in Python.
- How popular is Phantom?
- Phantom-video/Phantom has 1.5k stars on GitHub.
- Where can I find Phantom?
- Phantom-video/Phantom is on GitHub at https://github.com/Phantom-video/Phantom.