10,510 real scenes for when synthetic 3D isn't enough
It closes the gap between synthetic 3D benchmarks and messy reality with 10,510 annotated real-world videos and camera poses.

What it does
DL3DV-10K is a collection of 10,510 real-world videos—51.2 million frames in all—captured across 65 types of locations and annotated with scene attributes like indoor/outdoor, reflection, transparency, and lighting. Every video comes with COLMAP-calculated camera poses, and the project maintains a 140-scene novel-view-synthesis benchmark that reports quantitative results for several SOTA methods. The dataset is hosted on HuggingFace in tiers from 480P (~730 GB) up to full 4K (~44 TB), so you can match download size to your storage budget.
The interesting bit
The student-led project has already been adopted by Stability.ai and NVIDIA’s Cosmos for camera-control research, which is a decent signal that industry sees real-world, pose-annotated video as fuel for world models. The annotations go beyond simple labels: they explicitly tag bounded versus unbounded scenes and material complexity, which lets researchers filter for exactly the failure modes they want to study.
Key highlights
- 10,510 scenes and 51.2 million frames at 4K, with downsampled tiers down to 480P
- 65 point-of-interest categories covering both bounded and unbounded environments
- 140-scene benchmark with COLMAP poses; quantitative results reported for 3D Gaussian Splatting, ZipNeRF, Mip-NeRF 360, Instant-NGP, and Nerfacto
- Already used by Stability.ai and NVIDIA Cosmos for camera-control and world-model training
- Preview page and subset download script let you browse before committing terabytes
Caveats
- The full 4K frame dataset weighs roughly 44 TB; even the 480P version is ~730 GB
- Benchmark trained weights for the SOTA methods are marked as “coming soon”
- The preview page notes that some scene labels are still missing and will be updated
Verdict
Grab this if you are training or evaluating NeRFs, 3D Gaussian Splatting, or world models that need real-world camera trajectories. Skip it if your work is strictly synthetic or your storage topology cannot handle hundreds of gigabytes at minimum.
Frequently asked
- What is DL3DV-10K/Dataset?
- It closes the gap between synthetic 3D benchmarks and messy reality with 10,510 annotated real-world videos and camera poses.
- Is Dataset open source?
- Yes — DL3DV-10K/Dataset is an open-source project tracked on heatdrop.
- What language is Dataset written in?
- DL3DV-10K/Dataset is primarily written in HTML.
- How popular is Dataset?
- DL3DV-10K/Dataset has 650 stars on GitHub.
- Where can I find Dataset?
- DL3DV-10K/Dataset is on GitHub at https://github.com/DL3DV-10K/Dataset.