High-resolution joint video-and-audio generation is limited by the quadratic cost of full attention and by redundant token interactions that wash out informative audio-driven regions. Prism attacks this bottleneck by reorganizing the token sequence into spatiotemporal macro-zones and assigning dynamic, anisotropic block shapes per zone so attention focuses where visual variation and audio-visual coupling matter most. The result is much lower attention compute without sacrificing — and in reported experiments sometimes improving — generation quality.
What Sets It Apart
- Dynamic block-shape assignment: Prism estimates local video feature variance and audio-to-video cross-attention norms to choose finer partitions along axes with rapid visual change or strong audio influence, so blocks remain semantically coherent rather than using a fixed sparse pattern. This preserves cross-modal signals around sound-producing regions.
- Hybrid per-query sparsity: Combines top-k block selection with a top-p/CDF threshold to adapt sparsity per query, balancing compute reduction and expressiveness.
- Practical tooling and checkpoints: The project publishes preview alpha/beta checkpoints, training and inference scripts (supports FSDP and sequence parallelism), and clear dataset bucketing and latent-extraction utilities for native 720p/1080p/2K training and inference.
- Empirical trade-off: Demonstrated ~2.5× training speedup vs full attention in the authors' experiments while matching or surpassing generation quality for joint video-audio diffusion models.
Who It's For and Trade-offs
Great fit if you need native high-resolution joint video+audio diffusion training or inference and have access to large GPU resources (preview inference: single 80GB GPU for 720p; 4+ × 80GB GPUs for 1080p/2K; training: 32 × 80GB for 720p, 64 × 80GB for 1080p/2K). Choose Prism when attention compute is the bottleneck and preserving localized audio-visual interactions matters. Look elsewhere if you lack access to modern CUDA toolchains (requires Python ≥3.10, CUDA ≥12.4, FlashAttention) or cannot provision very large GPU clusters — integration and high-res runs assume heavy infra and a pretrained MOVA base model dependency. The repository is aimed at researchers and engineering teams, not as a lightweight drop-in model for single-GPU casual use.