AIAny
AI Video2026
Icon for item

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Enables native joint video-and-audio generation and training up to 2K using a dynamic block-sparse attention scheme that cuts attention cost while preserving audio–visual coupling; provides preview checkpoints, training and inference code. High-resolution training requires large multi-GPU (80GB) clusters and CUDA ≥ 12.4.

Introduction

High-resolution joint video-and-audio generation is limited by the quadratic cost of full attention and by redundant token interactions that wash out informative audio-driven regions. Prism attacks this bottleneck by reorganizing the token sequence into spatiotemporal macro-zones and assigning dynamic, anisotropic block shapes per zone so attention focuses where visual variation and audio-visual coupling matter most. The result is much lower attention compute without sacrificing — and in reported experiments sometimes improving — generation quality.

What Sets It Apart
  • Dynamic block-shape assignment: Prism estimates local video feature variance and audio-to-video cross-attention norms to choose finer partitions along axes with rapid visual change or strong audio influence, so blocks remain semantically coherent rather than using a fixed sparse pattern. This preserves cross-modal signals around sound-producing regions.
  • Hybrid per-query sparsity: Combines top-k block selection with a top-p/CDF threshold to adapt sparsity per query, balancing compute reduction and expressiveness.
  • Practical tooling and checkpoints: The project publishes preview alpha/beta checkpoints, training and inference scripts (supports FSDP and sequence parallelism), and clear dataset bucketing and latent-extraction utilities for native 720p/1080p/2K training and inference.
  • Empirical trade-off: Demonstrated ~2.5× training speedup vs full attention in the authors' experiments while matching or surpassing generation quality for joint video-audio diffusion models.
Who It's For and Trade-offs

Great fit if you need native high-resolution joint video+audio diffusion training or inference and have access to large GPU resources (preview inference: single 80GB GPU for 720p; 4+ × 80GB GPUs for 1080p/2K; training: 32 × 80GB for 720p, 64 × 80GB for 1080p/2K). Choose Prism when attention compute is the bottleneck and preserving localized audio-visual interactions matters. Look elsewhere if you lack access to modern CUDA toolchains (requires Python ≥3.10, CUDA ≥12.4, FlashAttention) or cannot provision very large GPU clusters — integration and high-res runs assume heavy infra and a pretrained MOVA base model dependency. The repository is aimed at researchers and engineering teams, not as a lightweight drop-in model for single-GPU casual use.

Information

  • Websitehuggingface.co
  • OrganizationsFudan University, Tencent Hunyuan, Zhejiang University
  • AuthorsShuyuan Tu, Qi Tian, Yinming Huang, Yue Wu, Xintong Han, Kaihang Pan, Weijie Kong, Jiangfeng Xiong, Jian-Wei Zhang, Zuxuan Wu …
  • Published date2026/09/24

More Items

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Hugging Face
AI Video2026

Replaces a selected person in a source video with a character from a reference image via a MiniMax H3 LoRA adapter, aiming to preserve scene, camera, and background. Trained for 1,000 updates; intended for Ref2VA runtimes and ComfyUI. Short (≈4–5s) continuous shots work best; distributed under the MiniMax H3 Community License.

Hugging Face
AI Video2026

LoRA adapters for MiniMax H3 that sharpen and enhance videos in ComfyUI by conditioning on source clips via guide latents for pixel-level alignment. Designed mainly for ref2va as a second-pass sharpening tool, includes a ComfyUI workflow and example before/after clips; requires aligned guide clips at the target resolution and valid clip lengths.