A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.
Packaged diffusers checkpoint of MiniMax H3 for image/text-to-short-video generation with native stereo audio; provided for direct use in diffusers image-to-video pipelines and aimed at easy integration into prototyping and production workflows.
Provides ComfyUI-compatible conversions and LoRA adapters of the MiniMax‑H3 video+audio generative model, with example presets and demo videos to run short stereo audio+video inference inside ComfyUI workflows.
Provides raw, unscripted first-person household video footage for training vision and embodied AI models. Released incrementally on Hugging Face in WebDataset shards with metadata parquets under Apache‑2.0; current raw tier contains ~7,834 hours (≈397k videos).
Provides 90,000 hours of head-mounted egocentric video paired with synchronized 3D hand pose and an optional 3D full‑body pose add-on, with event-level semantic labels available as a complimentary layer — designed for embodied AI and robotics training at scale.
Provides 1,000 five-second video clips generated by MiniMax H3 for lightweight evaluation of multimodal generation and understanding. Clips are roughly 768p base resolution with diverse aspect ratios and themes, produced with a pruned int8 minimax_h3_fl2va checkpoint at 30 steps.
A LoRA adapter for MiniMax H3 that improves photorealistic rendering of people—preserving skin texture, coherent micro-expressions, film-style lighting and subtle handheld motion. Trigger word: r34l1sm; intended for text-to-video portrait and close-up shots.
Encodes videos into a Film Knowledge Graph and reconstructs them to learn agent-native, editable video representations for agentic reasoning and manipulation. Uses agentic auto-encoding with dual-loop textual-gradient optimization, reports large reconstruction gains, and releases a benchmark and dataset.
Builds an editable, persistent 3D world state to drive iterative previsualization for film, games, and design — enabling local edits and recombinations instead of one-shot video regeneration. Uses separate stages for state construction, state evolution, and state access, with render-feedback camera refinement.
Externalizes persistent scene state into a camera-indexed world bank and designs a long-horizon teacher whose sparse-attention supervision is distilled into a three-step student, enabling responsive, low-latency interactive long-horizon video generation with bounded denoiser context.
Predicts future video frames conditioned on an observed frame, a language instruction, and a sequence of end-effector poses and gripper states for robot manipulation. Uses per-arm SE(3) geometric encoding (PRoPE-style), a lightweight depth branch, SAM3 masks with a frozen V-JEPA teacher, and distribution-matching distillation for efficient, consistent action-conditioned rollouts.
Systematically evaluates AI-generated video detectors and generators for real-world crisis scenarios using RA-Bench (17,886 clips: 1,830 real anchors, 16,056 generated). Shows detector families fail to generalize across generation conditions, and that human-misleading videos and social dissemination further degrade detection.