Provides a unified multimodal framework for large-scale 3D understanding, text-to-3D generation, and instruction-guided 3D editing. Trains on an 87M-sample 3D multimodal corpus (25M understanding, 50M generation, 12M editing) and pairs a vision-language model with a diffusion-based 3D synthesizer to preserve structure and enable part-aware edits; suited for researchers building text-driven 3D asset pipelines but requires large compute and data.
Generates multi‑speaker speech and environmental audio from textual instructions or a reference clip, supporting zero‑shot voice cloning and detailed scene/specification control. Combines a cleaned, captioned dataset with a VAE-based multimodal generator, reward-conditioned quality control, and staged training to improve expressiveness and multi-audio modeling.
Systematically evaluates AI-generated video detectors and generators for real-world crisis scenarios using RA-Bench (17,886 clips: 1,830 real anchors, 16,056 generated). Shows detector families fail to generalize across generation conditions, and that human-misleading videos and social dissemination further degrade detection.
An index of Cara App content: metadata and CDN URLs for ~3.43M posts, 8.52M master artworks (~12M image links). Includes an SQLite catalog and Parquet exports but does not include image bytes — only links and metadata for analysis and search.
Trains compact conversational agents to adapt at runtime to changing 'Harness' configurations (Skills, Hooks, prompts, tools) using Harness-Aware Training (HAT): Harness-State Augmentation, on-policy distillation, and RL to preserve generality while meeting low-latency deployment constraints.
Generates and edits speech from natural-language instructions plus optional reference audio, supporting zero-shot TTS, content/acoustic/paralinguistic edits, enhancement, and source separation. Open-source 1.5B-parameter base model with a 4-step distilled AuK‑Flash for faster inference.
Provides Parallel Decoding Distillation (PDD) LoRA adapters that accelerate MiniMax-H3 video generation into few inference steps. Includes official 8-step Acc LoRAs for FL2VA and Ref2VA (rank=64, network_alpha=64, BF16), demo comparison videos, and example scripts using Diffusers' MiniMax-H3 ModularPipeline.
Generates short multimodal videos from text, images, or reference clips using a fine-tuned MiniMax‑H3 fusion model; improves HDR clarity, motion fluidity, distant-face fidelity and VFX while preserving MiniMax‑H3’s prompt/style behavior. Best used via ComfyUI.
Generates long-form, text-controlled music with explicit arrangement and planning. Uses a 50 Hz single-codebook tokenizer, a flow-matching diffusion Transformer to predict VAE latents, and an MoE autoregressor with ABC‑CoT planning to produce 48 kHz audio up to 5m30s.