Unmixes green‑screen pixels with a neural model to recover straight (unmultiplied) foreground color and a clean linear alpha for every pixel, preserving hair, motion blur and translucency. Produces VFX‑standard EXR outputs, supports optional AlphaHint generators (GVM/VideoMaMa) and Docker/consumer‑GPU optimizations.
Provides 6,000 runnable, operator-level PyTorch tasks for training and evaluating CUDA kernel generation models; each sample includes executable code, operator descriptors, and provenance tags, with execution-driven filtering to ensure reproducibility and contamination control.
Lets an LLM autonomously propose, edit, run, and evaluate short single‑GPU LLM training experiments — fixed 5‑minute runs (~12 experiments/hour). Agent edits a single train.py; humans supply goals via program.md. Single‑GPU, val_bpb metric.
Generate text, images, video, audio and action/robot trajectories from combined text, image, video, audio and action inputs. A Mixture-of-Transformers omnimodal foundation model (Cosmos3‑Nano, 16B params) focused on Physical AI (robotics, AV, simulation) and optimized for NVIDIA GPU runtimes.
Instruction-tuned Mixture-of-Experts multimodal model that generates text from text+image inputs while activating a 4B subset of parameters for faster inference; supports a 256K context window, multilingual vision-language tasks, and is available under Apache-2.0.
Pretrained image-model checkpoint hosted on Hugging Face by Facebook (Meta) for vision experiments and transfer learning. Includes downloadable weights and metadata under CC BY‑NC 4.0 — suitable for research and prototyping but restricted for commercial use.
Converts text to natural-sounding speech across 600+ languages in a zero-shot way, with short-reference voice cloning and fine-grained voice-design controls; uses a diffusion language-model-style architecture to balance quality and very low inference latency.
Generates 48kHz multilingual speech from text using a tokenizer-free diffusion-autoregressive TTS architecture, supporting natural-language voice design, controllable cloning, and low-latency streaming. Notable for a 2B-parameter backbone and built-in AudioVAE super-resolution (16k→48k).
An 8B-parameter, instruction-tuned long-context LLM optimized for instruction following, tool-calling, and multilingual dialogue — supports 131072-token context and common NLP tasks such as summarization, QA, code, and RAG.
An open text-to-image generation model built on an 8B Diffusion Transformer that focuses on layout-sensitive, text-heavy, and instruction-following image synthesis. Notable for accurate text rendering, structured/compositional generation (posters, comics), and ability to run on consumer 24GB GPUs when paired with prompt enhancement.
Generates and reconstructs navigable, editable 3D worlds from text, single images, multi-view photos, or video; outputs meshes and Gaussian Splatting assets and includes WorldMirror 2.0 for fast multi-view reconstruction. Suited for research and production pipelines that import assets into engines; requires substantial GPU resources.
Generates text by iteratively denoising blocks of tokens with a two-tower design: a frozen autoregressive context tower and a trainable diffusion denoiser tower, trading minimal quality loss for higher wall-clock throughput.