Evaluates metric 3D spatial reasoning from single driving images via multiple-choice questions that require reconstructing scene geometry rather than relying on image-layout shortcuts. Each sample pairs a numbered-bbox image with a question, four choices, and the correct answer; images come from PlusAI and the dataset is CC BY 4.0.
Text-to-image model packaged for Diffusers that uses fp8 quantization to lower memory and speed up inference. Delivered as a safetensors checkpoint on Hugging Face with an Ideogram pipeline; created May 30, 2026 — license unspecified.
Provides 462 unrestricted long-form chain-of-thought reasoning traces distilled from the full Mythos V2 model (≈104.7M characters); intended for long-context evaluation, trace analysis and process-level supervision. License unknown—verify before reuse.
Generates and reasons about multimodal physical-world content—text, images, video, audio, and robot/action trajectories—conditioned on combinations of text, image, video and action inputs. The 64B “Super” variant targets Physical AI use cases and supports vLLM‑Omni, Diffusers, and action prediction.
Provides the renderer weights and inference code for Bernini’s video renderer, enabling text→video, image→video and video editing inference. Offers a ready diffusers-format bundle or safetensors checkpoints under Apache‑2.0; intended for multi‑GPU/Hopper inference and reproducible research.
Provides per-cell transcriptomes and five-day drug-sensitivity readouts for 1.83M single cells across 52 cancer cell lines and 91 drug conditions, with raw counts plus gene, cell-line, drug, and summary metadata for modeling drug response and context-dependent gene function.
Provides the gated, official OSWorld 2.0 Python task class files (task_*.py) required to run the benchmark; distributed via a Hugging Face gated dataset to reduce benchmark leakage. Download requires accepting gated access on Hugging Face.
A 20B retrieval subagent trained with reinforcement learning inside a stateful search harness that externalizes recoverable search state (candidate pool, curated evidence, verification records). The harness lets the policy focus on semantic search decisions, improving curated recall and transfer robustness.
Omnimodal world model that jointly processes and generates text, images, video, audio, and action trajectories for physical AI. Uses a mixture-of-transformers to combine autoregressive reasoning and diffusion-based multimodal generation; released open-source with checkpoints, datasets and benchmarks for robotics and simulation.
Benchmark dataset for evaluating long-horizon coding agents and software-engineering tasks, containing English code and tabular metadata in Parquet format; small scale (<1K examples) for fast prototyping and evaluation.
Generates minute-level, multi-shot synchronized audio+video from a single text prompt, using a paired cross-modal memory to preserve character appearance and voice across shots. Uses DMD-distilled few-step inference for ~7.5× speedup; requires high-GPU memory and is released under the LTX-2 community license.
Native multimodal model for image/text/video→text tasks with million‑token context support. Uses a sparse-attention operator to cut long‑context compute and latency, and targets agentic, coding, and long-horizon conversational workloads.