Generates persistent, explorable 3D worlds from a single image by synthesizing long-range, geometry-consistent video and reconstructing it into an explicit 3D Gaussian scene. Intended for internal research use under NVIDIA's research license.
A 30B-parameter, instruction-tuned language model built for long-context text generation, conversational agents, and tool-calling. It combines supervised fine-tuning and RL alignment, supports 131,072-token context, and is optimized for tasks like summarization, code, and RAG.
An open text-to-image generation model built on an 8B Diffusion Transformer that focuses on layout-sensitive, text-heavy, and instruction-following image synthesis. Notable for accurate text rendering, structured/compositional generation (posters, comics), and ability to run on consumer 24GB GPUs when paired with prompt enhancement.
Provides a compact GGUF export of a tuned Gemma‑4 26B variant for local inference, optimized for llama.cpp and Apple Silicon to deliver faster, less‑censored chat and coding outputs. Includes Q4_K_M quantization and a neutral embedded template for more reliable local deployments.
Multimodal agent model for long-horizon coding, image-text understanding, and autonomous task orchestration. Built as a 1T-parameter Mixture-of-Experts with 256K context and native int4 quantization — intended for coding-driven design, persistent background agents, and swarm-style sub-agent workflows.
Open-weight multimodal 35B Qwen3.6 model in Hugging Face Transformers format that supports image/video/text inputs and native long contexts (262,144 tokens). Emphasizes agentic coding and preserved reasoning traces (thinking), uses an MoE-backed architecture and is designed for self-hosting with vLLM/SGLang/KTransformers; requires multi-GPU resources for production.
Reconstructs camera poses and dense 3D point clouds from video streams using a feed‑forward foundation model. Combines a Geometric Context Transformer (anchor + local window + trajectory memory) with paged KV‑cache attention to enable stable, long‑sequence streaming inference (~20 FPS at 518×378).
Performs feed‑forward streaming 3D reconstruction from image sequences, combining coordinate grounding, dense geometric cues and trajectory memory to correct long‑range drift; uses paged KV‑cache attention for ~20 FPS inference at 518×378 and supports sequences >10,000 frames.
Unifies multimodal understanding, reasoning, and image generation in a single end-to-end architecture using the NEO-unify paradigm. Models pixels and words jointly without a separate visual encoder, and provides interleaved image–text generation, infographic editing, and GGUF/low‑VRAM inference options.
An uncensored, fully unlocked GGUF port of Qwen 3.6‑35B‑A3B for local multimodal (text+image) inference, offering K_P 'Perfect' quant variants (Q8–Q2) and an mmproj for vision. Suited for offline research and experimentation; not for use-cases requiring safety filters.
Generates English text matching pre-1931 style — a 13B language model trained on ~260B tokens of pre-1931 English, useful for historical-language generation and stylistic research. An instruction-tuned variant exists for interactive tasks.
Instruction-tuned 13B LLM post-trained on 260B tokens of pre-1931 English and finetuned with online DPO (LLM-as-judge) to improve instruction-following; suited for period-style English generation and etiquette/letter-writing formats, but not optimized for contemporary factual updates.