End-to-end multimodal model for native text↔image understanding, interleaved image-text generation, and image editing. Uses the NEO-Unify MoT architecture to avoid separate visual encoders/VAE. Suited for multimodal prototyping, demos, and research (Apache‑2.0).
Provides an NVFP4-quantized 27B Qwen3.6 checkpoint optimized for faster, low-memory multimodal inference on 24GB GPUs. Includes MTP (multi-token prediction), extended 262k native context, and deployment recipes for vLLM/SGLang/KTransformers; best used with recommended backends for peak throughput.
Provides a 30K+ problem multimodal, multilingual dataset of Olympiad-level math problems with expert solutions and a math-aware retrieval benchmark—includes images, hierarchical topics, provenance from official booklets, and LLM-assisted metadata (v0, CC BY 4.0).
A GGUF-format preview checkpoint derived from Qwen3.6-27B — a multimodal, image-text-to-text reasoning model fine-tuned for more structured reasoning and consistent answer style; packaged for local inference and compatible with engines like vLLM/SGLang/llama.cpp.
High-resolution vision transformers pretrained on one billion human images for human-centric tasks such as pose estimation, body-part segmentation, surface-normal and pointmap prediction. Provides multiple backbone sizes and task-specific checkpoints; released under the Sapiens2 license.
Provides satellite image tiles paired with per-tile land-cover captions and bounding-box overlays in SFT-compatible JSONL for supervised fine-tuning. Includes RGB chips, optional Mapbox context, metadata, and train/validation/test splits derived from Sentinel‑2 and Earth Engine labels.
Provides a lightweight assistant (draft) model for Gemma 4 E4B used in speculative-decoding pipelines — it predicts token drafts that the target model verifies in parallel, enabling up to ~2× decoding speedups while preserving identical final outputs. Useful for low-latency, multimodal assistant and on-device scenarios.
Acts as the assistant (drafter) checkpoint for Gemma 4 26B A4B on Hugging Face, used in Speculative Decoding to pre-draft tokens and speed up generation. Designed for long-context, multimodal workflows where lower latency and on-device or edge inference matter.
Unifies video, audio, image and text understanding for enterprise Q&A, summarization, transcription and document intelligence. The NVFP4 quantized variant reduces footprint to ~20.9GB for more efficient single‑GPU deployment and is tuned for NVIDIA runtimes (vLLM, TensorRT).
Provides ~55K multimodal VQA items with matched contrastive pairs and model‑generated rationales across five categories (General, Reasoning, Math, Graph/Chart, OCR), enabling research on faithful visual reasoning and robustness. Train split: 54,844 examples; license unspecified—verify before use.
Unified omnimodal foundation model for text, image, video and audio understanding and agentic workflows, with support for up to 1M-token context. Combines a sparse MoE LLM backbone, dedicated vision/audio encoders, multi-token prediction, and a hybrid sliding-window + global attention design to reduce KV-cache overhead.
Collection of 76 image-centric multimodal subdatasets (≈6.9M samples, ~39.56B estimated tokens) for training vision–language models, each published with a standardized conversation JSONL and dataset card. Media are referenced by path/URL and must be fetched separately; licensing is primarily CC-BY-4.0 with per-subdataset variations.