Pretrained image-model checkpoint hosted on Hugging Face by Facebook (Meta) for vision experiments and transfer learning. Includes downloadable weights and metadata under CC BY‑NC 4.0 — suitable for research and prototyping but restricted for commercial use.
Multimodal image-text-to-text fork of Gemma 4 (31B) using a 'CRACK v2' abliteration — tuned for conversational vision inputs and thinking-mode support in JANG v2 safetensors format. Recommended to run in vMLX; published by dealignai.
Generates persistent, explorable 3D worlds from a single image by synthesizing long-range, geometry-consistent video and reconstructing it into an explicit 3D Gaussian scene. Intended for internal research use under NVIDIA's research license.
An open text-to-image generation model built on an 8B Diffusion Transformer that focuses on layout-sensitive, text-heavy, and instruction-following image synthesis. Notable for accurate text rendering, structured/compositional generation (posters, comics), and ability to run on consumer 24GB GPUs when paired with prompt enhancement.
Provides 12.26M synthetically generated multilingual OCR samples (en/ja/ko/ru/zh) with word/line/paragraph bounding boxes and reading-order graphs, packaged as HDF5 shards for training detection, recognition, and layout models; licensed CC BY 4.0.
Generates and reconstructs navigable, editable 3D worlds from text, single images, multi-view photos, or video; outputs meshes and Gaussian Splatting assets and includes WorldMirror 2.0 for fast multi-view reconstruction. Suited for research and production pipelines that import assets into engines; requires substantial GPU resources.
Performs feed‑forward streaming 3D reconstruction from image sequences, combining coordinate grounding, dense geometric cues and trajectory memory to correct long‑range drift; uses paged KV‑cache attention for ~20 FPS inference at 518×378 and supports sequences >10,000 frames.
Unifies multimodal understanding, reasoning, and image generation in a single end-to-end architecture using the NEO-unify paradigm. Models pixels and words jointly without a separate visual encoder, and provides interleaved image–text generation, infographic editing, and GGUF/low‑VRAM inference options.
An uncensored, fully unlocked GGUF port of Qwen 3.6‑35B‑A3B for local multimodal (text+image) inference, offering K_P 'Perfect' quant variants (Q8–Q2) and an mmproj for vision. Suited for offline research and experimentation; not for use-cases requiring safety filters.
About 9,700 synthetic full-page web screenshots with YOLO-format, pixel-aligned bounding boxes for 14 UI element classes, generated by LLM-augmented HTML and Playwright DOM extraction. Includes CC3M image injection to reduce visual gap; released for non-commercial research (CC BY-NC-SA 4.0).
Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.
A 1.4M image–text style dataset for text-to-image generation and style transfer, produced by mapping 170K curated style prompts to 400K content prompts via Qwen-Image to yield strong intra-style consistency. Designed for training and evaluating style-aware generative models; license: other.