Provides 104.9M curated image–text pairs with precomputed embeddings, structured annotations and pre-encoded VAE latents for text-to-image pretraining and retrieval. Combines filtered web sources and synthetic samples with multi-model re-captioning, deduplication and safety filters; Apache-2.0.
Curated multimodal training corpus for spatial intelligence: ~8.16M QA-style samples paired with ~2.72M unique images (≈1.1 TB). Provides JSONL annotations, a 1,000-sample preview, and 52 independent image archives — used to train SenseNova-SI models.
Unified 4B vision-language model for document understanding that converts images or text into template-driven structured JSON or clean Markdown. Key features: multimodal inputs (image+text), template-based extraction, reasoning vs non-reasoning modes, and vLLM/OpenAI-compatible deployment for OCR, invoice/forms extraction, and RAG preprocessing.
Provides ~85K contrastive visual question–answer pairs where each example contains an anchor and a matched counterpart (image, question, answer). Pairs span General, Reasoning, Math, Graph/Chart and OCR categories to help train and evaluate fine‑grained, faithful visual reasoning in VLMs.
An uncensored, fine-tuned and GGUF-quantized variant of Qwen3.6-27B tailored for long-context, coding, vision and creative-writing use. Offers multiple NEO-CODE Di-Matrix quants (IQ2/IQ4/Q6/Q8), mmproj vision support and recommended inference settings for local servers.
Generates high-fidelity 3D assets from a single image by back-projecting pixel-aligned features into 3D, preserving fine geometry and PBR textures; includes inference code and a Hugging Face demo—best suited for single-view object reconstruction.
Provides the dataset and accompanying technical report for a DeepSeek project that interleaves spatial markers (points and boxes) into multimodal LLM reasoning. Includes a public subset of data and benchmarks under an MIT license; model weights are not included.
A 40B GGUF-quantized Qwen3.6 variant fine-tuned with Claude 4.6 Opus and Deckard/Heretic datasets for multimodal image-text-to-text tasks. Offers 256K context, custom NEO-CODE Di-IMatrix quants for long conversations and coding, optimized for local inference and creative/coding use cases; safety alignment removed.
Provides aligned urban driving sensor streams (camera frames, LiDAR, radar and HD‑map / lanelet2 annotations) for multimodal perception, tracking and mapping research. Expert-generated labels under CC BY‑NC‑4.0 and hosted on Hugging Face.
Transforms pretrained latent-diffusion priors into pixel-space diffusion models by removing the VAE and training shallow pixel layers on LDM-generated synthetic images — enabling fast convergence, native 4K output, and low-data training on 8 GPUs.
Provides paired images and English captions for vision–language research, curated by Stanford Vision Lab and hosted on Hugging Face; useful for training and evaluating multimodal models and reproducing related research.
Large-scale synthetic video dataset of physically simulated multi-object interaction scenes for training and evaluating models on physical reasoning, depth and optical-flow estimation, instance segmentation, and physics-grounded captioning. Provides RGB + lossless depth, per-frame instance masks, per-object physics annotations (NPZ), VLM-grounded captions, and USD scene files — useful for world-model and simulation-to-real work; commercial use permitted.