Provides 55 million scene-level video clips (each with captions, language labels, and timestamps) extracted from an 80M-video, 10-million-hour raw pool to support multimodal pre-training across video, audio, and frames. Access is gated for academic/non-commercial research.
Merges Unsloth UD XL quantized GGUF of Qwen3.6-27B with compact Q8_0 MTP heads to enable multi-token (speculative) decoding on llama.cpp builds that support MTP; aimed at image-text-to-text usage with reduced MTP overhead.
A small image-folder dataset for multimodal/vision model safety benchmarking, containing under 1,000 curated images with annotations to exercise safety-related model behaviours; licensed CC BY 4.0 and hosted on HuggingFace.
A reasoning-enhanced Mixture-of-Experts (MoE) LLM fine-tuned for multimodal image-text-to-text tasks and long-context reasoning; built on Qwen3.6-35B-A3B with LoRA and released as an experimental GGUF community model.
Open egocentric multimodal dataset for embodied AI and robot learning captured on commodity iPhone Pro: ~200 hours and ~10M RGB frames with LiDAR depth, ARKit 6‑DoF poses, IMU, two‑hand MANO mocap, room meshes, and hierarchical action captions.
Distilled dev checkpoint of an image foundation model that natively unifies raw pixels and text tokens for text-to-image, image editing, long-text rendering, and subject-driven personalization at up to 2048×2048. The Dev variant targets faster (28-step) inference for iterative use and research.
Generates and edits high-resolution images (up to 2048×2048) from text and reference images, plus subject-driven personalization. Implements a pixel-level unified transformer that encodes raw pixels and text in one token space and includes a reasoning-driven prompt agent for layout and text rendering.
Provides a curriculum-aligned knowledge graph extracted from Chinese K–12 textbooks and accompanying benchmarks and training data to evaluate and train educational LLMs. Releases a 23,640-question multi-select benchmark and a 7,335-sample graph-guided training corpus with multimodal VQA pairs and the full construction pipeline.
Processes text and images to produce conversational, reasoning-focused multilingual outputs for agentic workflows. Built as a sparse MoE decoder (25B active / 218B total parameters) with 128K context and available in BF16/FP8/W4A4 quantizations to balance quality and deployability.
Multimodal STEM problem set for verifiable, answer-supervised training and RL: contains single-image, multi-panel, and multi-image PhD-level questions across physics, math, chemistry and biology. Each example has a deterministic ground-truth answer, enabling reward modeling and automated evaluation.
Provides a quantized GGUF build of Qwen3.6‑27B with MTP (multi‑token prediction) support for faster local inference. Packaged for GGUF-compatible runners (llama.cpp) and Hugging Face/transformers workflows, with deployment notes for CPU/GPU and vLLM/SGLang integration.
A GGUF-quantized build of Qwen3.6-35B packaged by unsloth for local and accelerated inference. Adds MTP speculative decoding guidance and deployment notes for llama.cpp, vLLM, SGLang and long-context/multimodal use cases.