Provides a harness that lets language models control embodied manipulation via iterative perception–reasoning–action loops, semantic action abstractions, and multimodal observations. Demonstrates distilling capabilities into a 4B open-source model with under 2K simulated trajectories and shows sim-to-real generalization.
Proposes ZPPO, a distillation method that keeps the teacher inside prompts rather than injecting teacher gradients, using binary- and negative-candidate prompts plus a prompt replay buffer to recover learning signal on hard examples; shows gains for small Qwen3.5 students across 31 multimodal benchmarks.
Provides a dual-path approach for spatial vision-language models: a Language-Only Reasoning (LOR) path for stepwise linguistic deduction and a Detect-Then-Reason (DTR) path that detects 3D cues via region tokens before numerical inference. Trains with chain-of-thought cold-start supervision and reinforcement learning to improve 3D grounding and multi-step spatial reasoning.
Evaluates multimodal LLMs' ability to reconstruct past observations and act in controllable non-Markov games. Introduces RNG-Bench with two games (Matching Pairs, 3D Maze), three controllable difficulty axes, a head-to-head duel protocol, and a Memory Gap metric to separate forgetting from action errors.
Generates images from natural-language prompts as an 8-step distilled checkpoint of Krea 2, optimized for fast iterative text-to-image workflows with style references and 1K–2K resolution outputs.
Performs one-shot, long-horizon OCR and document parsing by using Reference Sliding Window Attention (R-SWA) to keep the decoder KV cache constant, enabling single-pass multi-page transcription; code, model weights and an accompanying arXiv report are provided.
Provides GGUF-quantized weights and runtime assets for running the Qwythos-9B reasoning LLM locally via llama.cpp and compatible runtimes. Key features include 1,048,576-token YaRN long-context, native function-calling, multimodal image input (requires mmproj), and multiple quantization/MTP variants tuned for different size/quality tradeoffs.
Provides multiview synthetic RGB video clips with per-frame depth, instance masks, dense long-range 3D point tracks, camera poses, and SMPL‑X human pose/shape labels for 4D reconstruction, tracking, and geometry-aware novel-view synthesis. Includes ~4.7K clips (1.4M frames) and is licensed for AI training.
Provides ~2 million instruction-aligned video-edit pairs for training and evaluating instruction-based video editing and generation models. Covers multi-task and structural edits (e.g., camera/subject movement), produced via a synthesis pipeline with progressive filtering; licensed CC BY-NC-4.0.
Multimodal video dataset for text-to-video and video-to-video research: about 2 million short English videos and extracted frames for instruction-based video editing and generation. Hosted on Hugging Face and licensed CC BY‑NC 4.0 (non-commercial).
NVFP4-quantized variant of Qwen3.6-27B that reduces parameter bits from 16 to 4, cutting disk and GPU memory requirements by ~2.5× while keeping comparable benchmark accuracy; ready for vLLM-based inference on NVIDIA hardware and supports long, multimodal contexts.
Provides a rubric-based benchmark that converts dense image captions into instance-specific atomic checks (Must-Right and Easy-Wrong) and a gated scoring rule, aiming to expose perceptual brittleness and better align multimodal model evaluation with human judgment.