Provides quantized GGUF weights and configs for Agents‑A1 — a 35B Mixture-of-Experts agent trained for long-horizon, tool-enabled reasoning; supports 262K-context serving and runtimes like vLLM and SGLang.
Adapts pretrained Vision-Language-Action (VLA) models to new camera poses and robot embodiments from a single demonstration by performing weight-vector arithmetic that injects domain-specific information. Filters noise via subspace alignment of singular components; designed for one-shot adaptation under visual and embodiment shifts.
Accelerates text-to-image diffusion for pretrained flow-matching models using a staged low-to-high-resolution pipeline: fast low-res sampling, pixel-space GAN super-resolution, light latent noising, and short high-res refinement — >10× end-to-end speedups without retraining.
Provides a portable C++ inference runtime to deploy embodied AI models (vision–language–action and world–action) on heterogeneous robot hardware, enabling latency-first batch-1 closed-loop control. Key features include modular multi-rate layers, fused low-latency inference, and extensible head/IO plugins.
Provides a systematic benchmark and design roadmap for video-based world models to evaluate robot policies, introducing WMBench and GigaWorld-1 optimized for long-horizon, action-faithful rollouts. Offers controlled comparisons across model families, action encodings, and 324k+ simulated vs real rollouts, with code, models, and datasets released for reproducible evaluation.
Detects when an action-chunked VLA policy drifts from expected visual dynamics and triggers lightweight corrective replanning via a latent-space vision monitor and online gradient guidance; creates an event-driven adaptive action horizon without retraining the backbone.
Evaluates generalist robot manipulation policies across simulation and real-world settings using 42 sim tasks and 18 real tasks; measures generalization, memory, precision, long-horizon execution and open-vocabulary instruction following, and provides a cloud-accessible real-world evaluation system with XPolicyLab integration and a public leaderboard.
Provides re-annotated academic video instruction data for captioning, video QA, and fine-grained motion understanding; rewrites short answers and concise captions into evidence-grounded, instruction-following responses and supplies JSONL annotation files (original videos not included).
Fine-tuned variant of Qwen3.6-27B that cuts internal reasoning (‘thinking’) token usage by roughly 46% on average while preserving benchmark accuracy and safety behavior. Targets lower latency and inference cost; ships on Hugging Face with GGUF quantizations for local use.
Trains a single diffusion model that unifies 3D scene reconstruction and generative modeling by operating directly in pixel/rendered-image space. Supervises diffusion on rendered views and adds a geometry-perception loss from a pretrained 3D foundation model, reducing latent information loss and improving 3D fidelity.
Expresses diverse computer-vision tasks as instruction-driven text, image, or mixed generation from a single unified multimodal model, producing outputs for detection, segmentation, depth, pose, OCR and more. Trained on a converted SenseNova‑Vision instruction–response corpus and requires no task-specific prediction heads.
Reconstructs historical experience into latent memory tokens and weaves short- and long-term latent memories directly into vision-language-action reasoning to improve long-horizon robotic manipulation. Uses a four-part pipeline (curator, seeker, condenser, weaver) so memory participates natively in multimodal action formation.