AIAny
AI Video2026
Icon for item

LingBot-Video-MoE (30B-A3B)

Generates videos from text and image+text prompts using a 30B Mixture-of-Experts model tuned for embodied intelligence; includes a refiner and structured prompt rewriter, and supports diffusers/SGLang runtimes with multi-GPU inference.

Introduction

Most video models focus on visual fidelity; this release focuses on aligning video synthesis with physical-world reasoning for embodied tasks. That shift matters because downstream robotics, simulation, and multi-step interaction tasks require temporally coherent videos that reflect plausible object dynamics and task completion, not just photorealism.

Key Capabilities
  • MoE scale: a 30B-parameter Mixture-of-Experts backbone plus a refiner, designed to increase capacity while keeping inference throughput tractable for large-video generation scenarios. This design targets higher long-horizon and multi-entity reasoning compared with dense counterparts.
  • Large embodied training signal: trained on a mixture that the authors report as including 70,000+ hours of web-sourced embodied video data, with multi-reward supervision for aesthetics, physical rationality, and task completion — meaning outputs are optimized for plausible interactions, not only appearance.
  • End-to-end inference workflows: provides a structured prompt rewriter (two-stage rewriter + LoRA adapter), auto-negative pruning, and ready scripts for single- and multi-GPU inference across diffusers and SGLang backends, plus FSDP/CP8 sharding options for MoE checkpoints.
  • Benchmarked leadership: reported top ranking on a public RBench leaderboard (as of early July 2026) on metrics combining manipulation, long-horizon reasoning, and embodied scenarios.
Who it's for — and tradeoffs

Great fit if you need research-grade video models that emphasize physical plausibility and task-oriented video generation (robotics simulators, embodied AI research, multimodal agents). The project is released under Apache-2.0 and bundles dense and MoE checkpoints plus rewriter components for production-style inference.

Look elsewhere if you need a tiny, single-GPU consumer model for quick mobile prototyping: MoE inference requires substantial system RAM and multi-GPU or specialized runtimes (SGLang/CP8/FSDP) for practical throughput. Expect engineering effort to set up the two-stage rewriter, LoRA adapters, and grouped-expert runtime for best performance.

Information

  • Websitehuggingface.co
  • AuthorsShuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang …
  • Published date2026/07/08

More Items

Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.

Hugging Face
AI Model2026

Performs a byte-level transplant of 144 tensors in an already-quantized GSQ-RCO Qwen3.8-Flash-Next to ablate the model's refusal direction while preserving GSQ-learned scales and the upstream per-tensor type assignment; multimodal, 262K context. Intended for local inference, red-teaming and quantization research; no retraining or built-in safety.

Hugging Face
AI Model2026

Maps multimodal inputs (text + images) to structured decisions (yes/no, choice, or scored rubric) in a single forward pass and returns calibrated probabilities. 3.1B parameters, long context (32,768 tokens), optimized for low-latency edge inference; not a text-generation/chat model.