Maps multimodal inputs (text + images) to structured decisions (yes/no, choice, or scored rubric) in a single forward pass and returns calibrated probabilities. 3.1B parameters, long context (32,768 tokens), optimized for low-latency edge inference; not a text-generation/chat model.
Runs a pruned, NVFP4-quantized GLM-5.3-Flash variant tuned for Blackwell GPUs: 224 routed experts per layer and ~141 GiB of weights. Retains the multimodal vision tower, activates 18B params/token, supports vLLM and optional MTP speculative decoding; fits 2× DGX Spark or a ≥180 GB B200.
Generates synchronized egocentric video streams for multiple interacting agents in a shared environment, enforcing cross-view action consistency, shared environment memory, and consistent propagation of interaction-induced state changes — aimed at embodied AI, VR, and multi-agent vision research.
Lets a pretrained multimodal LLM interpret navigation requests and orchestrate motion via tool calls for generalist robot navigation across unfamiliar scenes. Key features: an agent harness with Navigation Skills, a unified visual-point interface, task-progress tracking, and tool-based motion execution without navigation-specific fine-tuning.
Uses large-scale, mixed-task reinforcement learning to drive self-improvement of multimodal foundation models. Key features include fully-asynchronous large-batch RL (1,568 prompts × 16 rollouts, up to 3.7B tokens/update, 1M context), groupwise agentic grading, and sparse MoE architectures.