Enables closed-loop execution for embodied agents by evolving code-based runtime critics and recovery skills online while keeping the base policy frozen. Combines three timescale loops with Z-Infra rollout infrastructure; reports 90.8% on LIBERO-Pro, 93.6% on RoboCasa and an 11.1× inference speedup.
Estimates optimal learning rates for large-scale Mixture-of-Experts pretraining using a two-step, compute-efficient transfer: μP-based width transfer from small proxy models, then log-log linear extrapolation across token budgets to trillion-token horizons.
Explores a practical mechanism for recursive self-improvement by post-training LLMs: uses a routing harness to record agent executions and convert traces into curriculum-guided supervised fine-tuning and on-policy distillation data, closing an evaluation-selection-update loop and improving benchmark performance.
Serves token-level routed LLM inference by dispatching requests to per-model asynchronous subservers and using delayed-batching scheduling to reduce admission latency and step desynchronization. Exposes a request-centric route-send-receive API and reports 2.01–64.15× decoding throughput gains versus single-LLM servers.