Lets vision-language models control robots via a compact semantic action interface that maps intent to discrete action units; supports zero-shot use of closed-source VLMs, low-cost fine-tuning of open VLMs, and GUI-based demonstration collection.
Transforms scientific code repositories into executable, agent-learnable environments that support task generation, execution, and scientific verification. Agent-guided repository transformation produces verified interaction trajectories used to train the PhAI-IDE model family. Intended for research on agent learning, scientific-code repair, and training RL/SFT models.