Centralized enterprise platform to manage org-wide MCP servers with a private MCP registry, security guardrails, cost controls, and observability. Offers a Kubernetes-native orchestrator, built-in RAG knowledge base, security sub-agents, and tools for governed AI adoption.
Composes AI agent teams from a Ghost+Shell+Model formula: each Bot pairs a prompt/MCP/Skills Ghost with a Chat, ClaudeCode, or Dify shell and a model like Claude or DeepSeek. Bots form Teams that run as traceable Tasks, wired to GitHub and DingTalk.
Extends vLLM beyond text to serve omni-modal models — Qwen3-Omni, TTS like CosyVoice3, and diffusion image/video/audio generators — in one engine, adding the non-autoregressive Diffusion Transformer support the core project never targeted.
Turns papers, repositories, or natural-language research goals into local-first executable 'quests' that automate reproducible experiments, branching, and result-to-paper workflows. Preserves experiment history, supports web/TUI/connectors, and keeps human takeover and inspection simple.
Drives an LLM-powered agent to autonomously research, write, and ship ML code by accessing Hugging Face docs, datasets, repos, and cloud compute. Provides interactive CLI and headless modes, approval gates, tool routing, and integrations for HF, GitHub, and Anthropic models.
Equips AI coding assistants like Claude Code and Cursor with 75+ executable tools, an MCP server, reusable skills, and a Python library to build on Databricks—Spark pipelines, jobs, dashboards, Unity Catalog resources, and ML workflows—from your editor.
Generates daily LLM-powered decision dashboards for A/H/US stocks by combining multi-source market data, real-time news, technical signals and agent-style strategy reasoning; deploys via GitHub Actions or Docker and pushes reports to multiple channels.
Provides a framework to build, evaluate, and run AI SRE agents that investigate and remediate production incidents on your infrastructure. Includes a CLI, synthetic + end-to-end benchmark suites, and 40+ connectors for observability, infra, and LLM providers so teams can train agents and run investigations locally or in cloud.
Turns a PC, Mac, or Linux machine into a private AI server with one-command installers: local LLM inference, a ChatGPT-style web UI, voice, agents, RAG, workflows, image generation, hardware-aware model selection, and optional cloud/hybrid modes.
Local LLM inference server for Apple Silicon that exposes an OpenAI-compatible API and a macOS menubar app. Uses continuous batching and a two-tier KV cache (RAM + SSD in safetensors) to persist context across restarts, enabling practical multi-model serving and fast local coding workflows.
Acts as an OpenAI‑compatible local and cloud gateway that routes requests across 100+ LLM providers with smart routing, load balancing, retries and fallbacks. Adds policies, rate limits, semantic caching and observability for reliable, cost‑aware inference in Docker, Electron or npm installs.