Orchestrates low-latency, multi-stage pipelines for omni and multimodal models by running each stage with its own scheduler and using zero-copy shared memory for tensor transfer. Emphasizes per-stage bottleneck tuning and OpenAI-compatible streaming endpoints, suitable for TTS and multimodal serving.
Converts DeepSeek protocol calls into OpenAI/Claude/Gemini-compatible APIs with a Go backend and React admin UI. Offers account pooling, protocol adapters, tool-call translation, PoW, and multiple deployment modes (Docker, Vercel, standalone).
Management system for Xianyu sellers that handles multi-account operations, context-aware AI auto-replies, automated shipping/confirmation, and a web admin UI. Built with FastAPI, Playwright and Docker; intended for learning/research only, not for commercial use.
Runs background coding agents in isolated sandboxes to autonomously handle development tasks, create pull requests, and integrate with Slack, GitHub, Linear and webhooks. Supports multiplayer sessions, multiple LLM providers, fast startup via snapshots and prebuilt images; designed for single-tenant deployments.
Provides a REST server that lets AI agents browse sites while avoiding common bot-detection by running Camoufox (a Firefox fork with C++-level fingerprint spoofing). Returns compact accessibility snapshots, stable element refs, session isolation, proxy/geoIP support, and agent-friendly endpoints (click, type, snapshot, transcripts).
Provides a systematic, project-driven tutorial and runnable codebase for building AI agents, RAG pipelines, and multi-agent systems—focused on Python, LangChain/LangGraph, tooling, deployment, and an interview question bank for engineers aiming to ship production agent applications.
Unified API proxy and protocol gateway that translates and routes requests to Claude, OpenAI Chat/Images/Codex, and Gemini. Offers channel orchestration, multi-key rotation, failover, model routing, and a built-in web admin UI for consolidating multiple model providers behind a single endpoint.
Self-hosted coding assistant that runs frozen local LLMs with constraint-driven planning, energy-based verification, and self-verified repair to produce verified code. Emphasizes offline inference (no cloud), Docker/bare-metal deployment, and requires a 16GB+ GPU.
Self-hosted personal AI agent runtime that runs chats, tools, automations and long-term memory for persistent workflows. Small, readable core with a bundled WebUI, multi-chat integrations, an OpenAI-compatible API and a Python SDK for easy extension and deployment.
Equips AI coding agents with reusable AWS skills (deployment, serverless, Amplify, SageMaker) by packaging agent skills, MCP servers, hooks, and references so agents invoke vetted workflows instead of bloating prompts.
A TypeScript framework for building programmable, headless autonomous agents with a harness-centric runtime. Includes an SDK and CLI, virtual sandboxes (just-bash) with optional full container sandboxes, provider-agnostic model settings, and connectors for CI/Daytona/MCP—suited for deployable agent runtimes.
Generates high‑fidelity, expressive speech and environmental sounds from text. The MOSS‑TTS Family provides specialized models for long‑form TTS, multi‑speaker dialogue, voice design and realtime streaming, plus torch‑free inference paths (llama.cpp / ONNX) and Hugging Face releases.