Extends vLLM beyond text to serve omni-modal models — Qwen3-Omni, TTS like CosyVoice3, and diffusion image/video/audio generators — in one engine, adding the non-autoregressive Diffusion Transformer support the core project never targeted.
Automatically removes safety alignment from transformer LLMs via directional ablation, with Optuna's TPE optimizer tuning the parameters — no retraining or model-internals expertise needed; hit 3/100 refusals at 0.16 KL on Gemma-3-12b.
Lets AI agents describe interactive UIs as declarative JSON instead of executable code; client apps render the components with native widgets from a pre-approved catalog, keeping agent-generated UI safe across trust boundaries.
Provides a Gymnasium-style API and tooling to create, deploy, and interact with isolated execution environments for agentic RL training. Includes async/sync clients, a web interface, CLI, Docker-based deployment, and Hugging Face Spaces integration.
Worked examples and reusable abstractions for fine-tuning open LLMs via the Tinker training API: you write the training loop while distributed execution runs remotely. Covers SFT, math/code RL, DPO, three-stage RLHF, distillation, and tool use.
Turns documentation sites, GitHub repos, PDFs, videos and other sources into ready-to-use skill packs for Claude, Gemini, OpenAI and RAG frameworks like LangChain. Detects conflicts across sources, transcribes video, and exports to 21 formats.
Packages reusable GitHub Copilot building blocks — agents, prompts, instructions, and skills — to make AI-assisted coding repeatable and standards-aligned for a team. Built around an RPI (Research, Plan, Implement) workflow in VS Code.
Provides a frontend-design skill plus 20 steering commands and curated anti-patterns to steer LLMs toward clearer, accessible UI designs. Designed to plug into AI harnesses (Cursor, Claude/Gemini CLI, code agents) for auditing, critiquing, and polishing interfaces.
Wraps LangGraph in an opinionated harness giving an agent planning, file read/write, sub-agent delegation, and persistent memory out of the box. Aimed at long-horizon work where plain ReAct loops exhaust their context window.
Embeds into an app like SQLite, persisting to a local file with no server or separate process. Combines dense and sparse vectors, full-text search, and scalar filters in one hybrid query; C++ core with Python, Node, Go, Rust, and Dart bindings.
Normalizes wearable data — heart rate, sleep, activity, steps — from Garmin, Whoop, Apple Health and more behind one self-hosted API, so you write one integration instead of one per provider. Natural-language AI health automations are planned.
Enables parallel speculative decoding by using a lightweight block-diffusion draft model to produce multi-token drafts for faster, high-quality generation. Integrates with vLLM, SGLang and Transformers backends and ships draft models on Hugging Face.