Serves large language and multimodal models with low latency and high throughput using RadixAttention, continuous batching, structured outputs, parallelism, quantization, and broad accelerator support.
Performs document OCR, layout analysis, reading-order detection and table recognition across 90+ languages using a ~650M-parameter vision–language model; offers per-page and per-block modes and supports GPU (vllm) and CPU/Apple Silicon backends.
Clones a voice from a 5-second sample for zero-shot TTS, or fine-tunes on ~1 minute of audio for few-shot synthesis. Covers Chinese, English, Japanese, Korean, and Cantonese, with a WebUI bundling vocal separation, ASR, and dataset labeling.
Reworks AUTOMATIC1111's Stable Diffusion WebUI onto a custom backend that auto-manages GPU memory to speed inference and cut VRAM use. Adds native FLUX support with NF4/GGUF quantization and a UNetPatcher framework for model-agnostic extensions.
GPU-accelerated code editor written in Rust, organized like a game engine to render its UI via shaders. Includes native agentic coding over the open Agent Client Protocol, multiplayer editing, LSP/DAP, and an open edit-prediction model.
Runs GPT-4o-class vision, speech, and full-duplex audio-video conversation on a 9B model small enough to deploy on phones and tablets. The 4.5 release scores 77.6 on OpenCompass and adds real-time bilingual voice with voice cloning.
GPU kernel library for LLM inference attention, sampling, and KV-cache, built on block-sparse formats with JIT-compiled customizable templates. Reports 29-69% inter-token-latency cuts vs compiler backends; powers SGLang, vLLM, and MLC-Engine.
Runs AI-generated code in isolated, elastic sandboxes with SDK, API, and CLI access for agent workflows that need stateful execution and environment control.
A minimal GPU written in under 15 SystemVerilog files to teach how GPUs execute parallel kernels from the ground up. Includes an 11-instruction ISA, multiple cores with ALUs and load-store units, a fetch-decode-execute pipeline, and matrix kernels.
Runs a privacy-first, self-hosted answering engine that combines web retrieval with local and cloud LLMs to produce cited answers. Supports SearxNG search, file uploads, image/video search, and mix-and-match models with Speed/Balanced/Quality modes.
Provides local inference, fine-tuning, and a server/CLI for vision–language and omni (image/audio/video) models via MLX. Supports multi-image chat, audio/video inputs, activation quantization (CUDA), TurboQuant KV cache, and LoRA/QLoRA fine-tuning for on-device workflows.
Chains pre-trained AI weather and climate models like GraphCast, Pangu, and FourCastNet into composable inference pipelines. Swap prognostic or diagnostic components, plug in reanalysis sources, and add ensemble perturbations or in-loop metrics.