Official inference framework for 1-bit and ternary (1.58-bit) LLMs such as BitNet b1.58, with optimized CPU kernels. Delivers 1.37x-6.17x speedups and 55-82% lower energy on x86 and ARM, and runs a 100B model on a single CPU at 5-7 tokens/sec.
Chains four swappable open modules — voice activity detection, speech-to-text, an LLM, and text-to-speech — into a local voice agent that needs no proprietary APIs. Runs on CUDA, Apple Silicon, or Docker, with an OpenAI-compatible realtime WebSocket mode.
Runs local LLM, vision-language, ASR, OCR, and image-generation models across NPU, GPU, and CPU from one command. Differs from Ollama and llama.cpp with first-class Qualcomm Hexagon NPU support and day-0 coverage of new models like Qwen3-VL.
Runs and optimizes ML and generative-AI models on-device across mobile, desktop, web, and IoT. Successor to TensorFlow Lite, it adds automated GPU/NPU accelerator selection and zero-copy buffer interop to cut latency without cloud round-trips.
Turns PDFs and images into clean Markdown with a 7B vision-language model, keeping tables, equations, handwriting, and multi-column reading order while removing headers and footers. Runs on one 12GB+ GPU at about 1/32 the cost of GPT-4o APIs.
Runs AI models on user devices with native SDKs, optimized model management, hardware acceleration, and OpenAI-compatible APIs for apps that need offline, private inference.
A compact domain-specific language for writing high-performance GPU/CPU kernels (GEMM, FlashAttention, sparse kernels) with Python-like syntax. It provides tiling/pipelining primitives, a TVM-based compiler and multiple backends (CUDA/CuTeDSL, NVRTC, WebGPU, Metal, Ascend) for operator-level performance work.
Provides low‑latency on‑device speech-to-text, intent recognition, and text-to-speech for building real‑time voice agents and interfaces. Streaming-optimized models, incremental caching, multilingual TTS/ASR and cross-platform bindings (Python, iOS, Android, Linux, Raspberry Pi) target live voice use cases where sub-200ms responsiveness matters.
High-resolution image and video generation codebase and models that run with far lower compute and memory than typical diffusion systems. Uses linear-attention DiT variants, aggressive latent compression, and inference-scaling to support text-to-image (up to 4K), fast one/few-step generation, and efficient video pipelines.
Predicts 3D structures of proteins, nucleic acids, and small-molecule complexes, the first fully open-source model to approach AlphaFold3 accuracy. Boltz-2 adds binding-affinity prediction that nears FEP simulation accuracy at ~1000x the speed.
Generates video from text or images via a DiT-based latent diffusion model: text-to-video, image-to-video, frame extension, and multi-keyframe conditioning in one model. A distilled 2B variant runs near real-time on one H100; 13B for higher quality.