Applies quantization, pruning, distillation, speculative decoding and sparsity to compress and accelerate deep learning models, producing exportable checkpoints ready for deployment on TensorRT‑LLM, TensorRT, vLLM and similar inference runtimes.
Accelerates video generation with a unified framework for inference, finetuning, LoRA, distillation, sparse attention, and distributed execution for research and demos.
Runs open LLMs entirely on your own machine — discover and download models from Hugging Face, chat in a desktop GUI, or expose an OpenAI-compatible local server. Native Apple MLX and llama.cpp backends; headless deploy via llmster.
Segments each PDF page into 11 labeled regions — titles, tables, formulas, figures, footnotes and more — and recovers reading order. Offers two engines: an accurate VGT visual model (~0.96 F1) or a faster CPU-only LightGBM ensemble.
Streamlines the full lifecycle of foundation models — data prep, fine-tuning (SFT/LoRA/QLoRA/GRPO), evaluation, and deployment — with ready-to-run recipes, multi-engine inference support, and cloud/CLI workflows for both laptop experiments and large-scale runs.
Runs reproducible evaluations of large language models through a Python API with built-in solvers, scorers, and model-graded grading. Ships 200+ ready-to-run evals spanning capability and safety testing, and connects to most major model providers.
Stores and reuses LLM key-value caches across GPU, CPU, disk, and remote backends so vLLM and SGLang skip recomputing repeated context. Non-prefix reuse (CacheBlend) and PD disaggregation cut time-to-first-token for long-context and RAG serving.
Cloud-native control plane that scales vLLM on Kubernetes, adding the routing, autoscaling, and fault tolerance single-instance serving lacks. Brings high-density LoRA management, an LLM gateway, distributed KV cache reuse, and SLO-aware GPU serving.
Connects multiple Macs and Linux machines into one cluster to run models too large for any single machine. Auto-discovers peers, shards a model across them via tensor parallelism, and exposes OpenAI-, Claude-, and Ollama-compatible APIs.
Disaggregated LLM serving architecture that splits prefill and decode into separate clusters and pools spare CPU, DRAM, and SSD into a distributed KVCache. Powers Kimi in production, handling 75% more requests under the same SLOs.
Runs huge mixture-of-experts LLMs like DeepSeek-R1/V3 on a single 24GB GPU plus CPU DRAM by keeping attention on the GPU and offloading expert weights to CPU. Reports 3-28x speedups via Intel AMX/AVX512 kernels and fits 139K context in 24GB VRAM.
Generates images from text prompts using a 12-billion-parameter rectified-flow transformer trained with guidance distillation for more efficient sampling. Distributed with diffusers/ComfyUI support and multiple conditioning/editing variants; weights released under a non-commercial license.