A distilled 26M-parameter encoder–decoder LLM for on-device function-calling and tool use. Uses a pure-attention Simple Attention Network, provides open weights and local finetuning, and targets high-throughput inference on the Cactus runtime.
Compresses high-dimensional embeddings into low-bit TurboQuant indexes for fast, memory-efficient local vector search. Supports online ingest (no train/rebuild), SIMD kernels that match or beat FAISS, per-vector length-renormalization, and runtime allowlists — suited for privacy-sensitive, low-latency RAG.
A dense 128B multimodal model with a 256k context window, configurable reasoning effort, and native function-calling for agentic workflows. Supports text+image input, multilingual output, and is released on Hugging Face under a Modified MIT license with revenue-based exceptions.
Provides hardware-isolated, sub-60ms, ultra-low-overhead sandboxes to run untrusted LLM/agent code. Offers event-level snapshots, kernel-level egress control, credential vaulting, and drop-in E2B SDK compatibility for high-density AI agent deployment.
Provides a compact GGUF export of a tuned Gemma‑4 26B variant for local inference, optimized for llama.cpp and Apple Silicon to deliver faster, less‑censored chat and coding outputs. Includes Q4_K_M quantization and a neutral embedded template for more reliable local deployments.
Generates text by iteratively denoising blocks of tokens with a two-tower design: a frozen autoregressive context tower and a trainable diffusion denoiser tower, trading minimal quality loss for higher wall-clock throughput.
A Mixture-of-Experts instruct-capable LLM (295B total, 21B active) designed for long-context reasoning, code/agent workflows and instruction-following; released by Tencent Hy Team with safetensors weights on Hugging Face.
Drafts multiple tokens in parallel with a lightweight block-diffusion drafter to enable speculative decoding for faster LLM inference. Designed to pair with Qwen3.6-35B-A3B and reports up to ~2.9× throughput improvements on common benchmarks.
GGUF quantized files for a Qwen3.6-35B checkpoint fine-tuned with Claude Opus 4.6-style chain-of-thought distillation to improve reasoning. Offers multiple llama.cpp-compatible quant options (Q4/Q5/Q6/Q8) for local text-generation inference.
Fine-tuned Qwen3.6-35B-A3B MoE that reproduces Claude Opus 4.7-style chain-of-thought with explicit <think>…</think> blocks. Offers sparse activation (256 experts, ~3B active params), 64k context, and GGUF builds for local inference; best for long, multi-step reasoning but may emit very long reasoning traces.
Provides a GGUF-packaged, native-INT4 quantized build of the multimodal Kimi K2.6 model for image-text-to-text inference — packaged for local/self-hosted inference engines (vLLM, SGLang, KTransformers) to reduce footprint while keeping multimodal capabilities.
Produces 384‑dim multilingual (and code) embeddings with up to 32,768 token context, optimized for low‑latency production retrieval. Compact 97M model with ONNX/OpenVINO and vLLM/GGUF deployment options for edge and high‑throughput use.