Terminal-native AI coding agent that runs as a single static Go binary and preserves long LLM sessions using DeepSeek-aware prefix caching. Config- and plugin-driven: supports multiple providers, separate planner/executor sessions, and CLI/TUI, desktop and VS Code integrations.
Provides a single OpenAI-compatible /v1 API that aggregates the free tiers of 16 LLM providers into one unified endpoint. Features smart routing and automatic failover, per-key free-tier tracking, encrypted key storage, embeddings/media routing, and a Docker one-liner for local use.
A 284B-parameter Mixture-of-Experts LLM with only 13B activated parameters, designed for 1,000,000-token contexts. Uses hybrid compressed attention and mixed FP4/FP8 precision to reduce long-context KV-cache and per-token FLOPs; aimed at long-document QA, RAG pipelines, and local/high-capacity inference.
Provides a locally runnable, refusal-free variant of Qwen3.6-27B with multiple K_P GGUF quantizations and mmproj multimodal support. The Aggressive flavor skips preambles on edgy prompts—use when you want direct/raw responses for local research, red‑teaming, or offline workflows.
Generates conversational and reasoning outputs with support for million‑token contexts; uses a hybrid attention + MoE design to cut long‑context inference FLOPs and KV cache. Suited for long‑document retrieval, coding and complex reasoning; MIT licensed.
Unifies multimodal image understanding, text-to-image generation, and instruction-based editing in a single diffusion LLM using a Mixture-of-Experts backbone, SigLIP-VQ discrete tokenizer, and a distilled diffusion decoder enabling fast (8-step) decoding; full-generation needs ~47GB GPU RAM.
A 14B dense tri‑mode language model that supports autoregressive, diffusion‑based parallel decoding, and self‑speculation—designed to increase token throughput and acceptance length; best suited for researchers and engineers exploring decode‑efficiency tradeoffs on NVIDIA hardware under the Nemotron Open Model License.
Provides 150,000 synthetic Vietnamese patient personas to condition clinical text generation. Each persona bundles demographics, socioeconomic context, health and behavior fields, and prompt-ready narratives; intended for research and simulation, not clinical decision-making.
Drop-in Jinja chat templates for Qwen 3.5/3.6 that fix rendering errors, token waste, and tool-calling failures across runtimes (LM Studio, llama.cpp, vLLM, MLX). Adds a think-on/think-off toggle, auto-closes broken thinking tags, robust tool-argument handling, and a graceful fallback for missing user queries.
Provides an NVFP4-quantized 27B Qwen3.6 checkpoint optimized for faster, low-memory multimodal inference on 24GB GPUs. Includes MTP (multi-token prediction), extended 262k native context, and deployment recipes for vLLM/SGLang/KTransformers; best used with recommended backends for peak throughput.
A Qwen-3.6 27B model variant optimized for DFlash (speculative decoding) to reduce generation latency and increase throughput. Focuses on faster inference on serving stacks and is suitable for text-generation endpoints where lower latency and resource efficiency matter.
Provides a 30K+ problem multimodal, multilingual dataset of Olympiad-level math problems with expert solutions and a math-aware retrieval benchmark—includes images, hierarchical topics, provenance from official booklets, and LLM-assisted metadata (v0, CC BY 4.0).