Provides unquantized BF16 weights of Qwen3.6-27B with the base model's MTP head grafted in for high-fidelity, uncensored text (and multimodal) generation. Includes deployment guidance and hardware-tuned variants for A100/H100 and Blackwell-class GPUs.
Contains full chain-of-thought traces and final answers generated by DeepSeek-V4-Pro for use as distillation supervision. Key features: full CoT exposure, ~1,000 mixed-domain samples (JSONL/Parquet), Apache-2.0 license — suitable for training student models but watch for source contamination.
Provides multiple GGUF-quantized exports of Carnice V2 (a merged BF16 SFT of Qwen3.6-27B) optimized for llama.cpp and Hermes-style agent traces, with quant tiers targeted at 16–24GB local GPUs and agentic inference.
Cleaned dataset of reasoning-distillation examples derived from Claude Opus 4.7 outputs — 4,807 retained JSON chat rows after removing simulated-thinking, duplicates, and missing fields. Packaged for model distillation and reasoning evaluation; Apache-2.0 packaging with upstream Anthropic usage constraints.
GGUF-format, DS4-optimized quantized weights for DeepSeek-V4-Flash, offering q2 (≈80.8 GiB) and q4 (≈153.3 GiB) variants plus an optional small MTP file for speculative decoding. Built for the DS4 inference engine; MIT-licensed.
Open-source Mixture-of-Experts LLM designed for extremely long-context (up to 1M tokens) text generation and agentic workflows; uses a hybrid attention + MTP design to reduce KV-cache footprint while enabling 42B active parameters and FP8 mixed-precision training.
Unified omnimodal foundation model for text, image, video and audio understanding and agentic workflows, with support for up to 1M-token context. Combines a sparse MoE LLM backbone, dedicated vision/audio encoders, multi-token prediction, and a hybrid sliding-window + global attention design to reduce KV-cache overhead.
Provides a GGUF-quantized build of NVIDIA's Nemotron 3 Nano Omni 30B (Reasoning) for local inference — enables multimodal (video/audio/image/text) reasoning, transcription, and document understanding on compatible runtimes such as llama.cpp, Ollama, vLLM, and TensorRT-LLM.
Provides step-by-step guides to integrate DeepSeek V4 models (deepseek-v4-pro and deepseek-v4-flash) into 22 popular AI agents and coding-assistant tools. Each entry shows installation, configuration, and first-run steps for tools like Claude Code, Qwen Code, Codex, Cline, Deep Code, and more.
Multilingual on-device translation model compressed to 1.25-bit via the Sherry quantization, supporting 33 languages and 1,056 directions in a 440MB package for offline mobile translation and demos.
Provides reusable Docker compose files, scripts and benchmarked configs to serve modern LLMs (Qwen, Gemma, etc.) on 1–2 NVIDIA RTX 3090/4090/5090 GPUs. Multi-engine (vLLM, llama.cpp, ik_llama), measured TPS/context tradeoffs, and validated single/dual‑GPU recipes.
An instruct-focused LLM (104B total, 7.4B active) optimized for fast, token-efficient inference in agent workflows. Uses hybrid linear attention plus a sparse MoE to raise throughput and cut token use; suited for high-frequency production agents, with some trade-offs in very deep reasoning.