Fine-tuned variant of Qwen3.6-27B that cuts internal reasoning (‘thinking’) token usage by roughly 46% on average while preserving benchmark accuracy and safety behavior. Targets lower latency and inference cost; ships on Hugging Face with GGUF quantizations for local use.
Deployment-optimized hybrid MoE LLM (75B total / 9.3B active) produced via Iterative Puzzle compression and Multi-Token Prediction to double server throughput and raise single-GPU concurrency; designed for multilingual reasoning, long-context generation, and high-volume agentic/chat deployments.
GGUF-format quantized release of DeepSeek‑V4‑Flash for local inference — compatible with llama.cpp and Unsloth runtimes, with guidance for FP4/FP8 mixed precision and Q4/Q8 quantization; tuned for million-token long-context usage.
Provides pre-converted colibrì-format int4 weights so GLM-5.2 (744B MoE) can run by streaming routed experts from disk on a consumer machine with ~25 GB RAM. Includes MTP shard for lossless speculative decoding; requires the colibrì engine and ~400 GB NVMe.
27B multimodal LLM post-trained to prioritize agentic, weight-scaled reasoning over 64K-token contexts. Built on Qwen3.6-27B and released with BF16 weights plus several GGUF quants; aimed at coding, long-document reasoning, tool use and multimodal inspection.
Generates videos from text and image+text prompts using a 30B Mixture-of-Experts model tuned for embodied intelligence; includes a refiner and structured prompt rewriter, and supports diffusers/SGLang runtimes with multi-GPU inference.
A GGUF-distributed Qwen3.6 35B MoE model variant repaired with a
GGUF conversions of Laguna S 2.1 for llama.cpp, including quantized builds (Q4_K_M, Q8_0, F16) and a small DFlash drafter for speculative decoding; configured for a 256K default context window and intended for local inference and serving with Poolside's llama.cpp fork.
End-to-end 0.8B multimodal OCR and page-level document parser that converts page images into structured Markdown (text, LaTeX formulas, HTML tables, and image crops). Post-trained from Qwen3.5-0.8B using mixed real/synthetic data and SFT+RL+OPD; achieves 96.58 on OmniDocBench v1.6.
A GGUF-local variant of Qwen3.6-35B that applies a non-training 'Genesis' tensor-repair process and Hermes-agent fine-tuning to enable uncensored, multimodal (text+image) local inference. Highlights: MoE 35B spec, large native context, Hermes function-calling dataset transfer, and recommended quantization/runtime settings.
Provides low-bit quantized Hy3 (hy_v3) GGUF model weights and mixed-precision quantization recipes for running Hy3 on llama.cpp, with optional MTP self-speculative decoding and imatrix-based calibration for improved quality/speed trade-offs.
Agentic coding and long-horizon text generation via a 118B-parameter Mixture-of-Experts LLM with a 1,048,576-token context window. Features 256 routed experts, native preserved-thinking (reasoning) control, speculative decoding draft models, and quantized checkpoints for lower-cost serving.