AIAny
AI Model2026
Icon for item

Underdog Saluki 27B 1.0

A 2-bit quantized GGUF of Qwen3.8-27B that fits under 8 GB and runs on stock llama.cpp while preserving function/tool-calling behavior; includes an optional small vision add-on and is tuned for agent/tool workflows (Apache-2.0).

Introduction

Why this matters

Making a recent 27B-class model usable on commodity setups without sacrificing tool-calling is the core insight here. Underdog Saluki compresses Qwen3.8-27B into a ~7.89 GB GGUF so developers can run a model with Qwen-style conversational and function-calling behavior on standard llama.cpp stacks and deploy agent workflows on modest hardware.

Key Capabilities
  • Small footprint, llama.cpp compatible: the main GGUF is ~7.89 GB (text-only) so the model loads and runs in environments that cannot host the full 54 GB Qwen3.8-27B; the practical implication is easier local hosting and faster iteration.
  • Preserved function/tool calling: benchmarked on a 120-task function-calling suite, Saluki passes 88 tasks (vs 84 for the full Qwen3.8-27B), meaning tool/agent integrations remain reliable after aggressive quantization.
  • Optional vision add-on: a separate 0.6–0.9 GB mmproj file enables image inputs when needed, keeping the core text model minimal until multimodal capability is required.
  • Tuned for agents and parallel tool calls: retains strong performance on agent-style benchmarks and shows improved parallel tool-call throughput compared with the full-size model in the authors' tests.
Who it's for & trade-offs

Great fit if you need to run a Qwen3.8-compatible conversational/agent model on constrained hardware, want a drop-in GGUF for llama.cpp, or prioritize robust function-calling and tool integration over peak raw math/competition scores. The trade-offs: competition-math performance is reduced (roughly 82–85% of the original on some tasks), some letter-level instruction puzzles and a minority of parallel-call replies show small formatting slips, and the vision capability is an optional add-on rather than built into the main file. Use the full Qwen3.8-27B if absolute top-tier reasoning/math performance is essential.

Where it fits

Saluki sits between full-scale foundation models and ultra-small distilled models: it preserves many high-level behaviors of Qwen3.8 while enabling local deployment on machines that cannot host a 50+ GB weights file. It’s particularly useful for developers building local agents, tool-enabled chatbots, or workflows that require function calling without cloud dependencies.

Information

  • Websitehuggingface.co
  • OrganizationsConwayResearch (Underdog), Underdog AI, Qwen team, ISTA-DASLab
  • Published date2026/10/08

Categories

More Items

Hugging Face
AI Model2026

Post-trained multimodal Qwen3.8-27B variant that uses alternating SFT and RLOO to reduce pathological long reasoning tails; ships multiple quantization tiers (BF16, FP8, NVFP4, INT8, INT4, GGUF), supports MTP and DFlash2 speculative decoding, and includes detailed benchmark and runtime recommendations.

Hugging Face
AI Image2026

Runs text-to-image generation and instruction-guided image editing in 8 denoising steps. An accelerated checkpoint of Qwen-Image-2.1 that preserves the same 7B visual generator, native RGBA support, Diffusers QwenImage21Pipeline compatibility, a saved 8-step sampling schedule (CFG=1), and prefix KV cache reuse for multi-reference editing.

Hugging Face
AI Model2026

Performs a byte-level transplant of 144 tensors in an already-quantized GSQ-RCO Qwen3.8-Flash-Next to ablate the model's refusal direction while preserving GSQ-learned scales and the upstream per-tensor type assignment; multimodal, 262K context. Intended for local inference, red-teaming and quantization research; no retraining or built-in safety.