Why this matters
Making a recent 27B-class model usable on commodity setups without sacrificing tool-calling is the core insight here. Underdog Saluki compresses Qwen3.8-27B into a ~7.89 GB GGUF so developers can run a model with Qwen-style conversational and function-calling behavior on standard llama.cpp stacks and deploy agent workflows on modest hardware.
Key Capabilities
- Small footprint, llama.cpp compatible: the main GGUF is ~7.89 GB (text-only) so the model loads and runs in environments that cannot host the full 54 GB Qwen3.8-27B; the practical implication is easier local hosting and faster iteration.
- Preserved function/tool calling: benchmarked on a 120-task function-calling suite, Saluki passes 88 tasks (vs 84 for the full Qwen3.8-27B), meaning tool/agent integrations remain reliable after aggressive quantization.
- Optional vision add-on: a separate 0.6–0.9 GB mmproj file enables image inputs when needed, keeping the core text model minimal until multimodal capability is required.
- Tuned for agents and parallel tool calls: retains strong performance on agent-style benchmarks and shows improved parallel tool-call throughput compared with the full-size model in the authors' tests.
Who it's for & trade-offs
Great fit if you need to run a Qwen3.8-compatible conversational/agent model on constrained hardware, want a drop-in GGUF for llama.cpp, or prioritize robust function-calling and tool integration over peak raw math/competition scores. The trade-offs: competition-math performance is reduced (roughly 82–85% of the original on some tasks), some letter-level instruction puzzles and a minority of parallel-call replies show small formatting slips, and the vision capability is an optional add-on rather than built into the main file. Use the full Qwen3.8-27B if absolute top-tier reasoning/math performance is essential.
Where it fits
Saluki sits between full-scale foundation models and ultra-small distilled models: it preserves many high-level behaviors of Qwen3.8 while enabling local deployment on machines that cannot host a 50+ GB weights file. It’s particularly useful for developers building local agents, tool-enabled chatbots, or workflows that require function calling without cloud dependencies.