Why this matters
Large sparse MoE models concentrate most parameters in experts that rarely run on a single token. By pruning entire experts under a task-aware budget and combining that pruning with per-tensor low-bit scalar quantization, you can reduce resident memory dramatically while keeping the model useful for selected capabilities. This release shows that targeting the search to code and multimodal calibration data preserves coding and vision competence at high relative scores while cutting the working set to fit a single 32 GB accelerator.
What Sets It Apart
- Directed expert pruning + GSQ quantization: RCO is used to choose which experts to remove under exact per-layer budgets, and GSQ produces accurate low-bit scalar quantizations. Together they yield a non-uniform, deployable GGUF that balances pruning and quantization rather than applying a uniform precision.
- Measured, capability-focused tradeoff: the 176.9B-base model (354 GB BF16) becomes a two-shard GGUF of 58.4 GB total with 29.6 GB required resident (256 of 512 experts per layer). Reported effective rate is 1.89 bits/param (amortized); retained weights are stored at 3.5 bpw and the n-gram shard remains at higher precision.
- Empirical outcomes: coding benchmarks largely preserved (LiveCodeBench ~98.7% of base; SWE-bench Verified ~91.3%), showing the method can keep targeted abilities while reducing size by >6x versus BF16.
How it works (concise)
- RCO (Riemannian Constrained Optimization) enforces exact per-layer expert-count budgets and optimizes KL divergence between pruned and unpruned models on calibration data, yielding a joint selection across layers rather than independent heuristics.
- GSQ (Gumbel-Softmax Quantization) produces accurate low-bit scalar quantization per tensor, letting the release store retained weights at ~3.5 bits per weight (deployable in standard GGUF formats).
- The calibration mixture determines which experts are deemed important; omitting a capability (e.g., vision) from calibration will typically cause the search to prune the experts supporting it.
Who it's for / tradeoffs
Great fit if you need a locally runnable Qwen3.8-Flash-Next variant that prioritizes coding and multimodal outputs and must fit a single GPU with ~32 GB resident memory. It is also useful for researchers exploring expert-pruning, constrained budget allocation, and low-bit GGUF deployments.
Look elsewhere if you need a general-purpose, unpruned model for broad-domain accuracy — the pruning is intentional and directed, so capabilities outside the calibration targets can degrade. For general use, prefer the unpruned GSQ-RCO releases or the original BF16 weights.