AIAny
AI Model2026
Icon for item

Qwen3.8-Flash-Next · GSQ-RCO Coder

Capability-targeted compression of Qwen3.8-Flash-Next: half the experts are removed and remaining weights quantized to 3.5 bpw, producing a 58.4 GB GGUF (29.6 GB resident) that preserves coding and multimodal ability while trading off other domains.

Introduction

Why this matters

Large sparse MoE models concentrate most parameters in experts that rarely run on a single token. By pruning entire experts under a task-aware budget and combining that pruning with per-tensor low-bit scalar quantization, you can reduce resident memory dramatically while keeping the model useful for selected capabilities. This release shows that targeting the search to code and multimodal calibration data preserves coding and vision competence at high relative scores while cutting the working set to fit a single 32 GB accelerator.

What Sets It Apart
  • Directed expert pruning + GSQ quantization: RCO is used to choose which experts to remove under exact per-layer budgets, and GSQ produces accurate low-bit scalar quantizations. Together they yield a non-uniform, deployable GGUF that balances pruning and quantization rather than applying a uniform precision.
  • Measured, capability-focused tradeoff: the 176.9B-base model (354 GB BF16) becomes a two-shard GGUF of 58.4 GB total with 29.6 GB required resident (256 of 512 experts per layer). Reported effective rate is 1.89 bits/param (amortized); retained weights are stored at 3.5 bpw and the n-gram shard remains at higher precision.
  • Empirical outcomes: coding benchmarks largely preserved (LiveCodeBench ~98.7% of base; SWE-bench Verified ~91.3%), showing the method can keep targeted abilities while reducing size by >6x versus BF16.
How it works (concise)
  • RCO (Riemannian Constrained Optimization) enforces exact per-layer expert-count budgets and optimizes KL divergence between pruned and unpruned models on calibration data, yielding a joint selection across layers rather than independent heuristics.
  • GSQ (Gumbel-Softmax Quantization) produces accurate low-bit scalar quantization per tensor, letting the release store retained weights at ~3.5 bits per weight (deployable in standard GGUF formats).
  • The calibration mixture determines which experts are deemed important; omitting a capability (e.g., vision) from calibration will typically cause the search to prune the experts supporting it.
Who it's for / tradeoffs

Great fit if you need a locally runnable Qwen3.8-Flash-Next variant that prioritizes coding and multimodal outputs and must fit a single GPU with ~32 GB resident memory. It is also useful for researchers exploring expert-pruning, constrained budget allocation, and low-bit GGUF deployments.

Look elsewhere if you need a general-purpose, unpruned model for broad-domain accuracy — the pruning is intentional and directed, so capabilities outside the calibration targets can degrade. For general use, prefer the unpruned GSQ-RCO releases or the original BF16 weights.

Information

  • Websitehuggingface.co
  • OrganizationsDeep Algorithms and Systems Lab (DASLab), Institute of Science and Technology Austria (ISTA)
  • Published date2026/09/26

Categories

More Items

Hugging Face
AI Model2026

Provides compact mixed-precision GGUF quantizations of UkisAI's Swift 1.5 (derived from Qwen3.8-27B), using ISTA GSQ-RCO per-tensor allocations with Swift-specific refinement. Offers multiple 8–12 GB tiers, optional MTP heads, and KLD evaluation against the Swift BF16 baseline.

Hugging Face
AI Model2026

Provides a high-fidelity mixed-precision (≈3-bit) GGUF quant of Qwen3.8-27B tailored for long-horizon agent workloads and cyber-focused red-teaming. Preserves reasoning, thinking mode, MTP speculative decoding and vision (via a separate mmproj); released as an uncensored/abliterated research build under Apache-2.0.

Hugging Face
AI Model2026

Provides official pretrained VisionHOPE visual-backbone checkpoints for ImageNet classification, COCO object detection & instance segmentation, and ADE20K semantic segmentation. Includes hierarchical Tiny/Small/Base models with PyTorch-compatible downloadable weights.