AIAny
AI Model2026
Icon for item

Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2

Post-trained multimodal Qwen3.8-27B variant that uses alternating SFT and RLOO to reduce pathological long reasoning tails; ships multiple quantization tiers (BF16, FP8, NVFP4, INT8, INT4, GGUF), supports MTP and DFlash2 speculative decoding, and includes detailed benchmark and runtime recommendations.

Introduction

Why this matters

Large LLMs that can reason for long contexts sometimes fail to stop — they keep re-deriving answers until the context fills and the final answer is empty. This repository addresses that operational failure mode by post-training Qwen3.8-27B with alternating supervised fine-tuning (SFT) and reward-learning-from-LLM-output (RLOO), producing a multimodal 27B model family that reduces 94K truncations while preserving necessary long reasoning.

Key Capabilities
  • Targeted behavioral fine-tuning: alternating SFT + RLOO on a SimPO-informed base to penalize wrong or empty long trajectories while keeping long correct reasoning.
  • Multi-tier quantization and formats: official BF16 and static FP8 (Block128), NVFP4 W4A16/W4A4/W4A4-W8A8, INT8 W8A8, INT8 W8A16, INT4 W4A16, plus GGUF variants (Q2/Q3/Q4/Q6_K/Q8_0) and NInfer .ninfer builds for different runtimes and trade-offs.
  • Speculative decoding: two mutually exclusive options — DFlash2 (higher per-request speed: ~180 tok/s) and MTP (about 91 tok/s); no-speculation baseline ≈46 tok/s in the same environment.
  • Multimodal: text + image/video support; vision tower kept BF16 and provided where needed for runtimes that support it.
  • Benchmarks & measured outcomes: under the project protocol static FP8 achieves GPQA 178/198, MMLU 450/500, LCB 90/100 and substantially fewer 94K truncations vs. upstream Qwen3.8-27B.
Where it fits
  • Use it when you need a large multimodal Qwen-derived model with deployed-ready quantization tiers and careful behavior tuning to avoid runaway long outputs.
  • Choose BF16 or static FP8 for highest evaluated accuracy and parity; use NVFP4 / INT8 / INT4 / GGUF variants to trade model size and runtime throughput for small accuracy or behavior differences.
  • Use DFlash2 when speculative-draft support and maximum per-request throughput are available (SGLang / NInfer workflows tested); use MTP when you prefer an integrated BF16 MTP head and cannot load the DFlash2 draft.
Who it’s for and tradeoffs

Great fit if: you operate large-context reasoning workloads (100K contexts), need multimodal capabilities, and want multiple quantization/packaging options for deployment (SGLang, vLLM, NInfer, llama.cpp). The repo supplies runnable packages, speed/throughput recommendations, and evaluation artifacts.

Look elsewhere if: you require a turnkey cloud API service (this is a model release + launch recipes, not a hosted API), or you cannot accept occasional small accuracy drops on some quantized NVFP4 tiers. Known trade-offs: NVFP4 variants show slightly lower GPQA/MMLU in some runs; long-tail very-high-length cases (≥48K reasoning) remain challenging; certain formats are runtime-specific (INT8/INT4 validated on vLLM; .ninfer only loadable by official NInfer; GGUF targeted at llama.cpp/NInfer-all).

Practical notes
  • Runtime compatibility: SGLang was used to validate FP8 + DFlash2; vLLM 0.28 validated INT8/INT4 variants (with some kernel flags required); NInfer loads .ninfer packages with built-in MTP/DFlash2 and vision; GGUF packages are for llama.cpp and converted NInfer bundles exist.
  • Packaging & license: published tiers include BF16, FP8, NVFP4, INT8, INT4, and multiple GGUF variants; Apache-2.0 license.

In short: this release is a deployment-focused, behavior-corrected Qwen3.8-27B family that provides measured trade-offs (accuracy, truncation rates, throughput) and multiple runtime options to help teams run large-context multimodal reasoning reliably.

More Items

Hugging Face
AI Model2026

A 2-bit quantized GGUF of Qwen3.8-27B that fits under 8 GB and runs on stock llama.cpp while preserving function/tool-calling behavior; includes an optional small vision add-on and is tuned for agent/tool workflows (Apache-2.0).

Hugging Face
AI Image2026

Runs text-to-image generation and instruction-guided image editing in 8 denoising steps. An accelerated checkpoint of Qwen-Image-2.1 that preserves the same 7B visual generator, native RGBA support, Diffusers QwenImage21Pipeline compatibility, a saved 8-step sampling schedule (CFG=1), and prefix KV cache reuse for multi-reference editing.

Hugging Face
AI Model2026

Performs a byte-level transplant of 144 tensors in an already-quantized GSQ-RCO Qwen3.8-Flash-Next to ablate the model's refusal direction while preserving GSQ-learned scales and the upstream per-tensor type assignment; multimodal, 262K context. Intended for local inference, red-teaming and quantization research; no retraining or built-in safety.