Why this matters
Large LLMs that can reason for long contexts sometimes fail to stop — they keep re-deriving answers until the context fills and the final answer is empty. This repository addresses that operational failure mode by post-training Qwen3.8-27B with alternating supervised fine-tuning (SFT) and reward-learning-from-LLM-output (RLOO), producing a multimodal 27B model family that reduces 94K truncations while preserving necessary long reasoning.
Key Capabilities
- Targeted behavioral fine-tuning: alternating SFT + RLOO on a SimPO-informed base to penalize wrong or empty long trajectories while keeping long correct reasoning.
- Multi-tier quantization and formats: official BF16 and static FP8 (Block128), NVFP4 W4A16/W4A4/W4A4-W8A8, INT8 W8A8, INT8 W8A16, INT4 W4A16, plus GGUF variants (Q2/Q3/Q4/Q6_K/Q8_0) and NInfer
.ninferbuilds for different runtimes and trade-offs. - Speculative decoding: two mutually exclusive options — DFlash2 (higher per-request speed: ~180 tok/s) and MTP (about 91 tok/s); no-speculation baseline ≈46 tok/s in the same environment.
- Multimodal: text + image/video support; vision tower kept BF16 and provided where needed for runtimes that support it.
- Benchmarks & measured outcomes: under the project protocol static FP8 achieves GPQA 178/198, MMLU 450/500, LCB 90/100 and substantially fewer 94K truncations vs. upstream Qwen3.8-27B.
Where it fits
- Use it when you need a large multimodal Qwen-derived model with deployed-ready quantization tiers and careful behavior tuning to avoid runaway long outputs.
- Choose BF16 or static FP8 for highest evaluated accuracy and parity; use NVFP4 / INT8 / INT4 / GGUF variants to trade model size and runtime throughput for small accuracy or behavior differences.
- Use DFlash2 when speculative-draft support and maximum per-request throughput are available (SGLang / NInfer workflows tested); use MTP when you prefer an integrated BF16 MTP head and cannot load the DFlash2 draft.
Who it’s for and tradeoffs
Great fit if: you operate large-context reasoning workloads (100K contexts), need multimodal capabilities, and want multiple quantization/packaging options for deployment (SGLang, vLLM, NInfer, llama.cpp). The repo supplies runnable packages, speed/throughput recommendations, and evaluation artifacts.
Look elsewhere if: you require a turnkey cloud API service (this is a model release + launch recipes, not a hosted API), or you cannot accept occasional small accuracy drops on some quantized NVFP4 tiers. Known trade-offs: NVFP4 variants show slightly lower GPQA/MMLU in some runs; long-tail very-high-length cases (≥48K reasoning) remain challenging; certain formats are runtime-specific (INT8/INT4 validated on vLLM; .ninfer only loadable by official NInfer; GGUF targeted at llama.cpp/NInfer-all).
Practical notes
- Runtime compatibility: SGLang was used to validate FP8 + DFlash2; vLLM 0.28 validated INT8/INT4 variants (with some kernel flags required); NInfer loads
.ninferpackages with built-in MTP/DFlash2 and vision; GGUF packages are for llama.cpp and converted NInfer bundles exist. - Packaging & license: published tiers include BF16, FP8, NVFP4, INT8, INT4, and multiple GGUF variants; Apache-2.0 license.
In short: this release is a deployment-focused, behavior-corrected Qwen3.8-27B family that provides measured trade-offs (accuracy, truncation rates, throughput) and multiple runtime options to help teams run large-context multimodal reasoning reliably.