Provides a high-fidelity mixed-precision (≈3-bit) GGUF quant of Qwen3.8-27B tailored for long-horizon agent workloads and cyber-focused red-teaming. Preserves reasoning, thinking mode, MTP speculative decoding and vision (via a separate mmproj); released as an uncensored/abliterated research build under Apache-2.0.
Defines and evaluates AREX-2, an LLM agent that iteratively self-improves at test time via reflection and long-horizon execution. Trained on long-horizon improvement trajectories from ML engineering and algorithmic programming (built on Qwen3.8-27B), it scales with more rounds and achieves strong benchmark scores.
Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.
Analyzes how proposer–solver loops in self-evolving search agents can develop shared errors (co-cheating) that inflate internal rewards; introduces Multi-Sample Verification and CrossFit (cross-fitted scoring with partitioned sources) to reduce false agreement and improve downstream search performance.