AIAny
AI Model2026
Icon for item

Qwythos-9B

A 9B reasoning LLM fine-tuned from Qwen3.5 that ships with a 1,048,576-token context, native function-calling and tool-use, and notable benchmark gains (+34 MMLU, +30 gsm8k-strict).

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<tool_call>` (8:42-8:53) before the end of `paragraph` 6 | 7 | - 1M-token context out of the box: YaRN rope-scaling extends the native 262k window to ~1,048,576 tokens, enabling whole-repo reasoning, multi-document synthesis, and long agentic trajectories without RAG chunking. > 8 | - Tool-first design: emits Qwen3.5-style <tool_call> blocks natively. Evaluated with a python_executor + web_search harness and produced correct, source-cited answers on 7/7 hard factual prompts. | ^ 9 | - Measured reasoning lift: +34 points MMLU mean, +30 pts on gsm8k-strict vs. the Qwen3.5-9B base under matched evaluation settings. 10 | - Deployment-aware: includes sampling recommendations (T=0.6, top_p=0.95, top_k=20, repetition_penalty=1.05), large max_new_tokens budgets for the model's `<think>` reasoning block, and vLLM/SGLang serving notes for 1M contexts. More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • OrganizationsEmpero, Alibaba / Qwen team
  • Published date2026/06/19

Categories

More Items

Hugging Face
AI Model2026

Compresses Qwen3.8-27B into a 12.3 GB sensitivity-aware mixed-precision quantized checkpoint for long-horizon agent workloads; preserves BF16 fidelity (+0.02% PPL, 93.2% token Top‑1 agreement), supports 262K context and vLLM serving, text-only and Apache‑2.0 licensed.

Hugging Face
AI Model2026

Turns a context, a question, and 2–20 candidate answers into a single chosen option for classification, routing, ordered scores, and Boolean decisions. Built on mmBERT-small with a 144.3M-parameter decision head, runs on CPU with up to 8,192 combined tokens; FP32 weights occupy 550.5 MiB and are Apache‑2.0 licensed.

Hugging Face
AI Model2026

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.