AIAny
AI Model2026
Icon for item

Qwythos-27B-v1

27B multimodal reasoning model built on Qwen3.5-27B that preserves the base model's native multi-token-prediction head, full vision tower, and a 1,048,576-token YaRN context window. Designed for agentic tool use, long-context reasoning, and research deployments; released under Apache-2.0.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<tool_call>` (20:128-20:139) before the end of `paragraph` 18 | - Serving: supports vLLM for 1M-context serving and has GGUF quantizations for local runtimes; the YaRN factor creates a long-context vs short-context fidelity trade-off (factor 4.0 recommended only when you need the full 1M window). 19 | - Safety & usage: intentionally uncensored for research; deployers must provide review, access control, and policy filtering when exposing sensitive cybersecurity or biomedical capabilities. > 20 | - Format & tooling: ships a chat template compatible with Qwen3.5 function-call conventions so tool calls appear as structured <tool_call> blocks. | ^ 21 | More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • OrganizationsEmpero AI, Alibaba (Qwen team)
  • Published date2026/07/13

Categories

More Items

Hugging Face
AI Model2026

Enables local use of a GGUF-quantized DeepSeek-V4-Flash-0731 via Unsloth Dynamic quantizations; provides a Q8 (162GB) lossless option and smaller Q4 variants for lower-memory inference and agentic scenarios using Unsloth tooling.

Hugging Face
AI Model2026

An OpenAI-compatible LLM checkpoint optimized for agentic and long-context scenarios, shipping DSpark speculative decoding and vLLM/SGLang deployment recipes; tailored for code-agent and multi-step reasoning workloads and released under MIT.

Hugging Face
AI Model2026

Provides a 2‑bit quantized build of Qwen3.6‑35B‑A3B for local serving via an OpenAI‑compatible HTTP API. Key features: 12.3 GB on disk, eschamoe mixed 2/3‑bit expert quantization with int8 dense layers, runs on a single 16–24 GB NVIDIA GPU and ships with Escha SGLang and ZML runtimes.