AIAny
AI Model2026
Icon for item

Qwythos-27B-v1

27B multimodal reasoning model built on Qwen3.5-27B that preserves the base model's native multi-token-prediction head, full vision tower, and a 1,048,576-token YaRN context window. Designed for agentic tool use, long-context reasoning, and research deployments; released under Apache-2.0.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<tool_call>` (20:128-20:139) before the end of `paragraph` 18 | - Serving: supports vLLM for 1M-context serving and has GGUF quantizations for local runtimes; the YaRN factor creates a long-context vs short-context fidelity trade-off (factor 4.0 recommended only when you need the full 1M window). 19 | - Safety & usage: intentionally uncensored for research; deployers must provide review, access control, and policy filtering when exposing sensitive cybersecurity or biomedical capabilities. > 20 | - Format & tooling: ships a chat template compatible with Qwen3.5 function-call conventions so tool calls appear as structured <tool_call> blocks. | ^ 21 | More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • OrganizationsEmpero AI, Alibaba (Qwen team)
  • Published date2026/07/13

Categories

More Items

Hugging Face
AI Model2026

Provides a drop-in checkpoint of DeepSeek-V4.1-Flash with weight-level abliteration that removes safety guardrails to produce uncensored outputs; preserves vision, MoE routing, 1M-token context and native FP8 quantization. Intended for advanced self-hosted deployment; requires large NVLink GPU domains and careful serving setup.

Hugging Face
AI Model2026

A reasoning‑efficient fine-tune of Qwen3.8-27B that penalizes overthinking tokens to shorten internal reasoning traces — about 58.3% fewer thinking tokens with <1% accuracy loss and ~1.95× speedup; designed for long-context, multimodal and quantized deployments.

Hugging Face
AI Model2026

Open-weights preview checkpoint for a multimodal reasoning model that generates text from text and image/video inputs, exposes adjustable reasoning effort and tool-calling, and supports an extended 262,144-token context for long-horizon tasks and agent-style workflows.