AIAny
AI Model2026
Icon for item

Hy3

Provides a large Mixture-of-Experts instruct LLM (295B total parameters, 21B active, 256K context) optimized for reasoning, long-context retention and agent workflows; open-sourced under Apache-2.0.

Introduction

Hy3 is notable because it pursues production reliability and long-horizon agent workflows at MoE scale rather than purely maximizing benchmark scores. The design and fine-tuning prioritize tool-calling stability, hallucination reduction, and multi-turn intent retention so the model behaves more predictably in real-world pipelines.

Key Capabilities
  • Sparse MoE architecture with 295B total parameters and ~21B active parameters, enabling higher parameter capacity while keeping per-token compute manageable; this translates into stronger reasoning and coding performance relative to many dense models of similar active size.
  • Very long context support (256K tokens) and improved multi-turn intent tracking, so it can handle large documents, extended agent chains, and long conversational state without rapid drift.
  • Production-focused post-training and RL scaling that reduced hallucination and formatting/tool-call failures; practical benefits include more reliable tool invocation and fewer invalid loops in agent setups.
  • Integration-friendly deployment recipes (vLLM, SGLang) and quantization/finetuning tooling aimed at lowering inference cost for real deployments.
Who It's For and Trade-offs

Great fit if you need an open-source instruct LLM for long-context document processing, multi-step agent orchestration, or productized coding assistants and can provision multi-GPU inference (or use supported inference stacks). Look elsewhere if you require minimal-resource local inference (Hy3 expects substantial TPU/GPU resources), strict small-model latency/footprint constraints, or if you prefer purely dense architectures for simpler deployment and compatibility in very small-scale environments.

Information

  • Websitehuggingface.co
  • OrganizationsTencent Hy Team, Tencent
  • Published date2026/07/02

Categories

More Items

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.