AIAny
AI Model2026
Icon for item

Ornith-1.5-35B-A3B

A 35B mixture-of-experts LLM tuned for agentic coding and end-to-end self-improvement: it jointly generates tasks, scaffolds, and solution rollouts. Activates ~3B params/token, supports 256K context (extendable), and emits chain-of-thought plus OpenAI-style tool calls.

Introduction

Ornith-1.5-35B-A3B matters because it pushes self-improvement beyond fixed human-curated tasks: the model continuously generates new tasks, designs scaffolds, and learns solution rollouts via reinforcement learning, which drives better search trajectories and stronger agentic coding behavior than similarly sized dense models.

Key Capabilities
  • Self-improving agentic training loop: the model is trained to propose tasks, build solution scaffolds, and produce rollouts that are jointly optimized by RL — this reduces reliance on static human-written harnesses and can discover higher-yield strategies autonomously.
  • MoE efficiency for long-context agenting: a ~35B MoE that activates ≈3B parameters per token, enabling large-capacity behavior while keeping per-request compute comparable to smaller dense models; natively supports 262,144-token context windows and validated extensions (YaRN) toward ~1M tokens.
  • Agent & tool-first output design: emits explicit reasoning blocks (chain-of-thought) and well-formed tool-call function blocks compatible with OpenAI-style tool APIs, making it suitable for tool-enabled agent frameworks and terminal coding agents.
  • Production-friendly serving: provided recipes for vLLM/SGLang, GGUF builds for local inference (llama.cpp/Ollama), and recommended sampling settings for reproducible benchmark runs.
Who Should Use It and Tradeoffs

Great fit if you need a model focused on autonomous coding agents and tool-enabled workflows, want very large context windows for multi-file/codebase reasoning, and can allocate multi-GPU serving (or use GGUF for local inference). Look elsewhere if you need a tiny single-GPU dense model for low-resource devices, if deterministic short-answer tasks are primary, or if you cannot accept the infrastructure cost of MoE serving (recommended ~2×80GB GPUs for full bf16 serving with large context).

More Items

Hugging Face
AI Model2026

Provides an EXL3 3.0 bits-per-weight quantization of a weight-edited GLM-5.3 UNCENSORED FP8 model for self-hosted text generation and agent workflows. Key characteristics: 753B MoE architecture, 273 GiB on disk, converted with ExLlamaV3; tool-call parsing requires preserving string arguments.

Hugging Face
AI Model2026

Processes English and German text with long-context reasoning and structured tool-calling. Uses a 78B mixture-of-experts architecture that activates ~3.46B parameters per token, offers native 262k-token context (validated to 1M), and is released as Apache-2.0 weights — suited for RAG, document processing and human-in-the-loop decision support.

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.