AIAny
AI Model2026
Icon for item

Ornith-1.0-35B

A 35B mixture-of-experts LLM specialized for agentic coding and tool-enabled code generation, fine-tuned with self-scaffolding reinforcement learning. Supports very long contexts, OpenAI-compatible tool calls, and multiple serving runtimes under an MIT license.

Introduction

Ornith-1.0-35B is notable because it treats scaffolding (the search or reasoning scaffold) as a learnable object alongside solutions — the model learns to propose scaffolds and then generate rollouts that follow them, improving search trajectories for coding tasks. That design choice is why its creators position it toward agentic coding workflows rather than generic chat or instruction-following alone.

Key Capabilities
  • Agentic coding and tool calling: emits well-formed function/tool calls and can be deployed behind OpenAI-compatible endpoints; useful when you need an LLM that interacts with external tools, shells, or environment APIs.
  • Self-scaffolding RL tuning: the model was post-trained to jointly optimize scaffolds and solutions, which empirically lifts performance on agentic coding benchmarks (strong Terminal-Bench / SWE-Bench results reported), so it tends to search and iterate more effectively on multi-step coding tasks.
  • Long-context and flexible serving: confirmed support for very large context windows (262,144 tokens) and multiple runtimes (vLLM, SGLang, Transformers, GGUF/llama.cpp, Ollama), so it fits RAG/large-repo code-understanding and agent scenarios that need long memory.
  • Licensing and footprint: MIT-licensed and published for community use; the 35B MoE variant targets relatively efficient single-GPU deployment paths (GGUF/llama.cpp) while also offering high-throughput server recipes.
Who it's for and tradeoffs

Great fit if you build or evaluate coding agents, tool-enabled assistants, or long-context code-understanding pipelines and want an open, MIT-licensed model that was explicitly tuned for agentic behavior. Look elsewhere if you need a model with strict guardrails and commercial support guarantees or if latency/resource constraints demand a substantially smaller, denser model — MoE inference patterns and recommended server setups assume modern runtimes (vLLM/SGLang) or specific GGUF workflows. The project also expects users to handle tool sandboxing and safety when exposing shell/agent capabilities.

Where it fits

Ornith-1.0-35B sits between research-grade agent models and practical coding assistants: stronger than many open-source dense 30–35B models on multi-step coding benchmarks due to its RL-based scaffolding, and more deployable for long-context agent use than very large, monolithic models when you require local hosting, function-calling, and extended context support.

Information

  • Websitehuggingface.co
  • OrganizationsDeepReinforce (deepreinforce-ai)
  • Published date2026/06/21

More Items

Hugging Face
AI Model2026

Compresses Qwen3.8-27B into a 12.3 GB sensitivity-aware mixed-precision quantized checkpoint for long-horizon agent workloads; preserves BF16 fidelity (+0.02% PPL, 93.2% token Top‑1 agreement), supports 262K context and vLLM serving, text-only and Apache‑2.0 licensed.

Hugging Face
AI Model2026

Turns a context, a question, and 2–20 candidate answers into a single chosen option for classification, routing, ordered scores, and Boolean decisions. Built on mmBERT-small with a 144.3M-parameter decision head, runs on CPU with up to 8,192 combined tokens; FP32 weights occupy 550.5 MiB and are Apache‑2.0 licensed.

Hugging Face
AI Model2026

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.