AIAny
AI Model2026
Icon for item

Laguna XS 2.1

A Mixture-of-Experts causal LLM (33B total, 3B active) tuned for agentic coding and long-horizon workflows; offers 262K-token context, mixed sliding-window/global attention, FP8 KV-cache and native preserved 'thinking' for tool-assisted agents, with local-ready quantized checkpoints.

Introduction

Most coding and agent workflows hit two hard limits: context length for long-horizon state and latency/memory for local inference. Laguna XS 2.1 targets that gap by combining a compact MoE design with a very large context window and engineering optimizations (FP8 KV cache, mixed SWA/global attention) so you can run agentic coding agents locally without sacrificing multi-step tool reasoning.

Key Capabilities
  • High-context agentic reasoning: preserves explicit "thinking" blocks and supports interleaved reasoning between tool calls, so multi-step plans and tool interactions remain coherent across long sessions.
  • MoE with low activation footprint: 33B total parameters but ~3B activated per token via 256 experts, enabling strong capability with reduced runtime memory compared to dense models — useful for desktop or small-server inference.
  • Long context + efficiency features: 262k token context, sliding-window/global attention mix, and FP8 KV-cache reduce memory and enable long-horizon code and terminal-style tasks.
  • Local-ready ecosystem support: official support and quantized variants for vLLM, Transformers, TRT-LLM, Ollama and llama.cpp, making practical local deployment and tool integration feasible.
Who it's for and tradeoffs

Great fit if you need an LLM to run locally for agentic coding, multi-step terminal automation, or long-context reasoning where preserving internal "thinking" improves outcomes. It trades raw single-turn SOTA for a balance of multi-step capability, local resource efficiency, and tool-call fidelity. Look elsewhere if you need the absolute top leaderboard single-shot accuracy or prefer fully closed-source commercial models with managed hosting and SLA-backed inference.

Information

Categories

More Items

Hugging Face
AI Model2026

Provides an EXL3 3.0 bits-per-weight quantization of a weight-edited GLM-5.3 UNCENSORED FP8 model for self-hosted text generation and agent workflows. Key characteristics: 753B MoE architecture, 273 GiB on disk, converted with ExLlamaV3; tool-call parsing requires preserving string arguments.

Hugging Face
AI Model2026

Processes English and German text with long-context reasoning and structured tool-calling. Uses a 78B mixture-of-experts architecture that activates ~3.46B parameters per token, offers native 262k-token context (validated to 1M), and is released as Apache-2.0 weights — suited for RAG, document processing and human-in-the-loop decision support.

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.