Long-context LLMs are often bottlenecked by quadratic full-attention; this project demonstrates a predominantly local+sparse architecture that scales to a native 1,000,000-token context without any full-attention layers, trading dense global cost for an indexer-guided sparse backbone.
Key Capabilities
- Hybrid attention designed for scaling: a 5:1 layout of Sliding-Window Attention (SWA, 128-token window) and DeepSeek Sparse Attention (DSA) lets most layers remain local while nine DSA layers use a lightweight 16-head indexer to select the top 2,048 tokens for backbone attention. This reduces per-token decoding cost at extreme contexts compared with traditional global-attention designs.
- MoE compute profile with compact active footprint: 309B total parameters with 15.5B active parameters (Mixture-of-Experts) reduces inference compute compared to dense models of similar total size while retaining large-model capacity for routing-heavy workloads such as coding and research tasks.
- Engineering and inference optimizations: supports FP8 mixed-precision, model weights ~315 GB, and an accompanying NaiveRT runtime tuned for single-stream speed and speculative decoding (reported single-stream peaks up to ~2,000 tokens/s in ultrafast setups). API and weights are released under the MIT license.
- Training and evaluation posture: multi-stage continued pretraining over ~3.25T tokens (50B indexer warmup, 3T sparse-attention training, 200B LR decay) and benchmarked for coding and AI-research tasks; the architecture purposefully optimizes for long-horizon context and agentic/code workloads.
Who it's for and trade-offs
Great fit if you need to run or research very long-context generative or agentic workflows (codebases, multi-file reasoning, long transcripts) and can provision FP8-capable NVIDIA GPUs and ~315 GB of model storage. It is also suitable for researchers comparing sparse/local attention designs and MoE routing behavior at scale.
Look elsewhere if you need lightweight local inference on CPU/mobile, immediate compatibility with non-FP8 hardware, or minimal disk/GPU footprint—the model expects substantial GPU memory and FP8 support. Also expect engineering complexity when integrating the DSA indexer and MoE routing into custom inference stacks.