AIAny
AI Model2026
Icon for item

Naive-N0.5-Flash

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Introduction

Long-context LLMs are often bottlenecked by quadratic full-attention; this project demonstrates a predominantly local+sparse architecture that scales to a native 1,000,000-token context without any full-attention layers, trading dense global cost for an indexer-guided sparse backbone.

Key Capabilities
  • Hybrid attention designed for scaling: a 5:1 layout of Sliding-Window Attention (SWA, 128-token window) and DeepSeek Sparse Attention (DSA) lets most layers remain local while nine DSA layers use a lightweight 16-head indexer to select the top 2,048 tokens for backbone attention. This reduces per-token decoding cost at extreme contexts compared with traditional global-attention designs.
  • MoE compute profile with compact active footprint: 309B total parameters with 15.5B active parameters (Mixture-of-Experts) reduces inference compute compared to dense models of similar total size while retaining large-model capacity for routing-heavy workloads such as coding and research tasks.
  • Engineering and inference optimizations: supports FP8 mixed-precision, model weights ~315 GB, and an accompanying NaiveRT runtime tuned for single-stream speed and speculative decoding (reported single-stream peaks up to ~2,000 tokens/s in ultrafast setups). API and weights are released under the MIT license.
  • Training and evaluation posture: multi-stage continued pretraining over ~3.25T tokens (50B indexer warmup, 3T sparse-attention training, 200B LR decay) and benchmarked for coding and AI-research tasks; the architecture purposefully optimizes for long-horizon context and agentic/code workloads.
Who it's for and trade-offs

Great fit if you need to run or research very long-context generative or agentic workflows (codebases, multi-file reasoning, long transcripts) and can provision FP8-capable NVIDIA GPUs and ~315 GB of model storage. It is also suitable for researchers comparing sparse/local attention designs and MoE routing behavior at scale.

Look elsewhere if you need lightweight local inference on CPU/mobile, immediate compatibility with non-FP8 hardware, or minimal disk/GPU footprint—the model expects substantial GPU memory and FP8 support. Also expect engineering complexity when integrating the DSA indexer and MoE routing into custom inference stacks.

More Items

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.

AI Model2026

Explains how Jev turns input state into typed decisions and probabilities without generating text. Introduces parallel sampling and RLCD training, with workflow evaluations and caveats for software automation.