NaiveAI posts a 309B MoE under MIT

On 27 September 2026 NaiveAI published Naive-N0.5-Flash on Hugging Face. The repository was created at 22:14 SGT that day. The model is a 309-billion-parameter mixture-of-experts (MoE) model with 15.5 billion active parameters. NaiveAI releases the weights and inference code under the MIT licence.

NaiveAI's technical blog is titled "Building Frontier AI with AI". It says AI models designed the hybrid attention architecture and optimized training, inference and deployment. Human researchers set direction and made key decisions. The company says its R&D system serves close to ten million sandboxes a week, with 100,000 active at peak.

NaiveAI says it builds for coding and AI research work. It says the model opens "a path toward recursive self-improvement". That is NaiveAI's own framing.

NaiveAI on a Xiaomi MiMo-V2.5 base

  • NaiveAI built the model on Xiaomi's open-weight MiMo-V2.5 base. The model card thanks the Xiaomi MiMo team and the DeepSeek team.
  • NaiveAI also built NaiveRT, its inference system. The company says AI-centered R&D produced and tuned NaiveRT.
  • An FP8 repository holds the weights for the Transformers quick start.
  • NaiveAI says an API "will also be provided" at $0.10 input, $0.40 output and $0.01 cache reads per million tokens. The model card gives no launch date for the API.

1M context with no full-attention layers

The model has 48 layers in eight six-layer modules. Of these, 39 layers use sliding-window attention with a 128-token window. Nine layers use DeepSeek Sparse Attention (DSA), which selects the top 2,048 tokens. DSA uses grouped-query attention with four KV groups and a 16-head indexer. No layer uses full attention, yet NaiveAI states a native 1M-token context.

Training added 3.25 trillion tokens: 50 billion for indexer warmup, 3 trillion for sparse-attention training and 200 billion for learning-rate decay. NaiveAI says NaiveRT delivers 50 tokens per second per user in Standard mode. It says Ultrafast mode reaches up to 2,000 tokens per second.

NaiveAI reports these coding scores. The best comparison score on each chart follows in brackets:

  • DeepSWE v1.1: 67.8 (Muse-Spark-1.3, 75.4)
  • ALE-CLI: 32.4 (Opus-5.5, 34.3)
  • Terminal-Bench 2.1: 86.7 (DeepSeek-V4.1-Flash, 90.6)
  • SWE-bench Pro: 73.6 (Opus-5.5, 89.9)
  • ProgramBench: 17.5 (Opus-5, 37.0)
  • NL2Repo-Bench: 71.9, the top score on its chart
  • FrontierSWE v1: 78.2 (Fable-5 with fallback, 88.2)

NaiveAI bar charts comparing Naive-N0.5-Flash with open and closed models on seven coding benchmarks Caption: NaiveAI's coding results, with Naive-N0.5-Flash highlighted in yellow. NaiveAI ran or compiled every score. · Source: NaiveAI, Naive-N0.5-Flash model card, Figure 2, 27 Sep 2026 · link

On AI research tasks, NaiveAI reports 37.5 on PostTrainBench, 73.7% on MLE-bench-30 and 63.2 on PaperBench. Its MLE-bench-30 chart places it above Sonnet-5 at 66.9%. NaiveAI says the MLE-bench-30 figure is an average position score, following the Gemini 3.6 Flash model card protocol. On NanoGPT SpeedRun it reports 73.8 seconds, against 77.5 seconds for an undisclosed model from Recursive Superintelligence, Inc.

Vendor scores and a 315 GB FP8 footprint

  • All scores are NaiveAI's own runs or figures it took from other publishers. The Frontier found no independent reproduction.
  • NaiveAI ran its evaluations in Claude Code 2.1.207 with a 1M-token context, temperature 1.0 and top-p 0.95. Scores from other harnesses may differ.
  • Several comparison scores come from different sources and dates. The blog lists each source.
  • The model requires FP8-capable NVIDIA GPUs. The weights take about 315 GB before inference memory.
  • The Transformers path needs version 5.17.0 or later and trust_remote_code=True. Review the remote code before you run it.

Benchmark it on your repos before switching

  • Teams that self-host coding models should test Naive-N0.5-Flash on their own repositories. Compare it with DeepSeek-V4.1-Flash and their current model.
  • Infra teams should size FP8 GPU memory for 315 GB of weights plus long-context KV cache.
  • Security reviewers should audit the remote code in the repository before production use.
  • Cost owners should wait for the API to go live before they plan around the $0.40 output price.