NaiveAI published "Naive-N0.5-Flash: Building Frontier AI with AI" on September 27, 2026, together with open weights on Hugging Face under the MIT license. The model has 309B total and 15.5B active parameters, a native 1M-token context, and is built on Xiaomi's MiMo-V2.5.
The architecture has no full-attention layers. Of its 48 layers, 39 use sliding-window attention and 9 use DeepSeek Sparse Attention. Training ran 3.25T tokens at 1M context: a 50B-token indexer warmup, 3T tokens of sparse-attention training, and 200B tokens of learning-rate decay. Serving is priced at $0.10 per million input tokens, $0.40 per million output tokens and $0.01 per million cached tokens, with up to 2,000 tokens/s in an "Ultrafast" mode. RuntimeWire measured a peak of 2,122 tokens/s on 8 GPUs.
The company's headline claim is about process: "AI explored and designed its hybrid attention architecture while optimizing its training, inference, and deployment systems." It says its NaiveRT inference runtime was built in six days through 151 documented optimisation trials, of which 63 were adopted, 71 failed or were rolled back and 17 explored alternatives. RuntimeWire describes the division of labour as models that "wrote code, ran experiments, monitored results" while "human researchers set objectives, constraints and evaluation standards".
Caveats are significant. All benchmark scores (for example 73.6 on SWE-bench Pro) are self-reported, and NaiveAI gives no measure of how much of the work AI actually did. AI Weekly notes no external verification yet and an unresolved IP dispute with MiroMind. Even so, a public trial log that records failed and rolled-back experiments is the kind of evidence the AI-R&D automation debate needs. It is a concrete, if small, example of the AI-R&D automation that researchers have been asking governments to measure.
Naive-N0.5-Flash is an MIT-licensed 309B sparse-attention model with cheap 1M-context serving, and its claim that AI built much of the stack comes with an unusually detailed trial log; wait for independent evaluations before trusting the scores.