September 22, 2026
StepFun sells Kimi K3-level intelligence three times cheaper: $0,72 per task instead of $2
Step 5 Preview scores 44 points in the Artificial Analysis Intelligence Index, matching Kimi K3 (max), while a task costs $0,72 versus $2,00.

Artificial Analysis
@artificialanlys
StepFun's Step 5 Preview scores 44 points in the Artificial Analysis Intelligence Index, matching Kimi K3 (max), at roughly 2,8x lower cost per task, though it trails similarly ranked models on agent benchmarks Step 5 Preview is @StepFun_ai's new flagship: 600 billion total parameters, 27 billion active. It replaces Step 3.7 Flash (released in May 2026). It scores 44 points in the Intelligence Index, matching Kimi K3 (max) and sitting just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key points: ➤ An Intelligence Index task on Step 5 Preview costs roughly 2,8x less than on models with the same score. About $0,72 per task versus ~$2,00 for Kimi K3 (max) at the same score of 44, and ~$2,01 for GLM-5.3 (max) at 45. It comes down to pricing: $1/$2,70 per 1 million input/output tokens, cheaper than both on input and output. The only model that scores higher (46) at a lower task cost ($0,13) is MiMo-V2.6-Pro ➤ The model is strongest in frontier reasoning, which is also where it makes its biggest leap over Step 3.7 Flash. Step 5 Preview scores 46% on Humanity's Last Exam, close to Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both results rose sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Its AA-Omniscience accuracy is higher than GLM-5.3's with fewer parameters, but it also hallucinates more often. With 600 billion parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark for factual memory and hallucinations. That is ahead of GLM-5.3 (max, 34%, 753 billion) and behind Kimi K3 (max, 48%, 2,8 trillion). It attempts more questions than GLM-5.3 (68% versus 55%), and when it does, it hallucinates more often (43% versus 30%), giving it 16 points in the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agent benchmarks are where Step 5 Preview trails its peers in the Intelligence Index. On GDPval-AA, our main benchmark for agent work, it scores 1566 Elo, behind Qwen3.8 Max (1668) and GLM-5.3 (max, 1646). The gap remains on Terminal-Bench 4.0 (33% versus 39% and 42%), AA-Briefcase (1432 Elo versus 1640 and 1525), and AutomationBench-AA (51% versus 56% and 62%) Main model details: ➤ Size: 600 billion total parameters, MoE with 27 billion active ➤ Context window: 1 million tokens ➤ Multimodality: text, images, and video as input; text as output ➤ Pricing: $1/$2,70 per 1 million input/output tokens, cached input $0,05 per million ➤ Availability: StepFun's own API, with the weights scheduled to be released on October 15 ➤ License: weights are currently closed, with release scheduled for October 15

· 43.3K views
China's StepFun has unveiled a flagship model that performs on par with Kimi K3 while costing three times less.
Artificial Analysis tested Step 5 Preview and awarded it 44 points in its Intelligence Index on 22.09.2026. That matches Kimi K3 (max), while GLM-5.3 (max) and Qwen3.8 Max each score one point higher. One Intelligence Index task costs $0,72, versus $2,00 for Kimi K3 (max) and $2,01 for GLM-5.3 (max).
It comes down to pricing. StepFun charges $1 per million input tokens and $2,70 per million output tokens, with cached input priced at $0,05. That is cheaper than both neighboring models by score on both input and output. On cost per task, only MiMo-V2.6-Pro beats Step 5 Preview, scoring 46 points for $0,13.
Before and after. StepFun's previous flagship was Step 3.7 Flash, released in May 2026. On Humanity's Last Exam, the new model scored 46% — 25 points higher than its predecessor. On CritPt, it scored 21%, up 19 points. The first figure nearly catches Kimi K3 (max) at 47%.
On agent tasks, the model trails its peers by score. It has 1566 Elo on GDPval-AA, versus 1668 for Qwen3.8 Max and 1646 for GLM-5.3 (max). The gap remains on Terminal-Bench 4.0: 33% versus 39% and 42%.
The picture is equally mixed on factuality. Step 5 Preview reaches 42% accuracy on AA-Omniscience, versus 34% for GLM-5.3 (max), but hallucinates more often: 43% of answers versus 30%.
The model is available through StepFun's own API. Its context window is 1 million tokens; it accepts text, images, and video as input; and it has 600 billion parameters, with 27 billion active per request.
StepFun has promised to release its weights on October 15, 2026.
