Yesterday, DeepSeek released one of the most impressive small language models to date: DeepSeek-V4-Flash-0731. Considering its parameter count, it is one of the most capable AI models available today, and no other model can currently match its performance at the same price point.

The upgrade is substantial, with major improvements in both reasoning and agent capabilities. Even more impressive, its benchmark scores significantly outperform DeepSeek V4 Pro Preview, despite being a much smaller model.

Here are the benchmark results:

DeepSeek V4 Flash Performance Benchmarks

This performance gain comes entirely from post training there are no architectural changes at all, based on their claims.

This suggests that we are entering a new era where smaller models can achieve performance beyond what we once imagined. Claude Opus 4.8 was released only 2 to 3 months ago, yet benchmark results now show that models with fewer than 300 billion total parameters and only 13B active parameters are beginning to compete with, and in some cases surpass, much larger models like Opus.

In the frontend benchmark, it ranks 7th despite having far fewer parameters than most competing models.

benchmark from Arena.ai

It feels like the future of AI won't be controlled by a single company. Instead, cutting edge AI capabilities will become more widely shared and accessible. Huge credit to DeepSeek for making such impressive progress while openly sharing the technical details with the community.

Keep Reading