DeepSeek V4 Flash Is Beating Models That Cost 50x MORE!

Categories: AI

Summary

DeepSeek V4 Flash matches Gemini 3.6 Flash performance at a fraction of the cost—0.28¢ per 1M input tokens vs competitors charging 10-50x more. This signals a major shift: open-source labs are now competitive with frontier models on price-per-task metrics, forcing incumbents to cut prices aggressively.

Key Takeaways

  1. DeepSeek V4 Flash costs $0.0028 per 1M input tokens and $0.0028 per 1M output tokens, while Claude Opus costs 43¢-87¢. This 50x cost difference with near-parity performance on benchmarks fundamentally changes the economics of building with AI.
  2. The model achieves 82.7 on MMLU (10-point jump from Pro), 76.7 on CyberGym security benchmarks, and 25.1 on AutoBench—outperforming GLM 5.2 on software engineering tasks while remaining lightweight enough to run locally.
  3. Price-per-task metric is now the decisive competitive advantage, not raw benchmark scores. DeepSeek V4 Flash is 'currently unbeatable' on this metric, indicating builders should optimize for efficiency over marginal performance gains.
  4. Open-source labs have become competitive with Google and Anthropic. DeepSeek V4 Flash now matches Gemini 3.6 Flash—a development that would have been 'crazy to think about' one year ago, signaling rapid market consolidation.
  5. The model is 'not that heavy' and suitable for local deployment, making it a pragmatic choice for builders balancing cost control with inference performance—especially relevant as OpenAI just discounted Luna 80% and Terra 20%.

Related topics

Transcript Excerpt

We have a new model today coming from Deepseek. Yes, Deepseek version 4 Flash is officially live now and in public beta. They massively upgraded its agent capabilities and the benchmark scores might be surprising for a lot of us. So, what exactly is new? Let's get into it. So, if we look at the Deepseek version for Flash, you're going to notice one thing. This model is not trying to compete with Opus 5 or GPT 5.6 Soul. rather this model is trying to compete with itself. Number one is showing that they made a model which is much more efficient, faster and costs way less than their pro version which is a good release for them. But not only that, this model is quite competitive when it comes to GLM 5.2, Opus 4.8. So keep that in mind when we look at the benchmarks. You're not going to see thi…

More from In The World of AI