Why This Launch Matters
If you've been waiting for AI that's cheap enough to run all day, DeepSeek is trying to hand you the keys. The Hangzhou-headquartered lab rolled out V4.1 Flash on Thursday, a pared-back model it says beats Moonshot's Kimi K3 while costing below a penny per million tokens. The move turns up the heat on everyone from Anthropic PBC to Z.AI Co., right as Anthropic and OpenAI get ready for public listings.
DeepSeek's pitch is blunt: good enough at a rock-bottom price can win the next wave of users. The lab known for inventive tricks says this version brings "more efficient architecture" at lower cost.
The New Economics of AI Agents
Cheap, high-powered Chinese models are redrawing the map for how developers, startups, and enterprises plan spending on AI agents that can run for hours with little oversight. The industry's spotlight has swung to agents that code through long sessions, call tools, or comb the web with minimal human input. Artificial Analysis even labeled the new landscape a DeepSeek "death zone": to keep up, rivals must either undercut on price or clearly outclass on capability. Starting Sept. 14, DeepSeek will retire V4-Pro by automatically routing all inference to V4.1 Flash and billing at the lower rates.
In times of rapid change, steady strategies help protect and grow savings. Join Briefs Finance CEO Jaspreet Singh on September 29th for a FREE live investor workshop, How to Profit From A Dollar That's Losing its Value, where he shows how we're spotting investment opportunities as the dollar falls. Save your spot.
Market Moves and the Memory Squeeze
Investors noticed. MiniMax Group Inc. and Z.AI each tumbled more than 8% in Hong Kong, and Alibaba Group Holding Ltd. slid more than 2%. Memory names also took a hit: SK Hynix's US-listed shares were down as much as 5.8% Thursday, while Micron fell up to 5.3%. Part of the backdrop is a persistent shortage of memory chips driven by AI demand, which has buoyed profits for suppliers like Micron and SK Hynix.
Under the Hood and What It Means for Your Wallet
DeepSeek says V4.1 Flash employs a "causal encoder-decoder" setup that lights up only a tiny portion of the model's 552 billion parameters at any moment - a smaller slice than V4-Pro required. The lab also reduced the KV cache, trimming needs for high-bandwidth memory and solid state drive storage. On tests, the new model tops V4-Pro in coding and agentic tasks, though it still trails the flagship systems from Anthropic and OpenAI. Net effect: if your workload leans toward long-running agents that write code, call tools, and browse, paying far less per unit of work could stretch the same budget a lot further, even if you give up some peak performance.
Smart investors stay curious, seeking simple ways to stretch every invested dollar. Our CEO Jaspreet Singh is hosting a FREE live investor workshop, How to Profit From A Dollar That's Losing its Value, on September 29th. Sign up free to join him live.
