DeepSeek app
– Solen Feyissa/Unsplash

DeepSeek fired a warning shot at AI rivals by slashing API prices up to 90% amid soaring enterprise token usage.

The South China Morning Post reports that DeepSeek slashed prices on inputs for its latest V4 model that debuted just this month. Prices were also discounted on input cache hits, where tokens in a prompt correspond to data already found in the cache memory. If the required data is found in the cache, it’s retrieved and used for computation.

DeepSeek’s reduction sees the cost per million tokens drop from about $0.145 to just $0.036 for its flagship V4-Pro – providing developers the tantalizing prospect of having access to a frontier-level AI model at dramatically lower costs for running repeated requests.

DeepSeek’s price reductions come as token usage among enterprises is spiraling.

Take Disney. Per a Business Insider article, some of its engineers were using Claude around 51,000 times per day, leading to the entertainment giant having to employ an AI adoption dashboard to manage token usage across its various AI tools.

The splurge of AI usage has resulted in the phenomenon of "tokenmaxxing," with usage becoming an actual benchmark for productivity.

Meta, for example, had an internal dashboard of its own that morphed into a leaderboard where staff competed to see who could use AI the most, which was later shuttered. Visa was another big-name encouraging staff to "tokenmaxx," resulting in the firm spending 1.9 trillion tokens in March alone.

DeepSeek’s aggressive pricing push could entice enterprises like Disney and Visa that are running tens-of-thousands of workloads per day on the same platform with significant savings.

Val Bercovici, chief AI officer at Weka, said DeepSeek has left open model rivals no choice but to match its pricing model, saying: “Frontier labs will try to hold the line at first. But even with DeepSeek's 90% lower cached input token prices on the table, gross token spend will keep surging. Jevons Paradox is undefeated.”

DeepSeek went viral in early 2025, launching its frontier-level flagship model built on hobbled hardware courtesy of U.S. chip restrictions.

In December last year, the firm built DeepSeek V3.2, which showed performance levels on par with OpenAI’s GPT-5 on some industry benchmarks. The model was trained using Nvidia's China-market specific H800s hardware that has a reduced 400 Gb/s of bidirectional NVLink bandwidth.

In early January, the Chinese lab followed up with Engram, a "conditional memory" concept that bypasses graphics processing unit (GPU) memory constraints by offloading tactical knowledge (simple information lookups) to central processing unit (CPU) random access memory (RAM) in order to relieve an AI model’s core computational network and allowing it to focus on complex reasoning.