English
Back
Open Account
钛媒体APP
wrote a column · Aug 2 00:00

Major AI firms are locked in a 'token milk tea war'

(This article was authored by Alphabet AI and published by TMT Post with permission.)
By Alphabet AI
Over the past two months, AI companies suddenly started collectively 'issuing coupons.'
On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work.
Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out.
This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits.
In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket.
Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged?
A defining signal of this round of token promotions emerged around May 20.
At the time, OpenAI offered YC Ventures up to $2 million worth of tokens per company, conditional on receiving a portion of equity in return.
OpenAI effectively used its own computing power as an investment vehicle—betting on these ventures while also encouraging them to adopt OpenAI’s models from day one and become paying customers down the line.
(This article was authored by Alphabet AI and published by TMT Post with permission.) By Alphabet AI Over the past two months, AI companies suddenly started collectively 'issuing coupons.' On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work. Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out. This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits. In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket. Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged? AI companies have started 'issuing coupons' A defining signal of this round of token promotions emerged around May 20. At the time, OpenAI offered YC Ventures up to $2 million worth of tokens each...
By June, tokens had started to resemble e-commerce discount coupons. After Codex erroneously deducted extra quota due to a system glitch, OpenAI reset user quotas twice and additionally granted users one more reset opportunity that could be saved for future use.
When the system malfunctioned, OpenAI compensated users with additional quota—essentially handing out a 'free refill voucher.'
Anthropic took a different path.
Fable 5 was initially offered free for just one week, but the promotion was extended twice. Concurrently, the Claude Code weekly quota increase of 50% was also prolonged. Starting July 20, Fable 5 was officially added to the premium subscription tier—the initial free trial for product launch ultimately evolved into a permanent membership benefit.
(This article was authored by Alphabet AI and published by TMT Post with permission.) By Alphabet AI Over the past two months, AI companies suddenly started collectively 'issuing coupons.' On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work. Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out. This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits. In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket. Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged? AI companies have started 'issuing coupons' A defining signal of this round of token promotions emerged around May 20. At the time, OpenAI offered YC Ventures up to $2 million worth of tokens each...
DeepSeek adopted an even more direct approach.
V4-Pro first received a permanent 75% price cut to attract developers with low-cost access; by the end of June, DeepSeek further proposed time-of-day pricing to manage usage volume through variable rates. MiniMax, meanwhile, bundled tokens into three tiers priced at $20, $50, and $120, corresponding to different usage levels,reminiscent of medium, large, and extra-large cup sizes on a bubble tea menu.
Now OpenAI has directly joined the price-cutting trend. On the last day of July, OpenAI announced an 80% reduction in API input and output pricing for GPT-5.6 Luna, now costing $0.20 and $1.20 per million tokens, respectively; Terra prices were reduced by 20%, dropping to $2 and $12 per million tokens. Subscription pricing and total quota allowances for Codex and ChatGPT Work remain unchanged, though using Luna and Terra will now consume less quota. The flagship model Sol retains its original pricing.
This 'token bubble tea battle' is not limited to membership packages; for high-value enterprise clients, vendors are even more direct. According to The Wall Street Journal, the CEO of AI customer service platform Pylon revealed thatsince the beginning of this year, the company has received approximately $1.6 million worth of free tokens from one model provider, while two others have each offered tens of thousands of dollars’ worth of credits—and some vendors have even granted several months of unlimited free usage.
Model price cuts and token giveaways are nothing new.
In the past, however, such promotions were typically confined to new product launches or used as developer support initiatives, generating buzz for a short period before fading away. But in the past two months, token-giving schemes have proliferated. Tokens are now taking on a role similar to coupons—systematically deployed for user acquisition, retention, and compensation—and have also become bidding chips in the competition for enterprise clients.
No matter how vendors frame this practice, it all points to one thing: AI companies are no longer satisfied with usage-based pricing and have started adopting more 'internet-style' tactics to win users.
Behind this shift is the growing difficulty of covering vastly different usage intensities with a single subscription model.
Light users prioritize low barriers to entry, while heavy users suffer from quota anxiety when running long tasks. As a result, tiered memberships, data bundles, credits, and peak/off-peak pricing have all emerged. AI companies are no longer just selling models—they’re also studying the well-honed strategies of consumer giants: how to get users to place their first order, what incentives will prompt them to buy another cup, and what reason to give them to come back after they’ve finished.
The battlefield for AI companies hasn’t left the arena of technological competition, but it has clearly added a new layer of commercial rivalry.
Or, to put it in more internet-native terms: he who retains users, wins the world.
Vendors have suddenly ramped up their token allocations, driven by two underlying shifts: Agent users are growing rapidly, and per-user token consumption is rising sharply; meanwhile, performance gaps among leading models on many common tasks are narrowing, making it easier for users to switch platforms. As a result, pricing, quotas, and membership benefits have become key tools in the battle to shape user habits.
First, AI users continue to flood in at a rapid pace.
As of early June, Codex had over 5 million weekly active users—more than six times the number when its desktop app launched in February.
More notably, non-developers now account for roughly 20% of these users, growing over three times faster than developers. The use of coding Agents has expanded beyond writing code—organizing documents, drafting reports, and analyzing data have become common applications, with some users even delegating entire workflows to them.
(This article was authored by Alphabet AI and published by TMT Post with permission.) By Alphabet AI Over the past two months, AI companies suddenly started collectively 'issuing coupons.' On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work. Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out. This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits. In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket. Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged? AI companies have started 'issuing coupons' A defining signal of this round of token promotions emerged around May 20. At the time, OpenAI offered YC Ventures up to $2 million worth of tokens each...
The way users consume tokens has also changed. According to official OpenAI data, by May this year, 80.6% of Codex individual users had submitted at least one task equivalent to 30 minutes of human work, 25.6% had submitted tasks equivalent to more than eight hours, and over 10% of users managed three or more Agents simultaneously each week.
In the past, users would ask a chatbot a single question and receive an answer—a simple process with relatively low token consumption. Now, an Agent may run a single task for several hours. Token usage has abruptly leapt from the '2G era' to the '5G era,' rendering previous quotas clearly inadequate.
Anthropic’s analysis of approximately 400,000 Claude Code sessions also shows that users spend an average of about 20 hours per week on the platform. Between October 2025 and April 2026, sessions dedicated to document writing and data analysis doubled from roughly 10% to 20%, with the average task value increasing by 27%.
As Agents shift from occasional experimentation to integration into daily workflows, token allowances have evolved from a bonus perk into a necessity.
Vendors are now competing for high-frequency, long-term paying users. As of February this year, Claude Code’s annualized revenue exceeded $2.5 billion—double the figure from the start of the year—with enterprise subscriptions growing fourfold and enterprise customers accounting for more than half of total revenue.
Whoever enables this cohort of users to engage with Agents more frequently stands to capture sustained growth in token-based revenue.
The problem is that these users aren't loyal.
The CEO of AI customer service platform Pylon put it bluntly: 'I don't see any loyalty whatsoever.'
The era when businesses would pick one traditional software vendor and stick with it through recurring payments is over. Companies now typically integrate multiple models simultaneously, assigning specific tasks based primarily on performance and price.
Cursor's product is inherently 'model-agnostic,' allowing users to switch between OpenAI, Anthropic, Google, and open-source models. Its head noted that whereas the best model for a given task might have changed only once every few months in the past, it can now shift several times a week.
This also explains why vendors are rushing to give away more tokens.
Agent users are growing rapidly, but most haven't yet settled into fixed usage patterns, and enterprises haven't decided which model(s) to permanently integrate into their operations.Vendors are racing to become the go-to tool for individual users and the default model embedded in enterprise workflows.
In 2026, researchers from the University of Trieste, King's College London, and University College London analyzed 7,156 GitHub pull requests generated by coding agents and found that no single tool led across all tasks. Often, differences caused by task type were greater than those between tools.
For users, the best model increasingly resembles a dynamic multiple-choice question.
Switching is also becoming easier.
(This article was authored by Alphabet AI and published by TMT Post with permission.) By Alphabet AI Over the past two months, AI companies suddenly started collectively 'issuing coupons.' On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work. Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out. This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits. In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket. Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged? AI companies have started 'issuing coupons' A defining signal of this round of token promotions emerged around May 20. At the time, OpenAI offered YC Ventures up to $2 million worth of tokens each...
DeepSeek offers interfaces compatible with both OpenAI and Anthropic, allowing MiniMax token packages to be directly integrated into Claude Code, Cline, and OpenClaw.
Developers don’t even need to change their existing workflows—simply swapping the underlying model is enough to redirect tasks.
AI companies are willing to subsidize tokens because they’re competing for users’ long-term workflows. Once a model is embedded in an enterprise system, switching costs rise rapidly, creating a moat for the AI product. Hence, early-stage subsidies clearly serve as a strategic move to capture entry points.
Model capabilities attract users, while subsidies give them a reason to stay—at least for now.
But there’s a fundamental difference between token discounts and 'bubble tea coupons': giving away one extra cup of bubble tea has clear and controllable costs, whereas handing out an extra hundred million tokens makes it extremely difficult to predict the resulting compute expenditure in advance.
And compute capacity remains extremely, extremely scarce right now.
Thus, an interesting phenomenon has emerged: AI companies are simultaneously distributing credits and imposing usage limits.
(This article was authored by Alphabet AI and published by TMT Post with permission.) By Alphabet AI Over the past two months, AI companies suddenly started collectively 'issuing coupons.' On July 31 Beijing time, OpenAI announced an 80% reduction in API pricing for GPT-5.6 Luna and a 20% cut for Terra, with corresponding reductions in token consumption quotas for these models when used in Codex and ChatGPT Work. Tokens suddenly felt like discount vouchers—usable as free samples for new products or compensation after system outages, bundled into membership plans, or offered in 'medium,' 'large,' and 'extra-large' sizes, with users simply topping up once their quota runs out. This closely resembles the recent 'milk tea price war' in food delivery: platforms compete by issuing coupons that appear to benefit users but are actually vying for their spending habits. In the past, competition among AI companies centered on model parameters, benchmark rankings, and product launches. But as agent user numbers surged, performance gaps narrowed, and switching costs declined, technical capability has become merely an entry ticket. Today, AI companies face a more immediate question: when users can effortlessly switch between different AI tools, how do you keep them engaged? AI companies have started 'issuing coupons' A defining signal of this round of token promotions emerged around May 20. At the time, OpenAI offered YC Ventures up to $2 million worth of tokens each...
On May 6, Anthropic announced a compute partnership with SpaceX, enabling Claude Code’s five-hour quota to double and eliminating peak-time throttling. After Fable 5 was restored, its usage was capped at 50% of a paying user’s weekly allowance, with any excess requiring purchased Credits. From the outset, the free trial was designed with clear cost boundaries.
MiniMax’s large token packages are also not unlimited. They come with concurrent limits on five-hour quotas, weekly allowances, agent concurrency, and peak-time throttling, and the company recommends pay-as-you-go pricing for production environments.
OpenAI’s latest price reduction also maintains a clear cost boundary. The subscription prices and total quota allowances for ChatGPT and Codex remain unchanged; only the quota deducted when using Luna and Terra has been reduced. The flagship model Sol’s pricing is untouched, while the newly launched Fast mode via API offers up to 2.5x speed at twice the standard price.
OpenAI’s multiple Agent products still share a common quota pool, with more complex tasks consuming more quota. Once the quota is exhausted, users can only purchase Credits, reset their quota, upgrade their plan, or wait until the next billing cycle for restoration.
Tokens cannot be issued infinitely, and users may migrate at any time—forcing vendors to segment eligible user groups, models, and usage periods into increasingly granular tiers.
Anthropic differentiates quotas by tiers such as Pro and Max, with Max further subdivided into 5x and 20x usage levels; users can continue purchasing Credits once they hit these caps. DeepSeek splits V4 into Flash and Pro tiers, charging separately for cache hits, standard inputs, and outputs, with prices doubling during peak hours.
Model providers are adopting tiered pricing, and enterprises are beginning to adopt tiered usage patterns as well.
More challenging still, subsidies can attract new users but rarely buy loyalty.
Vendors hope free Tokens will secure long-term API calls, but enterprises are quickly learning to 'grab the freebies and leave.' Pylon, for instance, collects free quotas from multiple providers while explicitly stating it feels no loyalty whatsoever. Meanwhile, Cursor, Zoom, and Hex integrate multiple models into their workflows, dynamically reallocating tasks based on price and performance.
Vendors aim to lock users into a single ecosystem, while enterprises strive to avoid being locked into any one provider.
Microsoft’s study of tens of thousands of engineers found that initial adoption of programming Agents is often driven by colleagues, but continued usage depends primarily on whether engineers genuinely have sufficient coding needs. Engineers who consistently use these tools see their merged pull requests increase by approximately 24%. Ultimately, what truly retains users is whether the product becomes embedded in daily workflows.
In other words, coupons can put a cup of milk tea in the user’s hand, but they can’t guarantee he’ll drink it every day thereafter.
AI companies now have to run two businesses at once: on one hand, acquiring and retaining users with subsidies like a consumer platform; on the other, allocating computing power and protecting gross margins like a cloud provider.
They’re caught in an awkward position—afraid users will leave, yet equally afraid users will consume too much.
The more tokens they give away, the less certain they are of retaining users. And if users do stay, computing bills inevitably rise. The toughest part is that usage and costs can scale together, but revenue may not keep pace.
Discounts always come to an end. Once companies stop giving away credits to attract users, who will actually retain customers and sustain this business?
Risk Disclaimer: The above content only represents the author's view. It does not represent any position or investment advice of Futu. Futu makes no representation or warranty.Read more
Thumbs Up
2
12K Views
Report
Comments
Write a Comment...
2