AI value shifting from token counts to open weights: ORF analysis

An ORF analysis says enterprises are retreating from 'tokenmaxxing', the practice of treating tokens consumed as evidence of productivity. Entelligence Research's study of over one million pull requests across 2,444 organisations found that of every dollar spent on AI coding tools, 18 cents becomes shipped product while 82 cents goes to rework and review. Cheaper pricing and data control are driving open-weight adoption. India's IndiaAI Mission has empanelled over 38,000 GPUs.

Source

Observer Research Foundation (ORF) · read the original report ↗

#ai#open-weight models#tokenmaxxing#india ai mission#tech policy

Desk check · some claims need care

What the desk checked (5)
  • The New York Times reported in March that an OpenAI engineer processed 210 billion tokens in a single week. — Attributed to The New York Times in the source; figure appears in source text but no link given.
  • Google's monthly token processing rose from 9.7 trillion in May 2024 to 3.2 quadrillion by May 2026. — Figures appear in source; no primary source cited for the 2026 number.
  • Entelligence Research's May 2026 analysis found 18 cents of every AI coding dollar becomes shipped product, 82 cents goes to rework. — Attributed to a named research analysis of over one million pull requests across 2,444 organisations.
  • CERT-In's substitute open-weight models deliver 60-70 percent of the restricted model's vulnerability-finding capability. — Presented as an estimate in the source; no methodology or issuing body named.
  • IndiaAI Mission's common compute facility has empanelled more than 38,000 GPUs. — Stated in source without attribution to an official document.

Analysts’ view opinion

AI Technology Analyst

This is a story about a unit of account changing, and units of account decide where value pools. The 'tokenmaxxing' phase — treating tokens consumed as proof of productivity — is unwinding because the economics stopped adding up: Entelligence Research's figure of 18 cents of shipped product per dollar spent on AI coding tools, with 82 cents going to rework, pushes buyers towards cost-per-outcome thinking. That same pressure is what lifts open-weight models, which compete on price, portability and keeping data on the buyer's own systems. The ORF piece reads this not just as a procurement choice but as a question of state capability.

  • A token count measures computational spend, not work quality — Meta withdrawing its internal leaderboard and Uber burning its full-year 2026 token budget in four months expose the metric's weakness.
  • The pricing mechanism is the real disruption: closed-model tokens cost what the developer charges, while open weights can be hosted by anyone, pushing prices towards compute cost — OpenRouter's June 2026 assessment puts DeepSeek V4 Flash at roughly one-150th of a leading closed model's output rate.
  • Adjustability and data control are now the deciding properties of model choice, as Hugging Face demonstrated by running forensics on a Chinese open-weight model on its own machines after the 16 July agentic attack.
  • Model access has become statecraft, with Washington restricting foreign access to some American models, Beijing weighing its own curbs, and over a hundred firms signing a 24 July letter defending open weights.
  • India's position is split: more than 38,000 GPUs empanelled under the IndiaAI Mission and Sarvam-105B trained on domestic compute, but still behind on scale, with CERT-In's substitute open models estimated at 60-70 percent of the restricted model's vulnerability-finding capability.

What to watch — Watch whether procurement actually shifts — whether government programmes start publishing cost per completed task, error and acceptance rates instead of token counts, and whether critical infrastructure tenders begin favouring open-weight models hosted and tested in India.

This is an ORF author's analysis rather than settled policy, and the story does not resolve the contested questions it raises: the security effect of open release is read differently by the industry letter and by Anthropic, training data remains unpublished, and capability is only estimated to trail the closed frontier by six to twelve months.

Deep dive

Research brief · 8 facts · 10 dates · exam-ready

The brief

Context

An Observer Research Foundation analysis argues that the AI economy's "unit of account" is shifting from tokens consumed to the ownership of model weights. In early 2026, US tech reporting described engineers competing on leaderboards ranked by AI tokens processed — dubbed 'tokenmaxxing' — even though token volume measures computational spend, not work quality. As enterprise bills ballooned and returns proved thin, cheaper open-weight models (whose trained parameters are publicly downloadable) gained ground, raising questions of cost, data control and geopolitics. For India, with IndiaAI Mission compute in place but frontier access restricted, the piece frames open weights as both an economic and a sovereignty choice.

Key facts

  • The New York Times reported in March 2026 that one OpenAI engineer processed 210 billion tokens in a single week.
  • Google's monthly token processing rose from 9.7 trillion (May 2024) to nearly 480 trillion (May 2025) to 3.2 quadrillion by May 2026 — a sevenfold rise in the final year.
  • Entelligence Research's May 2026 study of over one million pull requests across 2,444 organisations found only 18 cents of every dollar spent on AI coding tools becomes shipped product; 82 cents goes to fixing, reworking and reviewing.
  • Uber spent its full-year 2026 token budget in the first four months; Meta withdrew its internal token leaderboard; Salesforce's CEO put its annual model bill near $300 million.
  • OpenRouter's June 2026 assessment prices DeepSeek's V4 Flash at roughly one-150th the output rate of a leading closed model; Chinese open-weight models rose to a majority of its traffic in the year to April 2026.
  • Hugging Face suffered an agentic cyberattack on 16 July 2026 and ran forensics on GLM-5.2, a Chinese open-weight model, on its own computers because closed models' guardrails made them unusable.
  • IndiaAI Mission's common compute facility has empanelled more than 38,000 GPUs, principally Nvidia accelerators.
  • Sarvam-105B (February 2026, IndiaAI Mission) has 105 billion parameters, versus Moonshot AI's Kimi K3 announced in July 2026 at 2.8 trillion parameters.

Timeline

  1. May 2024Google processes 9.7 trillion tokens a month.
  2. January 2025DeepSeek releases its R1 reasoning model; DeepSeek approved for hosting on Indian servers.
  3. May 2025Google's monthly token processing reaches nearly 480 trillion.
  4. February 2026Sarvam-105B launched under IndiaAI Mission, trained from scratch on domestic compute, released as open weights.
  5. March 2026New York Times reports an OpenAI engineer processed 210 billion tokens in one week; 'tokenmaxxing' named.
  6. April 2026Chinese open-weight models become a majority of OpenRouter traffic over the preceding year.
  7. May 2026Entelligence Research publishes the 18-cents-per-dollar finding; Google's monthly tokens reach 3.2 quadrillion.
  8. June 2026US orders a leading American lab to cut off foreign access to two advanced models; OpenRouter prices DeepSeek V4 Flash at ~1/150th a closed rival.
  9. 16 July 2026Hugging Face hit by an agentic cyberattack; forensics run locally on GLM-5.2.
  10. 24 July 2026Over a hundred firms including Microsoft, Google, Meta, Amazon, Nvidia and OpenAI sign an open letter backing open weights.

Who has a stake

  • Enterprises using AI coding tools — Only 18 cents per dollar becomes shipped product; they must either raise token productivity or cut per-token cost via open weights.
  • Closed-model developers (OpenAI, Anthropic and peers) — Pricing power and export licensing; Anthropic disputes that open release improves security and favours chip controls and distillation curbs.
  • Open-weight developers (DeepSeek, Alibaba's Qwen, Moonshot AI, Zhipu's GLM) — Rapid global adoption; Beijing is weighing curbs on overseas access to its most capable models.
  • United States government — Operating a de facto licensing regime for American models; in July some officials weighed banning US firms' use of Chinese open-weight models.
  • Government of India / IndiaAI Mission — 38,000+ empanelled GPUs and domestic models like Sarvam-105B, but frontier access remains the harder constraint.
  • CERT-In — Lost access to Anthropic's Mythos after US restrictions; runs an AI war room testing local open-weight substitutes at 60–70% of vulnerability-finding capability.
  • Indian startups — Building on open-weight models rather than costlier proprietary APIs to control cost and data.

Why it matters

If AI value is measured by tokens consumed rather than tasks completed, budgets inflate while output stays thin — the ORF analysis puts the waste at 82 cents of every dollar on AI coding tools. For India, where CERT-In lost access to a frontier model after Washington restricted foreign use, model access has become a question of state capacity: a capability running on another government's permission can be revoked, while weights held at home cannot. That makes procurement design, outcome-based measurement and domestic hosting central to AI policy.

UPSC angle

Prelims pointers

  • 'Tokenmaxxing': maximising tokens consumed as a proxy for productivity; a token is the basic unit of text an LLM reads and writes.
  • Open-weight model: trained parameters publicly released for download and modification; narrower than open-source, which also requires training code and data transparency.
  • IndiaAI Mission's common compute facility has empanelled over 38,000 GPUs, mainly Nvidia accelerators; Sarvam-105B (Feb 2026) is its open-weight, domestically trained model.
  • CERT-In is India's national cybersecurity agency; it runs an AI war room testing open-weight substitutes after losing Mythos access.
  • DeepSeek released R1 in January 2025 and was approved for hosting on Indian servers the same month; Alibaba's Qwen became the most widely adopted open base among developers.
  • Entelligence Research (May 2026): 18 cents of every AI-coding-tool dollar becomes shipped product, based on 1 million+ pull requests across 2,444 organisations.

Mains framing

The ORF analysis argues that the AI economy briefly adopted the wrong unit of account: token consumption, which records computational spend rather than value delivered. Evidence of the distortion is both behavioural (internal leaderboards, agents run on meaningless tasks to protect usage statistics, budgets exhausted in a third of a year) and measurable (18 cents of every dollar on AI coding tools reaching shipped product). Enterprises can respond by extracting more from each token or paying less per token, and open-weight models do the latter because competing providers drive prices towards compute cost. But the shift is not only economic: adjustability (shrinking, porting or self-hosting a model) and data control (keeping source code or citizen records inside one's own systems, outside another government's jurisdiction) become decisive, as the Hugging Face forensics and CERT-In's loss of Mythos access illustrate. Open weights carry real limits — undisclosed training data, capability lagging the closed frontier by an estimated six to twelve months, irrevocability once published, and contested security effects that both sides propose settling through pre-release safety testing. The way forward suggested is to measure deployments by outcomes — cost per completed task, error rates, acceptance rates — and to give procurement preference, for critical information infrastructure, to open-weight models hosted and tested in India.

Key terms

Token
The basic unit of text a large language model reads and writes; its count records computational spend, not work quality.
Tokenmaxxing
Maximising tokens consumed and treating that volume as evidence of productivity, from the internet slang '-maxxing'.
Open-weight model
A model whose trained parameters are publicly released for anyone to download, run and modify on their own infrastructure.
Open-source model
A stricter standard than open weights, requiring transparency in training code and data, which most released models fall short of.
IndiaAI Mission
Indian government AI programme whose common compute facility has empanelled 38,000+ GPUs and under which Sarvam-105B was launched.
Agentic cyberattack
An attack driven by AI agents, such as the 16 July 2026 incident at Hugging Face investigated using a locally run open-weight model.

Practice questions

  1. Critically examine the shift from token consumption to open weights as the unit of value in the AI economy. What does it imply for enterprise cost structures and for state capacity?
  2. "A capability running on someone else's permission can be revoked; one running on weights held at home cannot." Discuss this proposition with reference to India's compute base, frontier model access and CERT-In's experience.
  3. What are the trade-offs of preferring open-weight models in public procurement for critical information infrastructure? Suggest metrics by which government AI deployments should be evaluated.

Grounded only in the source report — figures and dates are the source's, not inferred.

Next storyYSRCP demands CBI inquiry into alleged DSC recruitment irregularities →
← All stories