Token literacy: a core capability for responsible AI adoption
Curia AI Perspective
An AI Daily Brief episode on tokens prompted a sharper question for charity leaders: not how much AI costs, but what it delivers. The evidence points to a strategic capability, not a technical detail, one that sits squarely inside Principles, People and Capability.
An episode of The AI Daily Brief on AI tokens sent me looking harder at a question I keep asking: how do you actually know whether an organisation is using AI well? I can tell you this: watching the invoice tells you almost nothing on its own.
The episode, “Everything You Need to Know About AI Tokens”, aired on 2 August 2026 with guest Nufar Gaspar, hosted by Nathaniel Whittemore. As a quick bit of detail: a token is simply the unit an AI model reads and writes in, usually a fragment of a word rather than a whole one. The detail matters less than the pattern the episode traces in how organisations respond once that invoice becomes visible for the first time.
The pattern runs in four stages. First, an all-inclusive era, when flat subscriptions kept consumption invisible. Then a token-maxing era, when usage itself became treated as a sign of AI maturity. Then, as the bill became visible, a token-anxious era, when organisations impose limits and staff quietly retreat to safe, low-value uses of AI rather than risk an expensive prompt. The aim, the episode argues, is a fourth stage: token-smart, where organisations understand consumption well enough to spend it deliberately.
That last shift is more important for charity leaders than the token mechanics that precede it. It reframes a question that looks financial into one that is really about governance and the operating system around AI.
Efficiency is not the same as minimisation
The instinct to control an unfamiliar cost is to reduce it. With AI, that instinct can misfire. An organisation that tells staff to use fewer tokens may be quietly reducing the value it gets from the investment already made, not protecting it, particularly if the tokens being cut are the ones spent on the research, analysis and problem-solving that AI was adopted to help with in the first place.
Sarah Friar, OpenAI’s chief financial officer, made a related argument in July 2026, proposing what she calls “Useful Intelligence per Dollar” as a better measure of AI economics than price per token, set out in “A scorecard for the AI age”. Her point is clear: “the lowest price per token does not always produce the lowest cost per outcome.” A more capable model may cost more per token and still be cheaper overall, because it gets a task right the first time rather than needing several attempts, more context, and human correction along the way.
One preliminary 2026 study examined 30 software-development tasks run through the ChatDev multi-agent framework using a GPT-5 reasoning model. It found that the iterative Code Review stage accounted for an average 59.4% of token consumption. That does not establish a universal ratio for agentic work, or prove that token use translates directly into financial cost. It does illustrate why the first answer is only part of the economics: repeated review and refinement can consume more resources than initial generation in some workflows. In a McKinsey interview, Pay-i chief executive David Tepper reports that its telemetry shows token prices for similarly sized models falling sharply since 2022 while the cost per completed task has risen as workflows have become more involved. That is one provider’s observation rather than a universal market finding, but it reinforces the need to measure the whole task rather than the token price.
The conclusion is simple: cost per token tells you almost nothing on its own. Cost per accepted task, work a person is actually willing to use, is a much more useful measure, particularly when considered alongside quality, dependability and the value of the outcome.
There is no fixed answer to “which model”
A second, related shift is in how organisations choose which AI model to use. The instinct is to look for a single answer, often the best model, adopted once, for everything. The evidence suggests that question is already out of date. Different tasks warrant different models, configured differently. Some need speed, some need depth of reasoning, some are cost-sensitive, some cannot tolerate an unreliable answer at any price.
This is not a gap to be embarrassed about. The practice of matching model, configuration and task is genuinely still being worked out, across the industry and inside individual organisations, and that is a normal condition for a fast-moving field rather than a sign that nobody has done the groundwork. The responsible response is not to wait for a settled playbook before adopting AI, nor to pick a model once and consider the decision closed. It is to build the discipline to test a small number of representative tasks, measure what actually happens to cost, accuracy and staff time, and revisit that judgement as models, tools and the organisation’s own needs change.
What the model sits inside matters as much as the model itself
A July 2026 study, “The Harness Effect”, makes that point concrete. Researchers kept six foundation models unchanged and instead altered the orchestration around them, how the AI’s context is assembled, which tools it can call, how steps are sequenced, and how the process is monitored, comparing a conventional setup against an alternative design of their own, which they called the Writer Agent Harness. Across 22 evaluation tasks, the alternative design cut tokens per task by 38%, cost per task by 41%, and typical completion time by 44%, while task quality stayed broadly steady.
The caveat matters. This is one study, on a specific set of tasks and models, comparing one team’s own orchestration design against a baseline they also defined; it is not evidence that any harness redesign will produce similar savings elsewhere, and the authors’ own product is the one being tested favourably. Read narrowly, though, the finding still stands: changing what sits around a model, not the model itself, moved the numbers more than the difference between the models being compared did. The model is one part of a system. The context, tools, permissions and oversight around it are not incidental detail. They are a large part of where the value, or the waste, actually happens.
Where this sits in Principles, People and Capability
This is not just about engineering. It maps cleanly onto how we think about responsible AI adoption at Curia AI, structured around Principles, People and Capability.
Principles set the boundaries: what a system is allowed to do on its own, where human authority has to remain, what it can access, and what counts as an acceptable outcome.
People supply the judgement that a benchmark cannot: domain expertise, an understanding of the organisation’s mission, and, ultimately, the decision about what “good” actually looks like in this context, for this task.
Capability is where the harness question lives. It increasingly means not just choosing a model, but knowing how to design, test and govern what surrounds it: context, tools, evaluation and ongoing oversight.
It is worth being clear about what this research does and does not change for how we work. The Responsible AI Accelerator has not been redesigned around the Harness Effect study, and orchestration design is not becoming a new stage bolted onto it in response to one paper. Capability was always meant to deepen as the evidence base matures and organisations gain their own experience of what works. This is exactly the kind of development that thinking was built to absorb, not evidence that it needed correcting.
What this means in practice
For a trustee or chief executive, the practical shift is less about technology and more about the questions an organisation is willing to ask itself. Ask for cost per completed, accepted task, not cost per token or a flat monthly bill. Ask which tasks currently use AI, and whether the model and configuration were chosen deliberately or by default. Ask how the organisation tests a change before it is rolled out, and how often those choices are revisited. None of this requires deep technical knowledge. It requires the same discipline any board already applies to a new supplier or capital decision: proportionate scrutiny, evidence over instinct and a genuine willingness to revisit a judgement as the picture changes.
None of that argues for unlimited AI spend, and none of it argues for tight rationing either. It argues for treating token spend the way a well-governed organisation treats any resource: worth understanding properly before deciding whether to spend more of it, or less.
If your organisation is starting to ask what its AI spend is actually buying, that is a Capability conversation worth having properly, not an invoice to quietly trim.
Sources
- The AI Daily Brief, “Everything You Need to Know About AI Tokens”, 2 August 2026, featuring Nufar Gaspar.
- OpenAI, “A scorecard for the AI age”, Sarah Friar, 17 July 2026.
- Preliminary 2026 study on token consumption in ChatDev agentic software-engineering workflows.
- McKinsey, “Cost versus value: Managing agentic AI system performance” (interview with Pay-i CEO David Tepper), 8 July 2026.
- “The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI”, submitted 8 July 2026.