Ninety-seven percent of OECD countries now use AI in at least one area of government. Only 28% measure whether those deployments produce any impact at all. The gap between those two numbers is where the interesting design failure lives.
Every product has four audiences. The user, the buyer, the regulator, and the surrounding ecosystem. Government AI dashboards, the ones tracking token consumption and chatbot conversations, are built for exactly one of them: the buyer.
The buyer's dashboard
A department head needs to justify a procurement. Tokens consumed goes up and to the right. Chatbot conversations launched goes up and to the right. Documents processed, prompts sent. All reportable in a quarterly review.
Kevin Rooney put it plainly: "Prompts sent. Users onboarded. Copilots deployed. Adoption rates. These don't tell you whether it's creating real value."
The Granicus benchmark found 85.7% of government staff report AI helps them simplify repetitive tasks and save time. That sounds like progress until you notice it measures how staff feel about the tool, not whether the citizen on the other end got a faster or more accurate answer.
The user's absence
The user public, the citizen, never experiences a token count. They experience wait times and wrong answers. The California Management Review found 53 to 77% of respondents had bad or frustrating chatbot experiences. Only 14% of customer service issues were fully resolved through self-service. A Gartner survey found 64% of customers preferred companies not use AI for customer service.
These are user-public numbers. They do not appear on the buyer's dashboard.
The ODI tested 11 large language models against over 22,000 questions drawn from GOV.UK material. The models rarely refused to answer, even when they should have. One model falsely told a user that Guardian's Allowance required the death of a child. No activity metric would have caught that. It is a trust failure, and trust failures are invisible to dashboards built for the buyer.
Why the metrics drift
A poorly structured prompt can force a model to iterate and consume more tokens without producing useful output. If token consumption becomes a performance indicator tied to evaluations, workers may optimize for AI interaction frequency rather than task quality. The metric creates its own gravity.
Bill Schmarzo framed this precisely: "Your metrics become AI's laws. An organization isn't what it values. It's what it rewards. AI simply takes that literally."
MIT's Project NANDA reported in July 2025 that 95% of organizations saw no measurable return from generative AI investments. The organizations asked AI to optimize activity rather than value. The metrics were the failure, not the technology.
The OECD's own numbers confirm the structural gap. Monitoring and evaluation across member countries scores 0.65, while strategy scores 0.87. Governments are better at planning AI than at knowing whether it worked. Only 28% measure the operational burdens their services impose on citizens.
The regulator and the ecosystem see nothing
The regulator public gets no signal from these dashboards either. A dashboard that counts chatbot interactions cannot tell an oversight body whether the AI gave citizens accurate answers about benefits eligibility or tax obligations. The ODI study showed models confidently producing wrong answers to government queries. No activity metric flags a confident wrong answer. The regulator needs error rates and accuracy audits, not throughput charts.
The ecosystem public, vendors and integrators building on government platforms, inherits the same blind spot. When 55.7% of government organizations use AI but only 42.9% have formal policies, the ecosystem builds against unstated rules. Vendors optimize for the metrics the buyer rewards, not for the outcomes the citizen needs. The dashboard shapes the market.
What a user-public dashboard would track
Think Digital Partners proposed the shift directly: "Rather than measuring chatbot conversations, organisations could measure first-contact resolution."
First-contact resolution is a start. But a real user-public dashboard would also track time-to-outcome from the citizen's side and error rates on high-stakes queries like benefits eligibility. These are harder to collect. They require instrumentation that follows the citizen's journey, not the system's throughput. That difficulty is exactly why they get skipped. The buyer public controls the procurement, and the buyer public picks metrics it can report.
The design problem
This is a defaults problem. The default dashboard template ships with activity metrics because activity metrics are easy to instrument. Defaults are political. When a government measures tokens consumed instead of problems resolved, it has made a design decision about which public matters. The decision is usually implicit. That does not make it neutral.
A government AI dashboard that tracks tokens consumed but cannot answer "did the citizen's problem get solved" is a procurement justification artifact. The longer the buyer public remains the only audience those dashboards serve, the wider the gap grows between rising activity charts and flat citizen outcomes.