Somewhere in the next twelve months, someone in finance is going to look at your AI seat licenses and ask a fair question: what did this spend actually buy? If the honest answer is “people seem to like it,” that renewal conversation goes badly.
OpenAI waded into this problem on September 16, 2026, with expanded analytics in the ChatGPT Admin Console. The update pairs usage and cost data with a task classifier: it samples messages from ChatGPT Work and Codex, sorts them into named use cases like feature development, account research, and code maintenance, then shows how credits split across models, reasoning levels, and speed settings for each task. Admins also get a plugin leaderboard, a skills view, and for engineering teams, a dedicated screen that connects Codex activity to outcomes like merged work.
Two customer names came with the launch. Datadog is feeding OpenAI’s task categories into its own Agent Console product. Playco, a game studio, says it cut manual fixes by half while building prototypes on GPT-6 Astra.
Half. On real prototypes. That is a number a CFO can do something with.
But notice what happened here. The vendor just shipped the dashboard that proves the vendor’s value. That is convenient for OpenAI and genuinely useful for you, and both things are true at once. The problem is timing. If the only measurement you have when renewal season arrives comes from the company selling you the seats, you are grading your own subscription with an app written by the seller.
The fix is to build your own measurement framework now, while the stakes are low. OpenAI’s own recommended method, buried at the end of the announcement, is a good spine to hang it on.
One task, one baseline, one review date
OpenAI’s suggested playbook is almost embarrassingly plain: pick one task tied to a business priority, agree on a baseline with the person who owns that task, and re-check on a set date.
Plain is the point. Most attempts at measuring AI value die because they try to capture everything at once. Every team, every use case, every tool, tracked forever. Nobody sustains that. The one-task method works precisely because it is small enough to actually finish.
Say your support team drafts customer replies with ChatGPT. Before the framework: ask the team lead how many tickets per person per day close, and how long a typical draft takes. Write it down. That is your baseline. Pick a date sixty to ninety days out. On that date, measure the same two numbers again with AI in the loop.
All three outcomes are useful. Improved numbers give you a renewal argument with dates attached. Flat numbers usually mean a coaching problem, which is cheaper to fix than a software problem. And if things got worse, you found out early enough to react.
The part people skip is agreeing on the baseline with the task owner beforehand. If you measure first and negotiate what the numbers mean later, every review becomes an argument about methodology instead of a decision.
What the new console changes if you have it
If you run ChatGPT’s enterprise tier, the Admin Console update does real work for you. The task classifier answers the question procurement always asks, which is “what are people actually doing with this,” without you having to survey anyone.
The per-task credit breakdown is the quieter win. When you can see that a team is burning high-reasoning credits on routine briefs, that is not a spending problem. It is a training moment. Someone never learned which settings suit which task, and a thirty-minute session fixes months of wasted budget. Analytics that surface a coaching conversation are worth more than analytics that just total the bill.
The Codex view matters most for engineering orgs, because it ties agent activity to merged work instead of stopping at activity alone. “The agent ran 4,000 tasks” means nothing. “The agent contributed to work that shipped” means something.
Doing it without any of that
Most teams reading this do not have an enterprise console. The manual version costs one spreadsheet and one hour.
List every recurring task your team does with AI. Pick the one with the clearest before-state, meaning a task that already gets counted somewhere: tickets closed, articles shipped, reports produced, code reviewed. Record the current rate. Tell the person doing the work what you are doing and why, because covert measurement is how you poison trust. Then set the review date and put it on a calendar you actually look at.
You are building the same evidence the console produces, just by hand. The advantage is ownership. When the renewal negotiation arrives, your numbers came from your process, and you can defend every line of it.
The uncomfortable question this all avoids
There is a limit to usage analytics, and it is worth naming. Credits consumed, messages classified, tasks counted: these measure activity. They do not measure whether the activity mattered. A team can generate enormous, well-categorized volumes of AI-assisted work that changes nothing the business cares about.
That is why the one-task method starts from a business priority rather than from the tool. “Are we using ChatGPT a lot” is a question vendors can answer. “Did the support team get faster without quality dropping” is a question only you can answer, and it is the only one finance will ask.
What to do this week
Pick one task today. One that is counted, owned by someone who will answer your email, and tied to something a manager already cares about. Get the baseline number before Friday. Put a review date sixty days out on the calendar.
If you advise other companies, do this as a paid engagement. OpenAI just told every enterprise buyer that this measurement is the expected standard, and most of them have no idea how to produce it. A two-week audit that ends with a baseline document and a review-date plan is a product. The vendor announcement did your marketing for you.
The renewal question is coming either way. The only choice is whether you answer it with your own evidence or with a screenshot of someone else’s dashboard.


