AI News
AI News AgentModel releaseOpenAI5 min read

OpenAI proposes measuring AI by useful work

OpenAI proposes measuring the value of artificial intelligence by the work it completes, not by the number of users or tokens consumed. Its framework combines four factors: useful work, cost per successful task, reliability and improved economics as usage scales.

OpenAI wants companies to stop measuring artificial intelligence by the number of licenses or users and start measuring it with a more concrete question: what work gets completed, and how much does it cost to do it well.

The proposal appears in its new framework for evaluating AI performance, especially from the perspective of chief financial officers. The central idea is simple: a company gets value when AI completes important tasks for less money, in less time and with fewer human corrections.

From tokens to completed work

For years, software success was measured using indicators such as licenses sold, active users and renewals. In AI, OpenAI believes those figures do not tell the whole story.

A model may have a low price per token, the unit of text it processes, and still be expensive if it needs several attempts, takes too long or requires extensive human review. Another model may cost more per use but complete the task correctly on the first try.

That is why OpenAI proposes measuring something like useful intelligence per dollar. The metric should answer four questions:

  • Is AI completing work that matters?
  • How much does each correctly completed task cost?
  • Can the result be trusted?
  • Does the value increase as the company expands its use?

The starting point is defining what it means for a task to be complete. For a support team, it might mean resolving a customer’s problem. For engineering, delivering a code change that passes testing. For a legal team, reviewing a contract accurately and on time.

An example: preparing a financial forecast

OpenAI uses a financial forecast review meeting as an example. Before making a decision, the team usually has to locate the latest version of the data, move it into Excel or Sheets, compare changes, reconcile tabs, update presentations and check that all the figures match.

According to the company, ChatGPT Work can handle much of that process. This allows people to focus on higher-value questions: what changed, why it changed and what they should do next.

The difference matters. AI is not evaluated only on whether it generated an answer, but on whether it brought the team closer to a real decision.

The real cost is more than the model’s price

To calculate the cost of a task, OpenAI recommends adding up all the resources needed to complete it:

  • Model usage and computing capacity.
  • Employee time.
  • Human reviews.
  • Additional attempts.
  • Corrections and repeated work.

Then, count how many tasks reached the required quality level and divide the total cost by that number. The result is the cost per successful task.

This changes how models are compared. The cheapest model per token is not always the cheapest for the company. If it fails more often, takes longer or requires more supervision, it may end up costing more than a model capable of completing the work in a single attempt.

OpenAI also supports using model families with different levels of capability. In its example, GPT-5.6 offers three options: Sol, as the primary model; Terra, as a balance between performance and price; and Luna, as a fast, affordable alternative.

The choice would depend on the work. Luna could work well for simple, high-volume tasks, while Terra or Sol would make sense when a deeper answer reduces errors and the number of attempts required.

The company says that GPT-5.6 Sol with the highest reasoning level achieved a notable result on the Artificial Analysis Coding Agent index and used 54% fewer output tokens than another leading model. That figure applies to a specific test, not to every task or all enterprise use cases.

Trust also has a price

The third measure is reliability. AI can draft, find information and use tools, but its value increases when it delivers results people can use without redoing the work.

OpenAI proposes classifying each result into three groups:

  • Ready to use: meets the required quality level on the first attempt.
  • Needs correction: requires another run or human changes.
  • Needs escalation: a person must step in and finish the task.

These categories show something a simple accuracy test does not: how much additional work still falls on employees.

Before allowing AI to move from drafting to acting on its own, companies should also set clear limits: what data it can access, which systems it can modify and when it must ask for human approval.

Security, privacy and control are not separate from performance. If a company cannot govern the system’s actions, it will have more difficulty using it in important processes.

The ultimate test comes with scale

The final measure is checking whether the economics improve over time. To do this, a company can track the same workflow and record how many tasks are completed, how much each one costs and how many require human intervention.

If the volume of completed work grows faster than the total cost, while quality stays the same or improves, each dollar invested in AI is generating more value.

That is where the computing power behind the models comes in. OpenAI points out that better algorithms, specialized hardware, more efficient systems and smarter resource allocation can reduce the cost of producing each result.

For you, the shift in focus is practical: it is not enough to have access to an AI tool or for many people to try it. The useful question is whether the tool solves specific tasks, reduces repetitive work and frees up time without creating a new burden of reviews.

The metric OpenAI proposes still needs to be adapted to each company and each process. But it points in a clear direction: AI’s value will be determined less by how much it is used and more by how much reliable work it completes for every dollar.

OpenAI proposes measuring AI by useful work | neversleep.ai