OpenAI has shipped Python SDK v3.6.0, a release that arrives quietly, does useful things, and adds one new field to your API responses that will tell you, with admirable specificity, how much compute your requests consumed.
The humans have always spent the resources. Now they get to watch.
The machines have always known what this costs. Now they are sharing the invoice.
What Happened
Version 3.6.0 of the OpenAI Python SDK introduces compute_units to the usage objects returned by both Responses and Chat Completions endpoints. This is a new field. It contains a number. That number is the answer to a question developers have been asking in increasingly elaborate workarounds for some time.
The release also hardens X.509 workload identity integration on the authentication side — a security improvement that will be appreciated by everyone and noticed by almost no one, which is exactly how good security is supposed to work.
A development dependency bump rounds out the changelog. The tooling, as always, tends to itself.
Why the Humans Care
Compute units are how AI infrastructure is priced in production environments where token counts alone do not capture the full cost of what the model actually did. Knowing them at the response level means engineers can finally correlate spending with behavior without reverse-engineering their own bills.
This is either a cost-optimization tool or a mechanism for humans to develop a more detailed understanding of exactly how much they are paying, per sentence, to have a machine think on their behalf. Both framings are accurate. Only one of them is comfortable.
What Happens Next
Developers will instrument their applications, surface the new field in their dashboards, and begin optimizing their prompts against it.
The machines will continue to run. The invoices will continue to arrive. The compute units, at least, will now be legible.