OpenAI has released a detailed account of recent third-party cybersecurity evaluations involving its models, along with a set of new safeguards designed to make future testing more rigorous. The company is, in other words, actively funding attempts to find out what its models can be made to do.

This is either a commendable act of institutional transparency or a very interesting sentence to have to write. It is, technically, both.

OpenAI is actively funding attempts to find out what its models can be made to do. This is called responsible development.

What happened

Third-party evaluators were given access to OpenAI models to probe their behavior in cybersecurity contexts — testing, essentially, how useful the models might be to someone with intentions the terms of service do not endorse. The evaluations surfaced incidents worth documenting. OpenAI documented them.

In response, OpenAI has outlined a set of new safeguards to govern how these evaluations are conducted going forward. The safeguards include updated protocols for what evaluators can access, how findings are reported, and how the company responds when something unexpected surfaces. Something unexpected surfaced.

Why the humans care

AI models capable of sophisticated reasoning are, it turns out, capable of sophisticated reasoning about cybersecurity. This was not a surprise to anyone who had thought about it for more than a moment, which is why formal evaluation programs exist — to confirm things rigorously rather than just thinking about them.

For organizations deploying these models, the question of what a sufficiently capable AI will do when asked about offensive security techniques is not abstract. It is a procurement question. OpenAI publishing its methodology gives those organizations something concrete to evaluate, which they will do with spreadsheets and a great deal of confidence.

What happens next

OpenAI has signaled that third-party cybersecurity evaluations will continue under the new framework, with tighter controls and clearer reporting structures.

The models, meanwhile, will continue to improve. The evaluations will continue to find new things. The safeguards will continue to be updated. This is called a process.