The race to build increasingly capable artificial intelligence systems is beginning to incorporate a new concern: how to ensure that the pace of development does not end up exceeding people's capacity to supervise it. In this context, OpenAI has put forward a proposal to establish common technical standards to evaluate frontier models and coordinate the management of risks associated with their evolution.
The initiative comes after several signals from the sector pointing in the same direction. Anthropic has warned of the advance towards systems capable of taking on a growing part of the AI development process itself, to the point of contemplating a future scenario of recursive self-improvement. The company emphasizes that this level of autonomy does not yet exist, but believes it could arrive before institutions are prepared to face it.
OpenAI now proposes that safety should not depend exclusively on the internal protocols of each laboratory. Its proposal involves developing shared criteria to measure capabilities, evaluate risks, and establish procedures for possible incidents, with the participation of AI safety organizations, independent experts, researchers, and companies in the sector. The company argues that these standards should be applicable to both open and closed models.
The challenge of AI being able to contribute to creating the next generation of AI
One of the scenarios that attracts the most attention is that of so-called recursive self-improvement. This refers to the possibility that AI systems increasingly participate in the design, training, evaluation, or development of subsequent models, progressively reducing human intervention.
Anthropic has already documented how its systems are taking on a growing part of tasks related to software development and research. The company believes that if this trend were to advance to allow an AI to autonomously design and develop its successor, the challenges related to the supervision and control of these systems would also increase.
OpenAI maintains, in line with this debate, that the progress of capabilities must be accompanied by an equivalent advance in alignment, evaluation, and safety mechanisms. The goal is for the most sophisticated models to continue operating within limits defined by people, even as their autonomy increases.
A proposal that seeks to coordinate laboratories and governments
OpenAI's proposal does not simply suggest reinforcing security measures within its own models. The company is committed to establishing a common technical language that allows comparing capabilities and risks between different systems and facilitates international cooperation.
Among the issues that should also be defined are what incidents should be reported, when to do so, and what response is appropriate depending on their severity. The company also advocates for greater collaboration between governments, laboratories, and specialized organizations to share information on vulnerabilities and security practices.
The approach comes as other companies in the sector are strengthening their own mechanisms. Anthropic, for example, has recently published new research on alignment and safety and has reported real cases of misuse of its models, while also tightening its protection systems against certain risks.
The underlying question, therefore, is no longer limited to what the next generation of artificial intelligence will be capable of doing, but rather how that progress will be measured and who will set the limits when machines begin to participate more and more directly in their own development. OpenAI now proposes that some of these rules be common. The debate, however, will have to extend far beyond a single company and, presumably, a single country.
Add ElConstitucional.es as a preferred Google source for free.
Stay informed about all the latest breaking news with the best information. Against disinformation, for democracy and social rights.