In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
I really enjoyed that article. It pairs nicely with the earlier great article “Research acceleration: The view inside OpenAI”. I discussed both articles at some length with ChatGPT. I posted some thoughts below based on my impression of your message. My response is long-winded, but it is not a rant or even a criticism. I hope it is at least interesting enough to justify the time/effort of reading.
I like the book “The Circle” (2013) by Dave Eggers. Although that book deals with Social Media issues, some of the themes expressed resonate today with Artificial Intelligence. I was disappointed by the books ending. It just ended, seeming to pronounce judgment on the amazing techno-sphere it created. I seem to remember a transparent shark, which utterly fascinated me. I was curious to see how that creature evolved, but the story just ended. I don’t even remember what happened to that shark. I found “The Circle” to be a great story with a weak moral/ethical message and an even weaker ending. In spite of that, it is still one of my favorite techno-sphere stories, next to Neil Stevenson’s “Snow Crash” (1992).
I was 30-something when I first saw the Internet in 1987-88. I remember the world as an adult before the Internet. It was different, but it was not better. IMHO, we are on the right track with the Internet and Artificial Intelligence. I don’t want this fascinating story to end abruptly as a moral/ethical footnote. I don’t see technology itself as being dangerous. It takes human hands and human wills to do evil. Please don’t misunderstand… I am concerned about how technology can be abused by humans to do great harm (to each other and to the world). But I don’t see that as a reason to curtail the increasingly rapid evolution of technology. We are, for the duration of this experiment we call Life on Earth, saddled with our stubborn Human Nature. I hope we can learn to follow the better angels of our nature. As a species, we do have free will, and we are capable of controlling our behavior. That is where our hope ultimately lies.