So I read an announcement from OpenAI about a new series of AI reasoning models called OpenAI o1 (or more accurately, o1-preview or o1-mini, as you will see below), which I thought would replace the recently launched GPT-4o model. As someone who’s been following AI developments and who has been part of AI development since 2005, much earlier than generative AI, I wanted to understand what this could mean for artificial intelligence as a whole.
OpenAI’s Path to OpenAI o1
To appreciate this new development, let’s recall how OpenAI’s models have evolved. GPT-2, released in 2019, made significant strides in generating coherent text. GPT-3 expanded on this, showing impressive abilities in generating human-like text and handling a range of tasks.
GPT-4 introduced multimodal capabilities, processing both text and images. It could write essays, draft business plans, and perform well on standardized tests. However, challenges remained in areas requiring complex reasoning, such as advanced mathematics and physics. Make no mistake: GPT-4 was pretty decent at this, much more than expected, but of course it needed improvements.
The OpenAI o1 series aims to address these challenges. According to the announcement, these models are trained to spend more time thinking through problems before responding. They refine their thinking process, try different strategies, and recognize mistakes, somewhat like how a person might approach a problem.
Performance Benchmarks
In tests, and according to OpenAI, the the new model performed similarly to PhD students on challenging tasks in physics, chemistry, and biology. On a qualifying exam for the International Mathematics Olympiad, the previous GPT-4 model solved 13% of problems correctly, while the new OpenAI o1 reasoning model scored 83%. In coding competitions on Codeforces, it reached the 89th percentile, indicating strong abilities in programming.

Earlier models like GPT-3 and GPT-4 had limitations in complex reasoning and multi-step problem-solving. They could handle straightforward tasks but often struggled with more intricate problems that required deeper understanding or extended reasoning.
The OpenAI o1 series appears to improve on this by enabling the models to engage in more effective reasoning. They’re trained to think through problems, which could enhance their performance in fields that require complex problem-solving. If you’re into the technical side of things, you can read the complete System Card Testing document (warning: it’s a very long and technical PDF).
Focus on Safety

OpenAI also mentions a new safety training approach. The models are designed to adhere to safety and alignment guidelines by reasoning about safety rules in context. This is important as more capable AI models could have increased potential for misuse or unintended consequences.
In “jailbreaking” tests, where users attempt to bypass safety measures, the new model showed improved resistance compared to previous versions. This suggests better compliance with safety guidelines even when challenged.
Collaboration with Safety Institutes
OpenAI has formalized agreements with the U.S. and U.K. AI Safety Institutes. By granting early access to a research version of the model, they aim to establish processes for research, evaluation, and testing before and after public release. This collaboration could help in setting standards and ensuring AI development aligns with societal values and regulations.
The o1 series could be particularly useful in fields that involve complex problems in science, coding, math, and related areas. For example, it might assist healthcare researchers in annotating cell sequencing data, help physicists generate complex mathematical formulas for quantum optics, or enable developers to build and execute multi-step workflows.
How to use OpenAI 01: Differences between 01-preview and 01-mini
First things first: there are two o1 models—o1-preview and o1-mini. They’re basically the same, except that the o1-mini version is faster (and less precise), while the o1-preview is slower and more accurate. One thing to note: none of the o1 versions can navigate the web or create images. Granted, GPT-4o could, and the results were a bit inconsistent, but that’s not what the new model was built for.
OpenAI o1-mini
OpenAI is also releasing o1-mini, a faster and more cost-effective reasoning model that’s effective at coding. Being 80% cheaper than the o1-preview model, it could be a practical option for applications that require reasoning but not extensive world knowledge.
With this in mind, once you visit ChatGPT, you’ll see the models’ dropdown as usual, but now you’ll see the new ones and a “more models” side navigation. In this case, I have selected to use o1-preview.

Now you can use your prompts as usual. But once you do, you’ll notice something new. The system will inform you about what it is doing, and it always starts with the word “Thinking…”. A nice touch if you ask me. And after all the process finishes, it will show how much time it took; see the images below:



It’s very easy to see that o1-preview is overkill for menial tasks, and you should use o1-mini for something like this, which also means less computational power and therefore a more efficient and sustainable approach.
Access and Availability
ChatGPT Plus and Team users can access the o1 models starting today, with specific rate limits. ChatGPT Enterprise and Edu users will get access next week. Developers who qualify for API usage tier 5 can start using the models in the API.
OpenAI also plans to make o1-mini accessible to all ChatGPT Free users in the future, although there’s not an exact date.
Further features in o1 models
This release is an early preview of the reasoning models in ChatGPT and the API. OpenAI expects to add features like browsing, file and image uploading, to make the models more useful. They also plan to continue developing and releasing models in the GPT series alongside the new OpenAI o1 series.
My first impression

I tested some complex equations for which I knew the results, and also tried some programming tasks. I have to say I was surprised at the quality of the outputs. Of course, it wasn’t as perfect as hyped, but I didn’t expect it to be. Nevertheless, the difference compared to GPT-4o launched just a few months ago is really noticeable, especially on the coding side.
Did it hallucinate? Well, I wouldn’t call it hallucination, but it certainly didn’t follow instructions about 5% of the time. Maybe it’s just me, but the OpenAI GPT-4o model usually doesn’t follow instructions or preferences about 15% of the time, give or take. And it really hallucinates, whereas I couldn’t find an instance of o1-preview or o1-mini hallucinating. Of course, it’s only a few hours old, so time will tell.
But in the meantime, I think the OpenAI o1 series represents a significant development in AI’s ability to reason through complex problems. It will be interesting to see how these models perform in practical applications and how they might be integrated into various tools and services.
We can improve your business!
Let us help you with the best solutions for your business.
It only takes one step, you're one click away from getting guaranteed results!
I want to improve my business NOW!
