Skip to content

Independent e-magazine

the OUTSPOKEN digest

Inside OpenAI's Pivot From GPT-4o to the o1 Reasoning Era

OpenAI spent 2024 rebuilding its flagship around two ideas: a model that sees and hears in real time, and one that stops to think before answering.

Outspoken Digest Technology Desk

Tuesday, December 10, 2024/3 min read

A laptop screen glowing in a dim room, displaying a conversational AI chat interface mid-response.
Photo: Afroswede via Openverse (CC BY 2.0)

Sam Altman likes to call things by their internal names in public, and this year the habit paid off. First there was the omni model. Then there was the model that thinks. Between May and December, OpenAI quietly rewired what "a new AI model" is supposed to mean, and the shift says more about where the company is headed than any single demo did.

It started on May 13, when OpenAI streamed a launch event and introduced GPT-4o, its new flagship. The "o" stands for omni, and the point was fusion: one model handling text, vision and voice natively instead of stitching together separate systems, as TechCrunch reported from the event. The demo that stuck in people's heads was a live voice conversation with almost no lag, the assistant picking up sarcasm, interrupting itself, even singing on request.

What made GPT-4o different from GPT-4

GPT-4o was not pitched as smarter than GPT-4 on paper. It was pitched as faster, cheaper and dramatically better at real-time audio, with response times low enough to feel like a phone call rather than a chat window. OpenAI folded voice, image and text processing into a single network rather than routing audio through a separate transcription step, which is what let it react to tone of voice and background noise instead of just words.

For millions of free ChatGPT users, GPT-4o was also the first time OpenAI's best available model stopped being a paywalled feature. That decision, more than any benchmark, is probably what will be remembered: the assumption that frontier AI sits behind a subscription cracked, at least for a few months, until usage limits crept back in.

Why OpenAI built a model that thinks before answering

Then, on September 12, everything about the interface changed again. OpenAI released o1-preview and o1-mini, a new model family built around a different premise: instead of producing an answer in one pass, the model generates an internal chain of reasoning first, effectively working through a problem before committing to a response.

The tradeoff is visible immediately. o1 is slower than GPT-4o on hard questions, sometimes taking many seconds where GPT-4o would have answered instantly. But on competition math, coding challenges and science questions that require multi-step logic, the difference in accuracy was large enough that OpenAI built an entirely separate product line around it rather than folding it into the main model.

What changed when the full o1 model shipped

The preview label came off on December 5, when OpenAI used the opening day of its "12 Days of OpenAI" event to ship the full o1 model alongside a new $200-a-month tier called ChatGPT Pro. According to VentureBeat's coverage of the announcement, the full model cut major reasoning errors by roughly a third compared to the preview and added the ability to reason over uploaded images, not just text.

ChatGPT Pro subscribers also get o1 pro mode, a version that spends even more compute per answer, aimed squarely at people running data analysis, legal review or research tasks where a wrong answer is expensive. Ten times the price of ChatGPT Plus is a steep ask, and OpenAI is betting there is a real market of professionals who will pay it.

How developers and researchers are using GPT-4o and o1 differently

What has emerged over the back half of the year is less a rivalry between the two models than a division of labor. GPT-4o handles the fast, conversational, multimodal work: customer support drafts, live translation, image description, anything where latency matters more than depth. o1 gets pulled in for the harder stuff, the kind of problem where a person would normally sketch out a plan on paper before writing anything down.

That split matters because it signals OpenAI no longer believes a single model curve can serve every use case. Scaling up parameter count produced diminishing returns; scaling up how long a model is allowed to think, apparently, did not. Competing labs took notice almost immediately, and reasoning-style models are already showing up as a category rather than a one-off experiment.

What comes next for OpenAI's model lineup

Heading into 2025, OpenAI has effectively split its roadmap into two lanes: a fast, multimodal, everyday assistant lineage descended from GPT-4o, and a slower, deliberate reasoning lineage descended from o1. Whether those two lanes eventually merge into one model that knows when to think and when to just answer is the obvious next question, and it is one OpenAI has not yet answered publicly. For now, the choice of model has become a decision users have to make themselves, several times a day.

Published in The Outspoken Digest

Share this story

the OUTSPOKEN digest

Beyond boundaries. Independent stories on technology, culture, and the trends shaping how we live.