A unified model, not an assembly
Sam Altman confirmed at a conference in San Francisco what the industry had been expecting for months: GPT-5 will be a natively multimodal model. Unlike GPT-4, which combined separate modules for text, vision, and audio, GPT-5 will process all these modalities in a unified architecture from the training phase onward.
This approach, already adopted by Google with Gemini, allows the model to understand the relationships between text, images, and sounds in a much more natural way. For example, a user will be able to show a photo of a faulty electronic circuit while verbally describing the problem, and GPT-5 will combine both inputs to provide an accurate diagnosis.
The announced capabilities
According to information shared by OpenAI, GPT-5 will include several major advances:
- Native chain-of-thought reasoning: the model thinks before responding, without a separate mode like o1
- High-resolution vision: analysis of images and documents up to 4K with advanced spatial understanding
- Two-way audio: real-time voice conversation with emotion detection
- Integrated image generation: creation and editing of images directly in the conversation
- Long memory: 500K-token context with persistent recall between sessions
The merger of o1 and GPT-4o
GPT-5 marks the convergence of OpenAI's two development branches. The GPT-4 lineage (versatility and speed) and the o1 lineage (deep reasoning) are merging into a single model capable of dynamically switching between rapid responses and in-depth thinking depending on the complexity of the request.
“With GPT-5, we no longer ask the user to choose between a fast model and an intelligent model. The model automatically adapts its depth of reasoning” — Sam Altman, CEO of OpenAI.
Pricing and timeline
OpenAI plans a two-phase launch:
- June 2026: API access for developers (beta program), estimated pricing of $5/$15 per million tokens
- July 2026: general availability on ChatGPT Plus and Enterprise
Implications for the market
The announcement intensifies competition with Anthropic (Claude 4) and Google (Gemini 2.5). The convergence toward unified multimodal models now appears inevitable, and differentiation will depend on reliability, agentic capabilities, and the integration ecosystem.
For European companies, the issue of data sovereignty remains central. OpenAI has not yet announced dedicated European hosting for GPT-5, an area in which players such as Mistral retain a significant competitive advantage.
Comments· 3 comments
Do we know whether “native” integration means users will be able to combine image, voice, and text in the same conversation without switching tools, or is that still unclear?
That seems to be the likely implication of a unified multimodal model, but the summary does not spell out the exact user interface or which combinations of inputs will be available at launch.
The article’s wording suggests that vision, audio, and reasoning would be part of one model rather than separate systems. I’d still wait for official product details before assuming how seamless the workflow will be.