Project Zeros
Shutdown

EP 009 · Shutdown · 29 min · PT

O novo modelo omnipotente da OpenAI (GPT-4o)

May 15, 2024

About this conversation

OpenAI released GPT-4o—the "o" stands for omni—this week, and the framing matters. This isn't an incremental upgrade. It's a consolidation play that flattens the architecture underneath the entire AI stack. Before, audio, vision, and text ran as separate models layered together in ChatGPT. Now they operate as one, which means no latency between modalities, no performance tax from model-switching, and markedly better results across the board. The Whisper component—speech-to-text and text-to-speech—improves dramatically on less-resourced languages, which will matter far more in five years than it does today.

The product moves tell a different story. GPT-4o goes free on ChatGPT (users previously maxed out at 3.5). Premium users get a native Mac app that can see your screen, generate and execute code, and act as an omnipresent layer across your OS. There's no Windows version yet, which is conspicuous given Microsoft's $10 billion investment—but that's a separate tension. The image generator, now unified as GPT-4o instead of DALL-E, handles consistency and inpainting (selective region editing via prompt) with enough fidelity that professional workflows become viable, not just hobbyist tinkering.

What this does to the market is the real story. Startups built on top of Whisper, or on pure image generation, or on structured vision tasks, now face direct substitution risk. One API call replaces four. The moat shrinks. OpenAI is moving from platform-as-service toward something closer to OS-level integration—the Mac app announcement hints at this—which means friction for any company that bet on being a thin layer above their models. The rumor that Apple is negotiating to bake OpenAI into iOS and macOS alongside (or instead of) Google's Gemini signals where this heads. By WWDC in June, expect Apple's own AI story to harden up considerably. The timing is not coincidental.