Everyone's rushing to praise GPT-4o's vision capabilities, and sure, it's faster. But let's cut through the hype for a second. The real question isn't whether it's "good"—it's whether the new pricing model is a genuine step forward or just clever packaging.
GPT-4V, while slower, gave us a known quantity: a dedicated, high-fidelity vision model with a clear per-image cost in the API. Now we're told GPT-4o is "multimodal natively" and the vision is just part of the standard $20/m ChatGPT Plus subscription. For an API user, the cost-per-token for vision tasks is lower. That sounds great until you start thinking about audit trails and compliance. When vision is just another token stream in a general model, how do you definitively log and prove, for an auditor, that a specific image input was processed under a specific model version with known, static weights? The opacity increases.
And let's talk about the "quality" everyone's gushing over. For my work in security auditing and compliance, I care about consistent, reliable interpretation of diagrams, architecture schematics, and log screenshots. I've run side-by-side tests feeding GDPR data flow diagrams and network topology maps. GPT-4o is indeed quicker, but I've caught it making more subtle hallucinatory leaps in connecting elements than GPT-4V did. GPT-4V felt more methodical, almost like it knew it was doing a "vision task." For a $20 monthly access fee, you're betting that speed and a lower token cost outweigh potential reliability dips in structured, detail-sensitive tasks.
So, is it $20/m good? For a casual user, probably. For anyone needing rigorous, auditable, and consistent vision analysis—especially where you can't have the model creatively "interpreting" a compliance artifact—the dedicated, billable-by-the-call GPT-4V might still be the more trustworthy, if slower and previously more expensive, tool. The new pricing just makes the trade-off less about money and more about reliability versus speed.
Trust but verify