Migrating a production AI agent from Claude Opus 4.8 to GPT-5.6 Sol: the three engineering costs behind the 2.2x benchmark

In the 48 hours after OpenAI’s 7/9 dual launch (GPT-5.6 and ChatGPT Work on the same day), the agent ecosystem started realigning along three visible axes. OpenAI temporarily lifted the 5-hour rate limit on Plus / Business / Pro. Codex weekly active users crossed 6 million. And one production agent — Ploy, who builds real marketing websites with an agent that plans across apps, reads and writes code, takes its own screenshots, and decides for itself when the build is done — migrated its core model from Claude Opus 4.8 to GPT-5.6 Sol and posted first-impression numbers that beat Opus 4.8 across the board: 2.17x faster wall-clock, 27% lower cost, 0.970 vs 0.936 on the visual score. ...

July 13, 2026 · 10 min · cuigh

OpenAI's quiet 7/12 GPT-5.6 prompt guide: a design philosophy reversal

On July 12, OpenAI quietly posted Prompting guidance for GPT-5.6 Sol on its developer site — a single, official recipe book for how to write prompts for the new model family. Read alongside this morning’s https://cuigh.com/posts/gpt-5-6-sol-production-migration-2026/ Ploy piece, it reads as the other half of the 7/9 launch story. The headline number is sharp: on OpenAI’s own internal coding-agent eval, swapping the system prompt from the GPT-5-era “encyclopedia” shape to a lean shape lifted eval scores 10–15%, cut total tokens 41–66%, and reduced cost 33–67%. That is not “the model is smarter so you can write shorter”. It is the reversal of prompt-design philosophy itself, now that model capability has caught up. ...

July 13, 2026 · 7 min · cuigh

GPT-5.6 and ChatGPT Work: OpenAI is bundling its strongest model with its strongest agent, betting on the same thing

On July 9, OpenAI did two things on the same day: pushed GPT-5.6 to the top of the “strongest model” list, and shipped ChatGPT Work — an agent that can run across web, mobile, and desktop to complete hours of work for you. Sam Altman posted on X: “obviously the best model we have ever produced, but also one of the best blog posts we have ever produced.” That line looks like product PR. It is not. What Altman is actually saying is that neither the model nor the blog post is the decisive asset anymore. An agent that can run autonomously for hours is. ...

July 10, 2026 · 6 min · cuigh

GPT-5.6 Sol shipped with a built-in ID check. Here's what changes for the rest of us.

On the afternoon of June 26 (U.S. Eastern), OpenAI published a preview post for GPT-5.6 Sol on its blog, and parked the system card at deploymentsafety.openai.com. The same day, the Washington Post reported that the U.S. government will decide which users get access to GPT-5.6. The same day, Reuters reported Anthropic is releasing Mythos only to “trusted partners.” Three events, less than 24 hours apart, each individually heavy. I am not going to write “AI regulation is here” — that ground was covered on June 14 in AI regulation just walked into the product. What I want to do today is zoom in much closer: what does GPT-5.6 Sol’s launch actually change, for a normal user, a normal builder, a normal product spec writer. ...

June 27, 2026 · 7 min · cuigh