Migrating a production AI agent from Claude Opus 4.8 to GPT-5.6 Sol: the three engineering costs behind the 2.2x benchmark

In the 48 hours after OpenAI’s 7/9 dual launch (GPT-5.6 and ChatGPT Work on the same day), the agent ecosystem started realigning along three visible axes. OpenAI temporarily lifted the 5-hour rate limit on Plus / Business / Pro. Codex weekly active users crossed 6 million. And one production agent — Ploy, who builds real marketing websites with an agent that plans across apps, reads and writes code, takes its own screenshots, and decides for itself when the build is done — migrated its core model from Claude Opus 4.8 to GPT-5.6 Sol and posted first-impression numbers that beat Opus 4.8 across the board: 2.17x faster wall-clock, 27% lower cost, 0.970 vs 0.936 on the visual score. ...

July 13, 2026 · 10 min · cuigh

OpenAI's quiet 7/12 GPT-5.6 prompt guide: a design philosophy reversal

On July 12, OpenAI quietly posted Prompting guidance for GPT-5.6 Sol on its developer site — a single, official recipe book for how to write prompts for the new model family. Read alongside this morning’s https://cuigh.com/posts/gpt-5-6-sol-production-migration-2026/ Ploy piece, it reads as the other half of the 7/9 launch story. The headline number is sharp: on OpenAI’s own internal coding-agent eval, swapping the system prompt from the GPT-5-era “encyclopedia” shape to a lean shape lifted eval scores 10–15%, cut total tokens 41–66%, and reduced cost 33–67%. That is not “the model is smarter so you can write shorter”. It is the reversal of prompt-design philosophy itself, now that model capability has caught up. ...

July 13, 2026 · 7 min · cuigh

GPT-5.6 and ChatGPT Work: OpenAI is bundling its strongest model with its strongest agent, betting on the same thing

On July 9, OpenAI did two things on the same day: pushed GPT-5.6 to the top of the “strongest model” list, and shipped ChatGPT Work — an agent that can run across web, mobile, and desktop to complete hours of work for you. Sam Altman posted on X: “obviously the best model we have ever produced, but also one of the best blog posts we have ever produced.” That line looks like product PR. It is not. What Altman is actually saying is that neither the model nor the blog post is the decisive asset anymore. An agent that can run autonomously for hours is. ...

July 10, 2026 · 6 min · cuigh

GPT-5.6 Sol shipped with a built-in ID check. Here's what changes for the rest of us.

On the afternoon of June 26 (U.S. Eastern), OpenAI published a preview post for GPT-5.6 Sol on its blog, and parked the system card at deploymentsafety.openai.com. The same day, the Washington Post reported that the U.S. government will decide which users get access to GPT-5.6. The same day, Reuters reported Anthropic is releasing Mythos only to “trusted partners.” Three events, less than 24 hours apart, each individually heavy. I am not going to write “AI regulation is here” — that ground was covered on June 14 in AI regulation just walked into the product. What I want to do today is zoom in much closer: what does GPT-5.6 Sol’s launch actually change, for a normal user, a normal builder, a normal product spec writer. ...

June 27, 2026 · 7 min · cuigh

Codex and Claude Code really did split: what a 350-like migration tweet tells us about two ecosystems quietly trading shifts

If you only read feature comparison tables, Codex and Claude Code differ in thirty places: model choice, browser support, rate limits, phone handoff. That is not what I want to write about today. Today I want to write about a 350-like tweet, and the two core maintainers standing next to it. The tweet that broke the silence Peter Yang (@petergyang) writes the Practical AI tutorials newsletter, 150K+ subscribers. At 3:32 AM UTC on June 20, he posted: ...

June 21, 2026 · 7 min · cuigh

AI regulation has landed: the 48 hours that moved it from slideware to product

On the evening of June 12 (US Eastern), the Wall Street Journal dropped a scoop. By the next morning, three stories occupied the top of Hacker News, The Verge had its own follow-up, and Times of India had pinned down the causal chain. This is not another round of “AI labs versus the regulators” commentary. Three things happened inside a 48-hour window, and they each hit a different layer of the two leading American AI companies. ...

June 14, 2026 · 9 min · cuigh

Apple's Three-Announcement Night: What It Means When Gemini Sits Behind Siri

Within the first 90 minutes of WWDC 2026, Apple did something very unlike Apple. It did not define a new device category on its own, did not reinvent an interaction paradigm on its own, and did not lock the entire supply chain inside its own walls. Instead, it let one of the two largest companies on earth, fold the model of one of the largest AI labs on earth, into its most strategic product. ...

June 9, 2026 · 11 min · cuigh

S&P 500 just shut the door on SpaceX — and took OpenAI and Anthropic with it

The bigger story is not the headline. On June 4, 2026, S&P Dow Jones Indices declined to bend its rules for SpaceX. SpaceX had asked for unusually swift entry into the S&P 500 on the strength of an “unprecedented market capitalization.” The answer was no: you are not yet profitable. Read in isolation, this is just a SpaceX IPO story. The interesting part is who got caught in the same net. By refusing to carve out an exception, S&P also closed the fast-track lane for OpenAI and Anthropic. There will be no six-month seasoning period, no profit waiver, no IWF shortcut for unprofitable AI mega-caps. The “use your size to skip the line” playbook is now off the table. ...

June 7, 2026 · 8 min · cuigh

Codex Mobile: Turning Your Phone Into a Remote Control for AI Coding Agents

OpenAI recently brought Codex into the ChatGPT apps for iPhone, iPad, and Android. At first glance, that sounds simple enough: now you can use Codex on your phone. I do not think that is the interesting part. Honestly, who wants to review code diffs on a phone for half an hour? That sounds more like punishment than productivity. The real point of Codex Mobile is that it turns the phone into a remote control for AI coding agents. The code still runs on your laptop, Mac mini, devbox, or remote environment. The phone is there for checking progress, answering questions, approving commands, changing direction, and dropping in a new task when an idea shows up. ...

May 15, 2026 · 7 min · cuigh

Voice Is Eating the Prompt: How Ordinary People Will Talk to AI Next

For the last two years, the default image of using AI has looked something like this: someone sitting in front of a text box, carefully writing a prompt as if they were briefing a very literal genie. That image is not wrong. In fact, it defined an entire phase of AI product behavior. But lately I have felt more and more strongly that this interface is starting to loosen. Prompting is not suddenly useless. It is simply moving into the background, while voice, screenshots, screen recordings, and raw documents are slowly becoming the more natural front door. ...

May 6, 2026 · 8 min · cuigh