I keep hearing the model labs tell a story that as models get better, you need to give the model less direction, and the coding agent harnesses around them should get simpler.
I like that story, it sounds so intuitive. But like a good engineer, I also do not quite trust it.
So I pulled together a small visual experiment around Claude’s published system prompts. The original version looked at Opus 4 through Opus 4.7. The chart now runs through Claude Opus 5.5, released September 22, with Fable 5 shown for context.
The short version: some prompt complexity really does move into training. But the total system does not simply get smaller. Complexity moves into safety policy, tool discovery, product context, memory, skills, and app-specific harnesses.
Updated September 23, 2026: Opus 5.5 reverses the recent dip. I recounted the published Opus 5 and 5.5 prompt bodies using the same rule. They went from 3,225 to 4,108 words, up 883 words or 27.4%. The increase is mostly in refusal guidance.
The interactive
The chart below is the experiment. Click a model tab to see the category notes, or switch the chart from absolute words to percent share.
What changed
words in the first Claude 4 prompt snapshot in this comparison.
words after safety, agentic behavior, and product scaffolding expanded.
words in the latest published Opus prompt body, using the same counting rule as Opus 5.
The curve went up, then down, then up again. Opus prompts grew from 1,714 words in May 2025 to 3,686 by April 2026. Opus 5 dropped to 3,225 on a fresh recount. Opus 5.5 rose to 4,108. That is the longest prompt in this comparison, and the new text is not spread evenly across the document.
The refusal section alone grew from 697 to 1,746 words. It now spells out boundaries for weapons, illegal drug protocols, copyrighted visual designs, and requests for recognizable characters. It also adds examples of original alternatives and a rule to ask a narrow question before declining some ambiguous requests. Anthropic removed 204 words of Fable safeguard-routing and default-stance text at the same time. The wellbeing and tone sections changed much less.
The pieces that moved
The behavior patches are the easiest part to understand. Older prompts carried little runtime hacks: count letters carefully, restate puzzle constraints, do not over-apologize, avoid certain linguistic tics. Those are exactly the sort of instructions you would expect to move into training as the model gets better.
The structural parts are stickier. Safety instructions remain large. Product and model identity keep changing as the product surface changes. Tool discovery becomes explicit in Opus 4.8: before Claude says it cannot do something, it is told to check for deferred tools, personal context, and skill files. Opus 5 and 5.5 do not carry that same dedicated block in their published prompt bodies. Opus 5.5 updates the product list, including Claude Tag and Claude Design, and moves the knowledge cutoff from the end of May to the end of June 2026.
This is Anthropic making product decisions. They are deciding what context and capabilities are visible, when to reveal them, and how Claude should behave when something might exist outside the visible prompt.
Why this matters
If you are building agents, you should not take the model labs at face value. You should measure what is actually changing.
This update makes the point sharper. Opus 5.5 is a stronger model, but its published claude.ai prompt is longer than Opus 5’s. Anthropic cut some older context while adding detailed rules elsewhere. Prompt length alone cannot tell you whether an agent harness got simpler.
Want to get hands-on?
If this framing is useful, the next step is to change a harness yourself and watch the trace move.
- Start with my hands-on course, Learn Harness Engineering with OpenHands. It walks through Agent Server and Agent Canvas, then turns model routing, retrieval, memory, security, critic loops, and goal scaffolding into runnable projects.
- The source repo is rajshah4/learn-openhands-harness, if you want to fork the exercises.
- The broader companion repo is rajshah4/harness-engineering, which collects the prompt-evolution experiment, references, and other small harness investigations.
- For the conceptual frame, the annotated talk is here: Harness Engineering: Why the System Around the Model Decides Agent Performance.
Data notes
- Opus 4 through Opus 4.7 use Simon Willison’s
simonw/researchmirrors of Anthropic’s published system prompts. - Opus 4.8, Fable 5, and Opus 5 were re-extracted from Anthropic’s published system prompt page on July 24, 2026.
- For the September update I compared Anthropic’s published Opus 5 and Opus 5.5 prompt code blocks. I concatenated each page’s blocks, stripped XML-like tags, then counted matches to
\b[\w'-]+\b. This gives 3,225 and 4,108 words. The earlier Opus 5 chart pass showed 3,221; use the paired recount for the percentage change. These are words, not model tokens. - Anthropic splits the Opus 5.5 prompt across two code blocks. The first block contains 1,877 words and ends at an open
<example>tag. The second begins with that example, adds 2,231 words, and closes</claude_behavior>. Copying only the first block produces an incomplete count. - The published prompts apply to claude.ai and Claude’s mobile apps. Anthropic says these updates do not apply to the API. Claude Code also has a separate prompt surface.
- Chart categories are editorial estimates based on primary function, not exact token accounting.