Skip to content

Update plugin-builder guidance for Opus 5.5 - #219

Draft
mxriverlynn wants to merge 9 commits into
mainfrom
opus-5-5-plugin-builder-guidance
Draft

mxriverlynn wants to merge 9 commits into
mainfrom
opus-5-5-plugin-builder-guidance

Conversation

@mxriverlynn

Copy link
Copy Markdown
Collaborator

Summary

han-plugin-builder taught authors to write for Sonnet 5, Opus 5, and Fable 5, as checked on 2026-07-31. Opus 5.5 has shipped, and it changes what a skill or agent author should write. This PR evaluates the whole plugin against Getting the most out of Opus 5.5 (primary). It also uses three baseline sources: How we use skills, The new rules of context engineering, and Anthropic's Prompting Claude Opus 5.5. It then applies the changes that survived review.

Every candidate change went through adversarial validation. 28 went in and 18 entries came out. The plan, its decision log, and the evidence are in docs/plans/opus-5-5-plugin-builder-guidance/.

What changed

Opus 5.5 facts (per-model-authoring.md, plus the four files that repeat its model list)

  • Opus 5.5 is named, and inherits the Opus 5 guidance except where a section says otherwise.
  • The warning against asking the model to show its reasoning now covers Opus 5.5 (reasoning_extraction refusals).
  • Opus 5.5 can't turn thinking off.
  • Effort names mean more thinking on Opus 5.5, so re-test a pinned effort value. Both frontmatter tables link to this.
  • Vague review limits ("only report high-severity issues") are still discouraged. A limit that names a checkable bar and asks for evidence on each item is fine.

Long runs and delegation

  • workflow-patterns.md: name the stops an unattended stretch should and shouldn't make. Keep that instruction out of any skill built to hand control back. Give a step that works through an unknown number of items a finish line and a checklist file.
  • multi-agent-economics.md: check each fanned-out subagent's evidence before accepting it. Don't finish while a run_in_background dispatch is still running.
  • Agent graceful-degradation.md: research agents mark claims they couldn't confirm and say where they looked.

From the baseline articles

  • A limit on "be specific" (writing-effective-instructions.md).
  • Instructions that restate what the model already does count as removable.
  • A gotchas category for reference files.
  • How to build a skill that verifies a running product.
  • Skill state lives in ${CLAUDE_PLUGIN_DATA}, never in the skill folder, because updates replace that folder.
  • Helper-script libraries.
  • Hooks scoped to a single skill's run.

Builders

  • skill-builder and agent-builder route to per-model-authoring.md when a target model comes up.
  • Their Step 6 reviews check for leftovers tied to a specific model.

The rule index (assets/rule-index-body.md) describes every widened file.

Behavior changes (all approved during planning)

  • The review-limit advice now allows a checkable limit (S-5).
  • "Be specific" now has a stated limit (S-11).
  • The builders read the per-model guide, and their final review checks one more item (S-18).
  • agent-builder's review checks for unconfirmed-claim marking, and skill-builder's review cuts restated model defaults (S-10, S-12).

Notes for review

  • Departure from the plan. The stop-naming section describes collaborative stops instead of linking han-core/references/collaborative-stop-rule.md. guidance init copies this guidance into other repos, where that link would break.
  • Left out on purpose. The state section doesn't say whether saving to ${CLAUDE_PLUGIN_DATA} triggers a permission prompt, because no source I used states it.
  • Consuming repos. Repos that ran /guidance init pick this up on /guidance update. Mention that in the release notes.
  • Out of scope. Cut: re-checking the Fable 5 guidance against Fable 5.1. Deferred: a full hooks guide.

Verification

  • npx prettier --check han-plugin-builder passes. The full npm run lint hook set didn't run locally because dev dependencies weren't installed.
  • Every relative link in the plugin resolves, except three pre-existing placeholder links in templates/plugin-readme-template.md.
  • No file still carries the old three-model list, including where the sentence wraps across lines.
  • /han-update-documentation ran on the branch and found one issue, a blocklisted word, which is fixed.

mxriverlynn and others added 9 commits September 23, 2026 08:24
han-plugin-builder teaches authors to write for Sonnet 5, Opus 5, and
Fable 5 as checked on 2026-07-31. Opus 5.5 has shipped, and it refuses
reasoning-echo instructions, cannot turn thinking off, reads effort
levels differently, and stops long runs early to report.

Evaluate the whole plugin against the Opus 5.5 article (primary), the
"How we use skills" post, the context-engineering post, and Anthropic's
Opus 5.5 prompting page. Twenty-eight candidates went through
adversarial validation; eighteen change entries survive, in five units:

- Per-model facts: add Opus 5.5 to per-model-authoring.md and its four
  mirrors, extend the reasoning-echo warning, the thinking clause, the
  effort re-test note, and separate vague review limits from a checkable
  bar.
- Long runs and delegation: name the stops an autonomous stretch should
  and should not make without touching collaborative stops, bound
  unbounded steps with a finish line and a task file, check fanned-out
  evidence, and wait for background dispatches.
- Instruction style: a limit on "be specific", a removal bullet for
  instructions restating model defaults, and a gotchas category.
- Baseline skill patterns: verification skills, skill state kept in
  ${CLAUDE_PLUGIN_DATA}, helper-script libraries, and a skill-scoped
  guardrail hook.
- Builder routing: skill-builder and agent-builder read and review
  against the per-model guide.

The operator approved all four behavior-changing decisions. The plan,
its decision log, the current-state findings, and the scope boundary
live in docs/plans/opus-5-5-plugin-builder-guidance/. No plugin file
changes in this commit.
The per-model guide covered Sonnet 5, Opus 5, and Fable 5 as checked on
2026-07-31. Opus 5.5 has shipped with behavior that changes what an
author writes.

- Name Opus 5.5 and state that it inherits the Opus 5 guidance except
  where a section says otherwise, per Anthropic's Opus 5.5 page.
- Extend the reasoning-echo warning to Opus 5.5 and name the
  reasoning_extraction refusal category.
- Note that Opus 5.5, like Fable 5, cannot turn thinking off.
- Tell authors that effort names mean more thinking on Opus 5.5, and to
  re-test a pinned effort value rather than carry it over.
- Keep the warning against vague review limits, and add that a limit
  naming a checkable bar with per-item evidence is fine.
- Add the Opus 5.5 sources and a Contents list.

Mirror the model list in the guidance router, the portable router, the
rule index, and the specialization guide, and point both frontmatter
tables' effort rows at the effort guidance.

Plan: docs/plans/opus-5-5-plugin-builder-guidance (Unit 1, S-1 to S-6).
Opus 5.5 works longer on its own and sometimes ends a turn with a
progress report instead of the next step. The guidance covered
deliberate human gates, but not that.

- workflow-patterns.md: name the stops an autonomous stretch should and
  should not make, and keep the keep-going instruction out of any skill
  built to hand control back. Bound a step that works through an
  unknown number of items with a checkable finish line and a checklist
  file that survives context summarization.
- multi-agent-economics.md: when fanning out, check each subagent's
  cited evidence before accepting its finding, and do not finish while
  a run_in_background dispatch is still running.
- agent graceful-degradation.md: a research or analysis agent marks a
  claim it could not confirm and says where it looked.
- per-model-authoring.md: name Opus 5.5's early-stop behavior and link
  the pattern.
- agent-builder Step 6 item 8 names both degradation rules, and the
  rule index describes the widened files.

The stop-naming section describes collaborative stops rather than
linking han-core's collaborative-stop rule, because this guidance is
vendored into other repos by `guidance init`, where that link would not
resolve.

Plan: docs/plans/opus-5-5-plugin-builder-guidance (Unit 2, S-7 to S-10).
Anthropic's skills and context-engineering posts report that Claude 5
generation models are over-constrained by blanket rules, that
restating default behavior adds tokens without changing output, and
that a gotchas list is the highest-signal content a skill carries. The
guidance pushed toward specificity with no limit, removed only
toolchain-enforced rules, and had no gotchas category.

- writing-effective-instructions.md: spend specificity where a miss
  breaks something, and leave out rules for details the model gets
  right unprompted. The rule governs instruction detail, not whether a
  step is fixed, and links the entity taxonomy's flowchart test.
- progressive-disclosure.md: list instructions that restate model
  defaults as removable, and name gotchas as a reference category.
- skill-reference-files.md: define gotchas as failure points found in
  use, distinct from a checklist or a canonical example.
- Rule index describes the widened files.

Plan: docs/plans/opus-5-5-plugin-builder-guidance (Unit 3, S-11 to S-13).
…s, and skill hooks

Anthropic's skills post describes four patterns the guidance lacked.

- success-criteria-and-testing.md: how to build a skill whose job is
  verifying a running product: pair it with a driver, assert state at
  every step, and record the run.
- plugin-json-options.md: keep a plugin-shipped skill's first-run
  answers and run history in ${CLAUDE_PLUGIN_DATA}, never in the skill
  directory, which ${CLAUDE_PLUGIN_ROOT} replaces on update. Name
  userConfig, hand-edited files, and repo-checked-in skills as the
  cases that need something else. allowed-tools-AskUserQuestion.md
  links to it.
- hardening-fuzzy-vs-deterministic.md: a library of helper functions
  the model composes into one-off scripts, which still prompt for
  permission. Adds a Contents list now that the file is over 100 lines.
- skill-frontmatter-fields.md: a skill-scoped hook lasts only for the
  run, which suits a guardrail such as blocking destructive commands.
- Rule index describes the widened files.

The state section does not say whether saving to ${CLAUDE_PLUGIN_DATA}
prompts for permission, because no source this plan used states it.

Plan: docs/plans/opus-5-5-plugin-builder-guidance (Unit 4, S-14 to S-17).
…r-model guide

Neither builder read per-model-authoring.md or reviewed against it, so
the Opus 5.5 guidance reached only people who asked the guidance skill
directly.

- Both decision tables route "instructions for a named target model"
  to per-model-authoring.md.
- Both Step 6 reviews gain a model-specific-leftovers item: no "think
  step by step" line, no instruction to reproduce reasoning in the
  reply, no pinned effort without a stated reason, and no vague
  limiting phrase in a review step, with one passing and one failing
  example quoted.
- skill-builder's progressive-disclosure item also cuts instructions
  that restate a default the model already follows, matching the
  guidance change in progressive-disclosure.md.

Also rewraps the specialization guide's per-model cross-reference,
which the Opus 5.5 addition pushed past 120 characters.

Plan: docs/plans/opus-5-5-plugin-builder-guidance (Unit 5, S-18).
…ection

Found by the /han-update-documentation pass over the branch.
…lder docs

The skill-builder and agent-builder long-form docs now say the review pass
strips model-specific leftovers and that per-model authoring is a governing
document in the design-tree map, matching the SKILL.md changes on this branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants