Do the Skills Earn Their Keep?
The skills my agents rely on are knowledge artifacts — so I measure whether they earn their keep: coverage, reliability signals, and skill-to-outcome quality. Evidence over vibes.
How the Hahn-Solo agentic pipeline actually runs — methodologies that produced the case studies.
The skills my agents rely on are knowledge artifacts — so I measure whether they earn their keep: coverage, reliability signals, and skill-to-outcome quality. Evidence over vibes.
I built UI from the prose spec and it drifted — stacked headers, missing decorative elements. The fix: treat the approved clickdummy as a binding contract, enforced by a screenshot diff before review.
Move the moment of truth forward. Before deep planning, run a two-day probe against a real tenant — and let it cancel the PRD if the foundation isn't there. Cancellation is a successful outcome.
I tried to run the agentic pipeline end-to-end on a four-SDK-surface PRD. Tests stayed green. The tenant didn't. The fix wasn't more autonomy — it was phasing the work, with checkpoints between.
I called them skills. The runtime never did. Three iterations, one cutover commit, and 46 markdown fragments finally became something a fresh agent could actually discover.
Ten numbered runs hardened Marketplace skills; Run 10 (PageShot) shipped first in Pages. Later apps dogfed the same learnings—patches, scale proof, or regression when skills already matched.