danielkov

cat writing/please-fix-your-agent-skills.md

Please fix your agent skills

Your agent skills suck. They cost you time and money. Fix them.

7 min read1,491 words

It’s a skill issue

Agent skills are an open standard for just-in-time context augmentation in LLM harnesses. A shortcut to instilling specific behaviour. You give the agent a task - it activates a skill. Except sometimes it doesn’t. Here’s how you fix that.

Anatomy

Skills are simple text files in markdown format. They have two parts:

  • frontmatter: holds metadata
  • body: raw skill contents the model sees after activation

A skill might look something like this:

name: write-blog-post
description: use only when your task requires the creation or modification of a blog post
---

## Tone of voice

...

The name

Name needs to be unique. Most harnesses apply a precedence-based rule to exclude duplicates. Name your skill well and this won’t be a problem.

Description

Dangerously misnamed attribute. Treat this as the trigger condition. This is not for you. It’s for the model. It’s the only way it knows whether it should read the rest of the skill. If you don’t give a skill a good description, the agent won’t use it, your skill becomes useless. More on this below.

Contents

The model only sees the content or body of the skill after it activates it.

If you read that and thought “no shit” you’ll be surprised to learn how many popular skills carry “When to use” or other useless narrative baggage. Skill are usually activated at the beginning of a session. This means that the token cost of the skill’s contents carries onto every subsequent model call. For the skill to earn its keep, it needs to offset that cost. So what should go into a skill?

Context

The agent sees something like this:

| name          | description                                        |
| ------------- | -------------------------------------------------- |
| pr-review     | Always use before reviewing pull requests          |
| write-like-me | Use ONLY when user asks to use their tone of voice |

Most harnesses either provide a dedicated tool for skill activation, e.g.: Claude Code’s Skill(<name>) or give the model the full path and nudge it to read the skills using built-in tools, e.g.: Codex read(<path-to-file>).

The best skill is the one you never write

Skills are criminally overused. Popular uses of skills that do more harm than good:

  • negative reinforcement, e.g.: pleading with Claude not to spew walls of slop: pick a model that works better for your use case / taste
  • how-to guides for MCP server tools: if the tool schema and description aren’t good enough, fix those instead
  • standards and conventions: enforce them instead - a skill is never prescriptive - use linters, formatters, test and CI
    • this also forces a split-brain problem - one more artifact that needs to be kept up to date
  • make a dumb model smarter: this never works, there are better ways to cut cost - this isn’t one
  • mechanical task instructions: write a script once, don’t use LLMs for scriptable workflows
  • commands: yes, Claude lets you active skills as a slash command. You only do this, because your description is inadequate.

Writing good skills

In my view there are two paths to authoring a good skill. Before you do, make sure you’ll have the bandwidth to maintain. An out-of-date skill is often worse than nothing.

Hand-write like it’s 1995

Some of the highest quality input tokens the model will see will come from a human. That human may be you. It’s worth investing your time.

The types of skills that benefit most from this are the ones where output needs to follow guidelines. Instead of:

Write like me...

actually take the time to spell it out. What makes post read like you wrote it? What are your quirks? What’s your style?

If you want the agent to write code like you, feed it examples you authored.

Rules to follow if you go this route:

  • keep it terse and concise
  • if you’re describing a process, make sure all steps are included, including contingencies
  • intricate processes require stop conditions - you don’t want your agent to run amok
  • try looking at things from the agent’s “point of view”, don’t rely on information you hadn’t shared with the model
  • positive exemplars are fine to use, but don’t overuse them. The model will overindex on specifics from your examples. Keep them as generic as possible.

I recommend frequently reviewing skill-driven outcomes. When an agent is following a skill and it gets lost, goes off script or does something undersired, I ALWAYS make sure the lesson gets encoded in the next version of the skill. There are tools to aid you in this. I’ve built the skills management feature in Speakeasy’s AI Control Plane with automatic incremental improvement at its core.

/skillify

This is my preferred way of crafting skills. Ever had a session go so well, you wish you could replicate it consistently? Now you can.

See the skillify meta skill I use
---
name: skillify
argument-hint: "[skill name or hint at which part of session to capture]"
description: >
  Turn workflow that just succeeded in this session into reusable skill.
  Trigger when user says "make this a skill", "save this workflow",
  "capture this as a skill", or asks that next session repeat what just
  worked without rediscovering it.
---

# Skillify

Distill finished work in this transcript into a skill. Source = transcript, not guesswork.

## Precondition — check before anything else

Need completed, successful workflow in transcript. Verify:

- Ran to end, result usable.
- Discovery happened — connections made, criteria settled, dead ends mapped, corrections applied.
- Pure one-shot answer / trivial edit → nothing to distill. Say so, stop.

Missing → state what's missing, ask which run to capture. Never invent a workflow.

Arg (`/skillify <name-or-hint>`) → scope to that thread. No arg → session's main thread.

Name collides with existing skill → update that file, don't duplicate.

## Extract

Walk transcript. Each step / roadblock / correction, triage:

- **Recurring?** Next run hits this again? No → drop.
- **Context fresh session lacks?** Yes → write it down: tool names, order of ops, auth shape, where things live, gotchas.
- **Dead?** Bug since fixed, one-off outage, deleted file, transient env break → drop. No posterity.

Keep: repeatable steps, live constraints, non-obvious decisions.

## User turns outrank yours

Mine every user message:

- Corrections ("no, do X", "actually…") → hard rules.
- Answers to your clarifying questions → those ARE the criteria. Bake in as defaults, or make skill ask them.
- Style / scope / output preferences → rules.
- Rejections → explicit do-not list.

## Ambiguity

Unclear whether something generalizes → ask user. Batch few concrete questions. No answer → pick default, note assumption inline. Never block. Never silently guess at user-supplied criteria.

## Strip

Skill must be publicly shareable. Remove:

- Names, emails, handles, org / customer / project names, ticket URLs with slugs
- IDs, keys, tokens, endpoints, hostnames
- Absolute local paths, machine names, home dirs
- One-off figures, dated numbers

Replace with placeholder or discovery step ("resolve from config / ask user"). Grep draft for leftovers before writing.

Keep: public tool + CLI names, generic conventions.

## Write

Skill-authoring meta-skill available in session (skill-creator or similar) → follow it for format. Else:

Frontmatter:

- `name`: kebab-case
- `description`: **activating condition only** — when to reach for it, plus trigger phrases. Not what it does, not how it works.

Body:

- Preconditions / required input first. Cheapest check first, fail fast.
- Then steps in run order.
- Terse. Fragments. No prose, no narration, no restating the task back.
- Exemplars: positive only, sparse, one where shape is hard to describe. Describe shape over pasting content — verbatim samples get echoed.
- Do NOT repeat activating conditions in body. Body only read after activation.

## Place

- Tied to one repo → `<repo>/.agents/skills/<name>/SKILL.md`
- Otherwise → personal skills dir, `<name>/SKILL.md`

Helper scripts beside SKILL.md, relative paths only.

Report path + description line. No process recap.

How to use it

  1. Steer or allow model to progressively navigate an ambiguous or complicated scenario
  • it’s fine if the model doesn’t get it right the first time, iterate until the outcome is satisfactory
  1. Tell agent to turn the session into a reusable skill
  • if you only care about a specific task, e.g.: generating an SVG logo, state that
  1. Review and scrutinise the outcome

Don’t trust the LLM to draw the right conclusions. LLMs are bad at creating skills. They’re rarely tuned for this sort of workflow.

Bottom line

Never start by asking an LLM to write you a skill. If it can be a prompt, it should be a prompt. If your agents aren’t using your skill, the description needs to be fixed. Skills should be the exception, not the rule.