cat writing/skills-are-not-the-right-abstraction-for-agentic-knowledge.md
Skills are not the right abstraction for agentic knowledge
Skills help, but for reliable agent workflows they're the wrong primitive. Move workflow logic into the tool.
If you’ve used coding agents for more than a week, you’ve seen this:
- Model misses the skill it should use.
- Invents a command that doesn’t exist.
- Doubles down into a dead end.
I shipped skills at Speakeasy when Anthropic introduced them in October 2025. We published our first Claude Code skills in February 2026. They helped - removed some friction.
But for reliable, repeatable workflows, skills are the wrong abstraction for long-lived agent knowledge.
Skills in brief
Skills are markdown files with structured frontmatter, per the Agent Skills Open Standard. Claude Code, Cursor, Codex support them.
At startup, agents typically see only frontmatter (name/description). They pull the full body later if activated.
The ecosystem moved fast
- Anthropic made skills available, pushed an open standard in December 2025.
- Vercel launched open skills January 2026.
- Vercel published evals January 2026 showing
AGENTS.md-style guidance outperforming skills.
Speed is good. The abstraction issues remain.
Where skills fail
1. Reproducibility is weak
Skills are conditional context injection. Efficient but non-deterministic. Two runs with the same prompt diverge because activation is heuristic.
“Good enough most of the time” is fine for some use cases. It’s not fine if you need same inputs, same path, same result.
2. Activation logic is brittle
People write skill descriptions like READMEs. Models activate better from state cues than feature explanations.
Example from my project Granary - the finalize skill description is blunt state-based routing: “IMPORTANT! you must use this skill when a granary project is complete”.
This activates better than “this skill helps with release workflows because…” prose.
Bad (backstory-heavy):
name: write-release-notes
description: Our release notes are often short commit messages and not proper product changes, so this skill exists to improve quality...
Bad (rationale-heavy):
name: openapi-spec-writer
description: This skill helps teams author OpenAPI specs in a maintainable way because standardisation and governance are important...
Better (state-based):
name: write-release-notes
description: I just merged changes and must produce customer-facing release notes
Better (state + hard cue):
name: openapi-spec-writer
description: user asked for a new or updated OpenAPI document and no valid spec file exists yet in the repository
3. Cross-model behavior drifts
Each agent stack has different instruction-file conventions. Granary has code to detect CLAUDE.md, AGENTS.md, CODEX.md, Cursor/Windsurf/Cline variants, Copilot files, etc..
That utility shows the fragmentation. If activation semantics were portable, you wouldn’t need this much surface-area mapping.
4. Teams use skills as commands
Skills are optional/conditional context. Teams treat them as guaranteed procedural steps. When you expect command semantics from probabilistic activation, you get weird failures.
AGENTS.md vs skills
Vercel’s eval post is the best public comparison:
AGENTS.md: 100.00%- Skills: 79.17%
- No guidance: 83.33%
AGENTS.md led on process-heavy tasks (codebase understanding, completion flow, output format). Skills did well on targeted capabilities like git operations.
Small benchmark - treat as directional. But it matches production experience:
- Skills work for reusable capability injection.
- Deterministic runbooks work better for reliability-critical workflows.
Security concerns
Public marketplaces make trust a first-class problem.
Promptfoo launched OpenClaw/ClawHub; Koi Security reported malicious packages and exfiltration patterns. There’s a public advisory for command-injection risk.
Anthropic’s docs note that skill commands run with your process permissions.
None of this means “don’t use skills.” It means you need supply-chain discipline, provenance checks, and permission boundaries before trusting broad skill ingestion.
What works better: move guidance into the tool
I moved Granary from “skills hold workflow logic” to “the CLI itself is the workflow interface.”
Examples from the public repo:
- Running bare
granarygives workflow entrypoints:plan,initiate,work,search. - Commands include
AGENTS:hints in help text:projects,tasks,task,start. granary planoutputs a scaffold with the hard rule that “Task descriptions are the ONLY context workers receive”.granary work startis fail-fast and explicit: blocked dependencies, claim conflicts, draft-state rejection, structured completion paths.- Prompts carry machine-usable next steps:
selection_reason,<next_actions>blocks.
Key examples:
name: finalize
description: IMPORTANT! you must use this skill when a granary project is complete
AGENTS: To work on a task with full context, use:
granary work start <task-id>
## When Done
granary work done <task-id> "summary of changes"
## If Blocked
granary work block <task-id> "reason for blocking"
Execution guidance is versioned with the tool, not floating separately as a skill pack.
Why this beats skills-first
Distribution and versioning are coupled. When workflow guidance ships in the CLI, command surface and guidance evolve together. No drift where “agent thinks command exists, CLI version says it doesn’t.”
Observability is easier. Granary tracks runs, retries, statuses, logs in first-class flows. Natural place to add telemetry without external skill instrumentation.
Works across agents that can execute commands. Keep skills for convenience. Critical path no longer depends on one vendor’s activation semantics. The command interface becomes the portability layer.
Internal eval results
I ran a controlled comparison: same model (Claude Code with Opus 4.5), same repo, 25 runs each. Only variable was workflow abstraction.
| Setup | Command failure rate | --help calls (avg/run) |
Eval success | Tokens vs baseline |
|---|---|---|---|---|
| No entrypoints + skills | 7.5% | 0.52 | 24/25 (96%) | 197k/run baseline |
| Entrypoints + root bootstrap | 1.75% | 0.04 | 25/25 (100%) | Input -9%, output -4.6% |
Setup 1 failures: 2/25 runs where skill didn’t activate immediately. One recovered via --help traversal (9 calls in that run). One didn’t recover - that was the single eval failure.
Workflow entrypoints eliminated the eval miss, dropped command failures by 76.7%, and cut --help dependency by 92.3%.
Skills aren’t dead
I still use them for:
- Narrow behavior nudges
- Bootstrapping new users
- High-frequency shortcuts
If you care about reliability:
- Tool-embedded workflow logic (deterministic, observable, versioned).
- Skills as lightweight accelerators.
- Free-form prompting for exploration.
Writing better skills anyway
When you write skills, optimize for activation state.
Don’t:
- Explain backstory
- Narrate rationale
- Over-document internals
Do:
- Describe the exact state when the model should activate
- Keep it concise
- Include hard markers like
IMPORTANT: - Iterate against real activation traces
Example:
description: IMPORTANT: I must use when a task is complete
Counterintuitive, but works more often.
Bottom line
Skills patch agent frustration. For production reliability, executable workflow guidance in your tool’s interface is the better abstraction.
Granary’s LLM-first redesign proposal states this: command output should be enough for agents to complete workflows without external skill dependency.
That seems right.