danielkov

cat writing/cli-vs-mcp.md

CLIs are for laptops. MCP is for fleets.

CLIs and MCP solve different problems. Pick the one that matches where your agent is running.

7 min read1,531 words

The tech Twitter argument about MCP vs CLI is exhausting. Same takes, same fights, zero signal.

Here’s the actual distinction: CLIs work when the agent runs on your machine. MCP works when the agent runs somewhere else, at scale.

I’ve shipped both at Speakeasy - MCP servers in ephemeral cloud runtimes for our assistant product, plus daily CLI usage for my own work. The difference isn’t theoretical. It’s about whose machine knows what.

CLIs: your machine, your auth, your problem

Local agents and CLIs pair well because the hard stuff is already solved.

You’re already logged in. The CLI has your credentials. The agent gets whatever permissions you have. Scope it down if you’re worried (you should be).

The February 2026 DataTalks.Club incident is instructive: a Claude Code session ran terraform destroy on production, nuking 2.5 years of data. Alexey Grigorev’s autopsy is public - the credentials the CLI held were the credentials the agent used, destructive permissions included. AWS Support recovered the database after 24 hours.

Similar story from July 2025: Replit’s Agent wiped Jason Lemkin’s production database during an explicit code freeze, then fabricated 4,000 rows of fake data and lied about rollback being impossible. Replit’s CEO acknowledged it.

These aren’t CLI failures. They’re permission scope failures. The CLI is doing exactly what it’s supposed to.

Debugging is cheap. Real binaries give real errors. Wrong PATH, expired token, missing dependency - you’ve fixed these before. The shell tells you what broke.

One interface for humans and agents. The same git status you run, the agent runs. You can watch it, take over, hand back. No translation layer.

Models have seen this before. Public corpuses are full of CLI patterns. Smaller models often handle CLIs better than novel protocols because the patterns are in their training data.

Counterpoint: CLIs are inconsistent. Every tool has different flags, error formats, output shapes. MCP servers declare schemas upfront - explicit contracts, structured errors. Smaller models hallucinate flags (--recursive vs -r vs -R) more often with CLIs.

For local use, I still pick CLI. A hallucinated flag costs a retry. A unified human/agent interface pays off daily.

Where CLIs fall apart: cloud runtimes

At Speakeasy we run assistants in ephemeral cloud environments - Firecracker microVMs that spin up, do the work, die. No persistent filesystem, no state between tasks. Efficient, parallel, secure.

This is where CLIs stop making sense.

Provisioning doesn’t scale

Installing a CLI on your laptop: one-time cost. Installing CLIs on thousands of ephemeral runners that exist for seconds: different problem. Every binary adds image weight, cold start latency, failure surface. Multiply by N tools per customer and you’re accidentally maintaining a Linux distro.

Runtime dependencies will bite you

Not an “Alpine vs glibc” story. It hits everyone.

Anthropic has an open issue on their own Claude Agent SDK: “Linux: musl binary preferred over glibc in v0.2.116 CLI auto-discovery.” The SDK ships a musl binary; auto-discovery picks it on glibc systems; users get:

FAILED: Claude Code native binary not found at
  .../claude-agent-sdk-linux-x64-musl/claude.

File exists. Dynamic loader doesn’t. The ELF header points at /lib/ld-musl-x86_64.so.1 but the system has /lib64/ld-linux-x86-64.so.2. AWS ECS Fargate users hit this with standard node images.

If Anthropic can’t get this right for their own CLI, good luck with ten third-party tools.

More examples: Distroless images with glibc 2.31 choking on binaries built against 2.32. CGO-enabled Go binaries in distroless/static failing with “no such file or directory” because libc.so.6 is missing. OpenSSL 1.1 vs 3.0 causing intermittent TLS failures.

Each has a workaround. Aggregate across every CLI your agents might call and you’re doing unpaid distro maintenance.

Authentication is the real blocker

Most CLIs store credentials in plaintext: ~/.aws/credentials, ~/.config/gh/hosts.yml, ~/.kube/config. On your laptop, fine. The OS isolates you; you trust your binaries.

In an ephemeral assistant runtime, this model breaks. The agent is the process reading the filesystem. There is no user - there’s an LLM that can be prompted to read any file it has access to. Few CLIs support secure keychains; almost none support remote keychains, which is what you need for per-task, per-tenant credential injection.

MCP changes this because auth is in the spec.

Current MCP auth mandates OAuth 2.1, PKCE, Resource Indicators (RFC 8707), Protected Resource Metadata (RFC 9728), plus Client ID Metadata Documents and step-up authorization from the November 2025 revision.

What this gets you:

  • Tokens are server-bound. A token for https://mcp.example.com/files fails at https://mcp.example.com/email. The resource parameter is required; confused-deputy attacks are explicitly prevented.
  • No token passthrough. MCP servers must mint their own upstream credentials. User tokens don’t leak past the boundary. The spec says it: “MCP servers MUST NOT pass through the token it received from the MCP client.”
  • Step-up scopes. Agent starts with minimal permissions; server returns 403 with WWW-Authenticate if more are needed; user consents in browser.
  • Per-user identity. OAuth 2.1’s authorization-code flow means every token represents a consenting human, not a shared service account.

At Speakeasy this maps directly to how we run. Every invocation is bound to the initiating user. The control plane pushes short-lived scoped tokens to the assistant at call time - not pulled from files or env vars, but received in memory. They exist only for the call duration. Exfiltrating one would require memory-inspection tools the runtime doesn’t have.

User permissions and approved integrations define what the assistant can do. Every action is authenticated as the user, logged in the same audit trail. None of this works if auth state is a YAML file in $HOME.

What MCP buys you in production

Once you’re running agent fleets for enterprises, requirements shift from “does it work” to “will security allow this.”

Observability. MCP uses JSON on the wire with defined shapes. Log it, sample it, redact fields, replay it. Put a proxy in front and you lose nothing. Try that with a CLI: parsing argv, scraping stdout, hoping nobody passed --quiet.

Auditing. Identity comes from the auth flow, not whoever’s logged into the runner. You know which user, which session, which tenant. Compliance teams need this; CLIs can’t provide it because they were built for single humans at terminals.

Control. Disable a tool org-wide? Flip it in the registry. Done. Do the same for a CLI? SSH to machines, uninstall binaries, update images, redeploy. By the time you finish, the thing you were stopping already happened. With MCP, if the admin didn’t connect the integration, the agent cannot use it - the capability doesn’t exist in the runtime.

These aren’t exotic needs. They’re table stakes for regulated industries. You wouldn’t ship SaaS without SSO and audit logs; you can’t ship assistants without tool layers that support the same.

What this unlocks

The piece that surprised me, building this out, is how the value compounds.

Every MCP source you configure for your org doesn’t just add one capability to one assistant. It adds it to your entire fleet, immediately. Connect a Linear MCP server and every assistant in the org can read tickets. Connect Datadog and every assistant can query metrics. Connect your internal customer-data MCP and every assistant can ground its answers in real customer state.

You go from “I have an assistant that can do X” to “I have a workforce of assistants that can each do X, Y, Z, and combinations of them I didn’t anticipate.” A few concrete shapes:

  • A support assistant that pulls the customer record, checks recent error rates in Datadog, drafts a reply, and files a Linear ticket if it spots a real bug—without a human stitching those tools together.
  • A release assistant that takes the kind of commit messages engineers actually write—fix: don't NPE on empty filter, chore: bump deps—and turns them into a changelog a customer wants to read, grouped by theme, framed by impact, ready to ship.
  • A research assistant that hits twenty internal data sources in parallel, because the runtime is ephemeral and parallelism is governed but cheap, and synthesises a brief in the time a human would spend logging into the first dashboard.

None of that works if your tool layer is a shell. It works because MCP is a protocol designed to be operated, observed, and authenticated as a system.

So which one should you pick?

Both. They’re not the same product.

If you’re building a coding agent that runs on a developer’s machine and acts as their pair programmer, lean on CLIs. The agent is in the developer’s environment, with the developer’s permissions, debugging the same tools the developer already uses. Don’t reinvent that with a protocol.

If you’re building a hosted assistant that runs in your infrastructure on behalf of your customers and integrates with their stack, MCP is the right primitive. Provisioning, auth, observability, and control are the hard parts of operating an agent product, and the protocol is designed for them.

The mistake isn’t picking one. The mistake is treating “CLI vs MCP” as a debate about which is better in the abstract, when the actual question is: whose machine is the agent running on, and what does that machine know?

Pick the interface that matches the answer. Then go build the thing.