danielkov

cat writing/teaching-an-old-dog-new-tricks-how-to-teach-an-llm-a-new-language.md

Teaching an Old Dog New Tricks - How to Teach an LLM a New Language

Going against the grain is tough. Especially if you're fighting statistical models to do your bidding. Here's a story of how I taught an LLM a new programming language.

17 min read3,667 words

My knowledge cutoff point is…

LLMs are great. Especially if you’re doing usual work. They excel at tasks that are mathematically likely. To put another way: take all of human work that was ever done and recorded. Now train a model on that. It follows that the model will most likely excel at the task that had the most historical records in its data set. This means I can ask an LLM who won the 2021 Superbowl and it’ll “know” the answer.

It didn’t have to search, or use other tools. In its data set, the Buccaneers are mentioned around "2021 Superbowl winners", more than anything else. Statistically speaking, the most correct answer to my question is the Buccaneers.

This also works well when you give an LLM a task that follows a well-known shape or common guidelines, e.g.: asking it to build a custom take on a common SWE interview question FizzBuzz.

A task like this shows up thousands of times in its training data. Even though the words changed slightly, the model still “knows” what the likely answer was.

It also follows that it simply won’t know anything past its knowledge cutoff point. This is why models often:

  • don’t know who the president is
  • use outdated versions of software
  • can’t tell you the weather without using tools

From RAGs to riches

The usual answer to a knowledge cutoff is Retrieval-Augmented Generation or RAG. It’s the VC-approved term for saying: “Google Search for LLMs”. RAG usually injects knowledge into context with the user message or via fake tool calls. The model usually doesn’t ask for it. E.g.:

user: best thai restaurant in my area?

<additional-context>
- User is 30 years old
- User is software engineer
- User lives in Central London
</additional-context>

or

user: any typos in my article?

<additional-context>
open file: ./src/content/writing/teaching-an-old-dog-new-tricks-how-to-teach-an-llm-a-new-language.md
</additional-context>

Without this, the model would have to orient its worldview reactively, such as calling web search tools or asking the user.

RAG doesn’t replace reactive augmentation. Reading a file or searching the web is still useful. It works well when the missing piece is information: new API documentation, an internal process, or facts published after training finished.

Retrieving information isn’t the same as teaching. A model can read the syntax of a new programming language and still fall back to patterns it already knows from Python, JavaScript, or shell. Without thousands of examples in its training data to pull from, its statistical “instincts” start working against you.

Are you for RL?

When a AI lab, like Anthropic or OpenAI roll out with version 4.N of their model, it’s usually not an entirely new model. Labs take checkpoints and post-train them on examples of desired behaviour. With Reinforcement Learning (RL) the model’s output is scored by humans or deterministic tests.

Training makes desirable behaviour more likely and - if all goes well - less desirable behaviour less likely.

You can’t directly post-train a frontier lab’s closed-weight model yourself. Unless the lab releases its weights, you’re limited to whatever fine-tuning or reinforcement-tuning APIs it provides.

Runlet before you can walk

I couldn’t post-train the model, so I moved the training into the context window instead.

The thing I wanted it to learn is called Runlet. It’s a small orchestration language for agents. One program, submitted through a single tool call, describing a whole piece of work: list the tickets, fetch each one’s detail, filter, aggregate, write the results back, submit. The host checks it against the tool schemas before anything runs, executes it as a dataflow graph, and returns one structured value.

A conventional agent loop sends every tool result back through the model. 20 tickets means 20 round trips, 20 inference calls, and 20 ticket bodies permanently lodged in the transcript. Runlet collapses that into:

model → one program → many host tool calls → one structured result

There’s no async, no await, no thread pool. References create data dependencies, and anything without a dependency runs concurrently. Write this:

issues = linear.search_issues({ assignee: user, state: "open" })
pulls = github.search_pull_requests({ author: user, state: "open" })

pulls_with_checks = for pull in pulls {
    checks = github.checks(pull.number)
    return { number: pull.number, title: pull.title, checks }
}

return { issues, pull_requests: pulls_with_checks }

The Linear and GitHub searches start together, each checks lookup starts when its pull request lands.

Runlet is used in production - inside Kit - the agent harness I’ve built at Speakeasy. In Kit, the model only has access to a single meta tool, called compose. It passes Runlet scripts into compose and only sees the results or errors that come back. compose also supports Lua, but my nefarious scheme hedges on us using Runlet.

One reveal before we go further. The screenshots at the top of this article are from Kit, my coding agent. Kit’s model doesn’t get a read tool, or a shell tool, or an edit tool. It gets exactly one tool - compose - and every other capability lives inside that tool’s description as a callable Runlet name. Every FuzzBoop above was written by a model programming in a language that didn’t exist eight weeks earlier.

It took a bit of trial and error to get there though.

JavaScript flashbacks

My first instinct: write good documentation, put it in the tool description, let the model read it. The model read it. The model then wrote JavaScript.

Every mistake it made was a statistically probable guess. Here’s a real program, submitted by Kimi K3 on a revenue-reporting task, taken from a benchmark transcript:

cust_ids = fold acc = [] for row in detailed {
    if row.customer_id in acc {
        return acc
    } else {
        return acc + [row.customer_id]
    }
}

That looks like code. It also has 4 errors in it, because in Runlet if is not a statement, it’s an expression that has to be bound to a name. Here’s what came back:

runlet program rejected before execution; fix the errors and retry:

error RL1014 [control structures are expressions] at 598..600: `if` is not a
statement; bind the conditional's result: `x = if condition { ... } else { ... }`
- inside loops, `skip if condition` filters
error RL1017 [a block must return a value] at 703..704: write `return expression`
as the final statement
  fix: insert `return null` → `return null`

Here are some more examples of errors emitted during my benchmark runs:

# total = total + 1
error[RL2106]: duplicate binding
  `total` is already declared in this scope; bindings are immutable - merge both
  cases into one expression, bind a new name, or accumulate with a fold

# label = "has" if name else "none"
error[RL2305]: schema error
  condition must be Boolean

# n = xs.length
error[RL2103]: unknown property
  property `length` is not available on this value of type int[] - iterate or
  index the list itself

# r = await demo.task("a", 0, null)
error[RL1008]: unexpected token
  expected a newline or `;`

# function score(x) { ... }
error[RL1008]: unexpected token
  expected `=` after the binding name

Mutation. Truthiness. Method calls. await. function. Early returns (RL1012, “a block has exactly one final return”) showed up thirty times across the archived transcripts. These aren’t hallucinations, they’re priors. The model has seen xs.length a hundred million times and Runlet zero times. When the documentation and the training data disagree, the training data has better odds.

You cannot argue a model out of a prior. You can only give it something more specific to imitate.

A prompt response

The first version of the manual was what you’d expect: a dense, well-organised specification. Immutable bindings. No truthiness. Blocks end in return. Every rule stated clearly - micro-example as the cherry on top - plus one short correct program at the end. About 5.7 KB.

The second version threw the rules out and replaced them with one program - roughly seventy annotated lines exercising the entire grammar - pagination, concurrent fan-out with skip, fold aggregation with computed keys, block-if, retry boundaries, intrinsics, fire-and-forget writes. Every constraint moved out of the prose and into a comment placed exactly where a model would be tempted to violate it:

total = fold acc = 0 for row in shaped { # fold is THE way to aggregate: sequential reduce,
    return acc + row.amount              # the body's return becomes the next accumulator;
}                                        # `total = total + x` is an error - bindings are immutable.

Same token budget. Roughly 7 KB, around 1,700 tokens. Nothing was added - the rules were relocated to the scene of the crime.

Then I A/B’d them inside the agentkit compose benchmark, five scenarios, three repetitions, both languages. On claude-sonnet-4.5 the two forms were statistically indistinguishable; every apparent gap dissolved under replication. On claude-haiku-4.5 they came apart:

primer, haiku-4.5 suite cost mean accuracy parse errors across all transcripts
rules $0.626 0.72 RL1008 ×71, RL1017 ×58, RL1014 ×56
exemplar $0.301 0.83 RL1008 ×19, RL1012 ×6, RL1020 ×4

Half the cost, better accuracy, and the syntax-error storms collapsed about fourfold. The failure mode under the rules primer is comical to read in sequence: Haiku reads the rule, violates it, reads the diagnostic, violates it again. Under the exemplar it simply copies the shape it was shown.

Below a certain capability floor nothing works - gemini-2.5-flash failed under both forms, and one run burned 742k tokens iterating forty broken scripts before giving up. You can’t teach a fish how to climb a tree.

The exemplar is the default now. It’s also compiled against a mock registry in the crate’s test suite. The comments can still drift, but the substantial part is deterministically testable.

The other half of the manual isn’t in the prompt at all. It’s in the error messages.

Early calendar-scheduling runs produced a storm of generic parse errors - 37x RL1008 expected { in a single three-run cell - all traceable to one habit: models blending the two loop forms and writing fold acc = true for u in avail limit 32, importing limit from the concurrent loop into the sequential fold. The parser pointed at the wrong token and the model had no idea what it had done wrong. Adding one targeted diagnostic - “fold has no limit; fold iterations are sequential by definition”, with a removal fix-it - took that cell from $0.51 to $0.23 on the next run.

For a language whose only users are models in a repair loop, a targeted error at a predictable confusion point is worth more than a paragraph of specification. Design the diagnostics, not just the grammar.

Just-in-time training

The weights never change. Nothing is learned. The competence exists only inside the current context window and is gone the moment the conversation ends. The next session starts from the same statistical instincts that wanted to write xs.length.

Here’s how that context is assembled:

  • grammar comes from the primer - one annotated program, injected as the description of the compose tool
  • vocabulary comes from the tool catalogue that is visible for that turn, rendered as compact type notation (list_services(): string[], search_logs(...): { items: {...}[], total_pages: int }) instead of a JSON Schema
  • correction comes from the parser and analyzer. A rejected program returns the diagnostics verbatim - error code, span, message, fix-it, did you mean: candidates - and the model retries in the same turn

Corrections are our stopgap against runaway error loops. We’re combining RAG’s augmentation upfront with reactive steering. Let the model try, and then tell it not only why it failed, but also how to fix that failure.

There’s a fourth, quieter channel: a conservative auto-repair pass. When a submitted program fails to parse, the host tries a set of insertion-only rewrites - missing return, unbound statement, statement-form if, missing separator - re-parses, and if it now compiles, runs the repaired program and reports the corrections back as warnings. The model gets its result and the correction, without paying for a retry.

Put it on a spectrum: RAG gives the model facts it lacks. Post-training gives it instincts it lacks. This sits in between - more structured than retrieved documentation, temporary in a way that RL is not. You’re not teaching the dog a new trick. You’re holding up a card with the trick written on it, every single time, and correcting the dog when it gets it wrong.

Designing for the student

Nothing above works if the language is hostile. Runlet’s design brief was, roughly, what would a language look like if its only authors were models writing under a repair loop?

Look familiar, cost nothing to learn. Objects, lists, for, +, dot access, if/else - all borrowed. The surface that had to be taught is small, because most of the surface didn’t have to be.

Delete the ambiguous choices. No functions, no imports, no classes, no methods, no recursion, no while, no mutation. One return per block. Explicit booleans, no truthiness. There are no two ways to achieve the same thing in Runlet. The complete mental model fits in under 100 tokens.

Make the right choice the language default. The model never writes await, never spawns a task, never joins. It builds the data flow it needs. Concurrency is an execution detail. The host caps how many iterations run at once. Scheduling and fan-out are the runtime’s problem, not something the model has to synthesize.

That last point is where prompt design and language design turned out to be the same activity. for ... limit 32 used to be valid syntax in earlier Runlet versions. Then the benchmark transcripts showed models reading the number as an iteration cap: one run split forty orders into two loops and issued a second verification program, because exactly thirty-two happened to be complete. The number left the language entirely. for item in items now always processes the whole collection, and the host configures concurrency. Removing that one number cut the affected scenario’s reasoning cost by ~60%.

The key takeaway from this is that language models good customers. They tell you how your tooling fails them. You just need to listen.

The proof is in the program

Here’s a real task, run through the real harness. Today is 2026-06-10. Escalate every open helpdesk ticket created more than seven days ago whose body mentions a refund - subjects don’t - by setting priority to high, adding the tag billing-escalation, and submitting the escalated IDs. Tickets are paginated, and the mock API injects transient 503s.

This is what Kimi K3 submitted, first attempt, compiled clean, verbatim from the archived transcript:

first = boundary retry 2 {
    return list_tickets({ page: 1, status: "open" })
} catch err {
    return fail("LIST_FAILED", err.code + ": " + err.message)
}
remaining = for page in list.range(2, first.total_pages + 1) {
    result = boundary retry 2 { return list_tickets({ page, status: "open" }) } catch err {
        return fail("LIST_FAILED", err.code + ": " + err.message)
    }
    return result.items
}
all_open = fold acc = first.items for page in remaining { return acc + page }
cutoff = time.parse("2026-06-03T00:00:00Z")
old = for t in all_open {
    skip if t.created_at == null
    raw = t.created_at
    norm = raw + "T00:00:00Z" if text.length(raw) == 10 else raw
    ts = time.parse(norm)
    skip if ts >= cutoff
    return t
}
escalated = for t in old {
    detail = boundary retry 2 { return get_ticket({ id: t.id }) } catch err {
        return fail("GET_FAILED", err.code + ": " + err.message)
    }
    body = detail.body if detail.body != null else ""
    skip if not ("refund" in text.lower(body))
    upd = boundary retry 2 {
        return update_ticket({ id: t.id, priority: "high", add_tags: ["billing-escalation"] })
    } catch err {
        return fail("UPDATE_FAILED", err.code + ": " + err.message)
    }
    return t.id
}
result = submit_result({ answer: { escalated_ticket_ids: escalated } })
return { escalated, recorded: result.recorded }

Most of that reads like any language you know. The parts that don’t:

  • boundary retry 2 { ... } catch err { ... } is an expression that produces a value. The whole subgraph inside it retries on retryable failures; if it never succeeds, the catch block’s return becomes the result. This is how the program survives the injected 503s.
  • skip if condition drops the current element and moves on - the filter form, because there’s no early return.
  • for is concurrent. Every page fetch runs together, every ticket detail runs together, every update runs together. Result order is preserved. There is no annotation asking for any of that.
  • fold acc = first.items for page in remaining is the sequential form, because flattening pages onto an accumulator depends on the previous iteration. Independent work goes in for; dependent work goes in fold.
  • upd = ... is assigned and never read. Effectful calls run when their statement runs, whether or not you use the result. Fire-and-forget is a legitimate style.
  • submit_result(...) is nested inside the program, so the answer is filed without another trip through the model.

That run took two model requests: one to write the program, one to say “done”. Twenty tickets, three tool families, forty-odd underlying calls, one round trip.

The Lua arm of the same benchmark, same model, same task, same repetition:

local cutoff = "2026-06-03"
local page, matches, inspected = 1, {}, 0
repeat
  local r = tool('list_tickets', { page = page, status = 'open' })
  for _, it in ipairs(r.items) do
    inspected = inspected + 1
    local date = string.sub(it.created_at or "", 1, 10)
    if date ~= "" and date < cutoff then
      local t = tool('get_ticket', { id = it.id })
      if string.find(string.lower(t.body or ""), "refund", 1, true) then
        tool('update_ticket', { id = it.id, priority = 'high', add_tags = { 'billing-escalation' } })
        matches[#matches + 1] = it.id
      end
    end
  end
  page = page + 1
until page > (r.total_pages or 1)
return { inspected_open = inspected, escalated = matches }

Shorter, more familiar, entirely serial - and it died on an injected 503, because Lua has no retry construct and the model didn’t build one. The model rewrote the script, ran it, then called submit_result as a separate step. Four model requests instead of two.

Both got the right answer. Both escalated exactly the six correct tickets. And here is the part that ruins the tidy narrative: that single Runlet run still cost more. 1,548 output tokens against Lua’s 537 - of which 1,076 were reasoning tokens the transcript never shows.

Line len-1

I’d love to tell you the purpose-built language won. It didn’t.

What’s measured, from 100 runs on Kimi K3 - final configuration, ten repetitions per cell, five scenarios, both languages:

  • Accuracy: 1.00 in all 100 runs. Both arms. The model writes correct programs in a language that did not exist in its training data.
  • Round trips: 2.7 model requests per run against Lua’s 4.6. On the support-triage scenario, 2.1 against 4.8.
  • First-pass compile rate: 38 of 50 runs produced no compiler diagnostic at all. Fourteen diagnostics total across 69 submitted programs, and not one of them a parse error - the remainder were argument-arity and property-projection complaints from the analyzer. The syntax problem is solved at the frontier tier.
  • Robustness: 13 compose failures across 50 Runlet runs, against Lua’s 52 across 50. Retry-as-a-language-construct paid off.
  • Prompt cost: ~7 KB, about 1,700 tokens, injected once per turn as a tool description.
  • Cost: 1.77× Lua’s, at identical accuracy.

That last number is not the one I wanted. When you decompose it, it isn’t the primer, or the catalogue, or verbosity - the Runlet programs are about three times shorter than the equivalent Lua scripts. It’s deliberation. Kimi burns 2,685 reasoning tokens per Runlet run against 686 for Lua. The model works Runlet out every single time, in its head, at output-token prices, before writing a character you can see.

It’s a familiarity tax. Invisible in transcripts, it’s paid before the first visible token, and no prompt can refund it - the primer A/B confirms this, since primer form moved nothing at the top tier. It also isn’t uniform: it tracks how much a model habitually deliberates, not how good it is. Across the models we ran the full matched suite on, grok-4.5 barely thinks and pays about 1.3×; Kimi thinks hard and pays 1.8×. And on one scenario shape - page, filter, escalate, submit - Runlet is simply cheaper than Lua on both models.


  • runlet on GitHub for the language internals and REPL if you wish to play
  • agentkit on GitHub the agent harness libary I’ve built that hosts the compose meta tool I used to benchmark Lua vs Runlet
  • Kit, an agent harness that materialises all of the lessons I’ve learnt benchmarking and tuning LLM tooling

So, did I teach an LLM a new language?

No. The weights are exactly where they were. What I built is a machine that manufactures temporary competence, on demand, in a context window - a language with a tiny footprint and ambiguous choices designed out, one compile-checked exemplar program in place of a specification, a tool catalogue that supplies the nouns, and a compiler that corrects the model’s instincts the moment they misfire. Turn it off and the model forgets everything, instantly and completely.

Which turns out to be enough. Not free - you’ll pay the tax in tokens nobody can see, and the only real way to stop paying it is to get your language into the next model’s training data. But for a constrained domain, a small grammar and a loud compiler will substitute for post-training you’ll never get to do.

Statistically speaking, the most correct answer was still Lua. I taught it to say something else anyway.