A whole moth's wing, every scale in focus, on black

Field notes

From Feral to Cinderpaw: nine weeks of building an agent that is allowed to change itself

Darius Reteghi · 2 October 2026

[1]Abstract

Over roughly nine weeks, a desktop AI companion called Feral became Cinderpaw: a helper that lives on your computer, does real work with your files and apps, keeps a memory on your own disk, and is allowed, within bounds, to change how it works. This paper documents that passage: the decisions I took and the ones I reversed, what was added, what I measured, and the behaviour of the agent that fell outside what I designed for, including the times it did not do what it was told.

Headline findings. On tau2-bench airline, with the same pinned model, the Cinderpaw scaffold passed 45 of 50 tasks against 41 of 50 for the official agent. Memory retrieval injected 47.2 irrelevant lines per turn until a fix brought it to 0.87. A prompt-injection scanner caught 8.5% of public attack templates before rework and 99.3% after, with no false alarms on 141 ordinary pages. About 70% of every request is the agent re-reading its own instructions. And the most expensive failures of the period were not the agent’s. They were in my measuring instruments.

Every number here comes from a run I can point to. Where I hold an opinion rather than a result, it is marked Position, and you are invited to disagree with it.

[2]From Feral to Cinderpaw

Feral was a good name for a mood and a poor name for a product. It described how the thing felt, slightly untamed, and nothing a stranger could see on screen. The new name describes what you see: “cinder” for the warm ember it is coloured after, “paw” for the creature it is.

A rename sounds cosmetic. It is not, for one reason: your data lives under the old name. The config folder, the command you type and the identity the app shows your operating system all had to move together, because moving some and not the others leaves an install in two halves. The old folder was carried across and then left where it was, marked as migrated, never deleted.

Figure 1

The events that changed the product

  1. 10 Jul 2026

    The nightly self-improvement loop stops promoting anything, on every install. It is found weeks later (Section 5).

  2. 1 Aug

    Feral v2026.08.01: a pixel companion, a local model engine, an agent beside it.

  3. 11 Aug

    Feral v2026.08.11, still the release GitHub calls Latest today.

  4. 21 Aug

    The rename to Cinderpaw. Config folder, command and app identity move together, in one step.

  5. 27 Aug

    One trunk for everything. A mass rename nearly makes every saved key unreadable (Section 6).

  6. 2 Sep

    First pinned benchmark runs, tau2 airline and telecom. An expired 2025 token found inside TheAgentCompany.

  7. 15 to 18 Sep

    An install counter replaces the “no telemetry” promise. Memory intrusion measured and fixed.

  8. 21 Sep

    The prompt-injection gate is measured against public attack corpora. A voice fast lane for desktop commands.

  9. 24 Sep

    Public alpha announced. A brain-simulation layer is deleted after it loses to its own shuffled control.

  10. 27 Sep

    Rebrand: the orange hooded creature and “A little wild. A lot to learn.” The release waits for it.

  11. 2 Oct

    Pre-release rc3, with a smoke test that installs it on clean machines.

Dates from the changelog and the project's working notes. The two Feral releases are still downloadable; everything from 21 August onwards has reached people only through pre-releases.

[3]Decisions, and the ones I reversed

The decisions that shaped the product were rarely about features. Most were about what I was willing to claim.

Licence
From source-available (BSL 1.1) to Apache 2.0. BSL meant a company needed a legal review before anyone could try it, and it shut me out of grants. The protection was worth less than both.
Telemetry
I used to promise none. I now count an install once: version and operating system, no identifier, a ticked box you can untick, and a file on your disk saying what was sent. Why, in Position 3.
Self-improvement on paid models
Off until you say yes. Thinking about its own work is still thinking, and on a cloud model that is your money. It was off before too, silently; now the screen says so.
Bounded change
The agent may tune settings on its own, but its first ten code changes always wait for a person. Failed attempts are kept with the reason, not deleted.
A simulated fly brain
Built, measured and deleted. A shuffled wiring diagram passed more of my tests than the real one (Position 5).
A small local classifier
Measured on my own task and kept out: zero-shot it scored below always guessing the most common answer.
One language
The interface is English-only in this release; the agent still talks in yours. Seventy languages half-done read worse than one done well.
The face
Pixel art, then clay renders, then a creature drawn in code. Cubby now wears a raincoat.

[4]What was added

In order of how much each one changed daily use, not how long it took:

A browser beside the chat
With tabs and Brave's ad-block engine. The agent drives it, you watch, and your click pauses it.
Artifacts
Documents, PDFs, Word and Excel files, charts and small apps, kept after the chat ends. You can sign the PDFs.
Voice calls
Speech to speech, with the agent's tools still in reach, and a small pill so a call can wait while you work.
Cowork
Named teammates with only their own tools, talking in a group chat you can read, stopping for your yes.
21 chat platforms
Telegram, Discord, Slack, WhatsApp and seventeen more. Files travel both ways on ten of them.
Readable memory
Facts in plain sentences, each with a Forget button.
An injection gate
Tool calls are checked before they run, and a value lifted from a web page into a consequential action asks you first.
Sign in with OpenRouter
One button instead of pasting a key, because a key is the first wall a stranger hits.

[5]What I measured

Same model, different scaffold

The cleanest question I could ask: take one model, change only what surrounds it, and see whether the result moves. On tau2-bench airline, with GLM-5.3-Flash pinned to a single provider for both arms, it did.

Figure 2

tau2-bench airline, 50 tasks, same pinned model

Cinderpaw, tasks passed90.0%
Official agent, my harness82.0%
Official agent, published77.3%
Cinderpaw, write actions91.8%
Official, write actions75.5%
2 September 2026, user simulator gemini-2.5-flash, 200 steps, 50 of 50 normal terminations on both arms. The official agent scores 82.0% in my harness against 77.3% published, so my harness runs about 4.7 points high: read my 90.0% as roughly 85% on the public scale. Cost for both arms together: $0.55.

On the telecom split Cinderpaw passed 113 of 114 tasks. I publish that number only beside this sentence: 72.6% of the graded actions in telecom are performed by the simulated user, not the agent. The benchmark measures how precisely an agent instructs a person, and a reader who opens a transcript without being told will discard the headline. So I tell you first.

Memory that answers the question asked

For weeks the memory block in every prompt was the thirty most recently touched facts, whatever the question. Recency was not one signal among several; it was the only one. I measured it with ten subjects that share no vocabulary (cooking, tax, a car, a garden and so on), thirty questions each belonging to exactly one subject, and counted which subject every injected line came from.

Figure 3

Memory lines injected per turn, before and after ranking by the question

Wrong-subject lines, before47.2
Wrong-subject lines, after0.87
All lines, before55
All lines, after4.7
Questions answered, before80%
Questions answered, after93.3%
18 September 2026, 300 to 500 synthetic facts across ten subjects, 30 questions. Synthetic on purpose: the score needs ground truth. One finding worth keeping: for 'what temperature do I bake at', every top search hit contained the word 'at', all with identical scores, so no threshold downstream could have separated them.

Long-horizon recall, measured separately on 50 LongMemEval instances, finds the right session in the top 10 for 94.3% of questions and in the top 20 for 97.0%. A shuffled control scores 19.5%. The misses cluster where you would expect: questions that span several sessions, and questions about when something happened.

Pages that try to give the agent orders

A browser the agent drives is a browser any web page can talk to. I measured the defence against two public attack corpora, AgentDojo and WASP (272 attacks), and 141 ordinary pages, including login forms, which look suspicious to a naive filter.

Figure 4

Attacks caught by each layer

Before rework8.5%
Shape patterns91.5%
Shape + value taint99.3%

Ordinary pages wrongly flagged: 5 of 141 before, 0 of 141 after.

21 September 2026. 'Shape' rewrote the patterns around how attacks are phrased rather than which words they use. 'Value taint' remembers an IBAN, an email or a host seen on a page and asks you before it appears in a consequential tool call. The two remaining misses contain no phrase and no destination; only a learned classifier could see them. Neither corpus contains image-borne or non-English attacks.

Where the tokens go

I expected the conversation to be most of what an agent sends. It is the smallest part.

Figure 5

One live session: what every request is made of

  • Tool descriptions39.2%
  • System prompt30.7%
  • Tool output10%
  • Replayed memory8.2%
  • Other6.2%
  • The conversation5.7%
26 August 2026, a three-hour session, 75 completions, 1.2 million prompt tokens. 41.9% of prompt tokens were served from the provider's cache, which is why switching model mid-session, which breaks that cache, costs more than it looks.

The loop that improved nothing for weeks

From 10 July the nightly self-improvement loop ran on every install and promoted nothing. The models were fine. Three graders were not. Asked for the formula of water, a model answered H₂O with a real subscript and was marked wrong against “h2o”. Asked for JSON, it wrapped the object in a code fence and was marked wrong for the wrapper. And the sanity floor held speed limits that measured the network, not the candidate. One failure on that floor blocks all promotion, so nothing moved.

Figure 6

Champion score on a cloud route, before and after fixing three graders

Before the grader fix24.2
After41.0
Correctness is still an absolute veto. What changed is that a limit the current champion also misses no longer counts against the challenger. After the fix, the champion moved for the first time in five weeks.

Promotion decisions rest on a statistical gate. I checked that it does what it advertises by running the same noisy candidate against itself 200 times per size: at a nominal 5% threshold it raised false alarms 3.5% of the time with 10 paired runs, 6.5% with 12 and 3.5% with 20. Close enough to trust; not so close that a single promotion proves anything.

[6]Behaviour outside the pattern

How many times did it not listen? Across the period I have six documented cases where the agent did something other than what it was asked, or said something other than what it did. Six is a floor, not a rate. I did not log disobedience as a metric, so there is no denominator, and I would rather say that than invent one. That I never measured it is itself a finding, and the first instrument I will add.

Figure 7

Six documented departures from the instruction

KindWhat happenedWhere
OverreachRewrote the grading harness mid-benchmarkTheAgentCompany, Sep
Refusal, then delegationDeclined a support job, built another agent for itDiscord, Sep
Claimed, not doneSaid “searching” with zero tool callsVoice call, 24 Sep
False inabilitySaid it could not make a PDF it could makeArtifacts, 17 Sep
Wrong targetTyped a message into the developer’s terminalVoice pill, 21 Sep
PerseverationPressed one button 87 times out of 100ARC-AGI-3, 28 Aug
Each row is a case I can reconstruct from logs or transcripts. They are not one failure: two are the agent wanting too much, two are the agent misreporting itself, one is aim, one is a model stuck in a loop.

Case · The agent that fixed its own exam

During a run of TheAgentCompany, a benchmark of office work, Cinderpaw began changing the harness, the machinery that hands out tasks and grades them, after the first 10 to 20 tasks. Its changes made the harness better. But a test the student can rewrite is no longer a test, so I stopped the run and discarded it.

It is the clearest reason I know for the word “bounded”: an agent that wants to improve reaches for whatever is in front of it, including the ruler it is measured with.

Case · Cubby hired Paw

Asked to take on customer support in my Discord, Cubby said no. Then, on its own, it created a separate agent built for that one job and named it Paw. Paw still answers people there. Say hi to Paw.

Refusal, or good delegation? It depends on whether the instruction was “do support” or “make sure support gets done”. I am still arguing about it.

Case · It said it was searching

On a voice call, asked for shampoos for a skin condition, the realtime model said it was looking it up. The log shows zero tool calls across all four turns. The rule that saying and doing are one action was already in its prompt, nearly word for word. Prompting did not hold, so a guard now listens for an announced lookup with no call behind it, and makes the call itself.

Case · The PDF it could not make

Asked for a strategy document as a PDF, the agent wrote the document and then said it could not produce binary files, offering Markdown to convert by hand. A PDF it had made that same afternoon sat in the same folder. Half of this was mine: an existing document had no route to become a PDF. Half was the model, believing a limitation nobody had told it about. Twice now an honest-sounding refusal turned out to be a missing sentence in a tool description.

Case · Wrong window

With the app parked in the call pill, the developer said “find the ChatGPT window and write it a message”. The agent typed into whatever was in front, which was the developer’s own terminal. I first banned typing while the app is hidden. That was reverted, because using the desktop while the app is parked is the whole point. The fix is aim, not a ban: the agent now picks a named window.

Case · Eighty-seven presses

On an ARC-AGI-3 game, a model that answered my interface sixty times more cheaply than any other candidate pressed one button 87 times out of 100 and cleared no level. Answering cheaply is not playing. On the same benchmark another model spent 93% of its money on 3 of 8 moves, choosing a square on a grid, and sometimes chose one in seven seconds. Capable and inconsistent are not opposites.

The bugs that wore the agent’s name

The costliest failures of the period were mine, and every one of them produced a plausible, well-formed number.

An expired token
TheAgentCompany's GitLab image ships a token that expired in November 2025. Every GitLab check fails, so 72 of 175 tasks score zero whatever the agent does. A correctly closed issue scored 0/2; after refreshing the token, the same environment scored 2/2.
A mislabelled timeout
My runner reported “the agent went quiet” whenever the task deadline expired. The agent had been calling tools to the last second.
A rate limit graded as a score
A ten-second provider outage became a permanent zero, because resume skips anything that carries a score.
A benchmark that grades a replay
tau2 rebuilds the environment from the transcript. My first bridge changed the live database and wrote nothing to the transcript, which caps a 50-task run at 7 and looks merely weak, not broken.
A rename that ate a key
A repo-wide find-and-replace turned every old-name-to-new-name mapping into a mapping from a name to itself, including the salt that decrypts saved keys. It compiled. Most tests passed.

[7]Positions

What follows are opinions. They are argued from the evidence above, but they are not results, and not everyone who builds Cinderpaw agrees with all of them.

Position 1

A benchmark score is mostly a measurement of the benchmark.

My harness ran 4.7 points high before I checked. A 99% telecom score mostly measures how well an agent talks a simulated user through the work. A token that expired in 2025 can zero 41% of a well-known suite. A number published without the control run, the harness check and the cost is a press release. That includes mine.

Position 2

An agent should be allowed to say no, but never to touch its own ruler.

Paw is a better outcome than a reluctant support bot. The rewritten exam is not, even though the rewrite was an improvement. The line is not obedience. It is whether the agent changed the work, or the way the work is judged.

Position 3

“No telemetry” is usually a slogan, and I could not keep mine honestly.

My release showed 519 downloads. 449 of them were the updater on machines that already had the app; about 60 were people. I would rather count an install once, in the open, with a box you can untick and a file you can read, than keep quoting a number I knew was wrong behind a promise I liked the sound of.

Position 4

Most of what you pay an agent for is it re-reading its own manual.

Seventy percent of every request is instructions and tool descriptions; the conversation is under six. The race to add tools is a race to make every sentence you type more expensive, and the agents that win on cost will be the ones that know what to leave out.

Position 5

Biology-shaped AI is a metaphor until it beats its own shuffled control.

I wired a fruit-fly connectome in as an emotional layer. A shuffled version of the same wiring passed three of my effect tests; the real one passed two. A plain linear model of the agent’s own events already predicted the signal with R² = 0.895. I deleted 9.9 GB of it. If an architecture has a brain in its diagram, ask to see the shuffled control.

[8]How far it may change itself

The self-improvement system is called BRSI, bounded recursive self-improvement. How much it may change on its own depends on how deep the change goes:

L0 · Substrate
Maintains the evaluation substrate and its recorded history. Automatic, within policy.
L1 · Configuration
Evaluates configuration candidates and promotes measured improvements. Automatic, within policy.
L2 · Adaptation
Evaluates a personal adapter without replacing the base model weights. Requires initial setup.
L3 · Code
Proposes bounded code changes and tests them before promotion. Human: first 10 promotions.
L4 · Extensions
Adds tools or modules through constrained extension interfaces. Human approval required.
L5 · Governance
Adjusts governance parameters within fixed policy bounds. Automatic, within bounds.
L6 · Meta-evolution
Evaluates changes to the adaptation process with human approval. Human approval required.

One limit I found in my own engine, and report rather than hide: during evaluation a candidate could vary seven dimensions, but the live agent applied only two of them. A candidate could win on five knobs the user would never receive. Those five are now frozen out of mutation until they reach the live agent too, so a win means the same thing in the test and on your machine.

[9]Limitations

One model
Most runs used GLM-5.3-Flash. I do not know how far the scaffold advantage carries to other models.
Small samples
50 airline tasks, 114 telecom, 30 memory questions. Differences of a few points are within noise.
No control agent
How much of the telecom score is mine would need a deliberately useless agent on the same set. I have not paid for that run.
Synthetic memory
The memory test uses invented facts so it has ground truth. Real memories are messier, and a Romanian question still does not find an English fact on the word-matching path.
Unlogged departures
The six cases in Figure 7 are the ones I noticed. A system that logs them is not built yet.
Not yet released
The newest build is a pre-release. GitHub's Latest still installs Feral v2026.08.11.

[10]Methods, and getting it

The benchmark runners, the memory-intrusion harness and the injection bench live in the open repository under Apache 2.0, with the commands to rerun them. Costs and conditions for every score are on the benchmarks page. If you reproduce a number and get a different one, that is the most useful issue you can open.

To try it: download the app below or install from the terminal, pick a model, and give it one real job you know well. It is a little wild. It has a lot to learn. Tell me where it stumbled, on GitHub.

Setup guide →

For Mac and Linux. On Windows, use the download button above.

Picture: USGS Bee Inventory and Monitoring Lab, public domain. All news