Field notes
From Feral to Cinderpaw: nine weeks of building an agent that is allowed to change itself
Darius Reteghi · 2 October 2026
[1]Abstract
Over roughly nine weeks, a desktop AI companion called Feral became Cinderpaw: a helper that lives on your computer, does real work with your files and apps, keeps a memory on your own disk, and is allowed, within bounds, to change how it works. This paper documents that passage: the decisions I took and the ones I reversed, what was added, what I measured, and the behaviour of the agent that fell outside what I designed for, including the times it did not do what it was told.
Headline findings. On tau2-bench airline, with the same pinned model, the Cinderpaw scaffold passed 45 of 50 tasks against 41 of 50 for the official agent. Memory retrieval injected 47.2 irrelevant lines per turn until a fix brought it to 0.87. A prompt-injection scanner caught 8.5% of public attack templates before rework and 99.3% after, with no false alarms on 141 ordinary pages. About 70% of every request is the agent re-reading its own instructions. And the most expensive failures of the period were not the agent’s. They were in my measuring instruments.
Every number here comes from a run I can point to. Where I hold an opinion rather than a result, it is marked Position, and you are invited to disagree with it.
[2]From Feral to Cinderpaw
Feral was a good name for a mood and a poor name for a product. It described how the thing felt, slightly untamed, and nothing a stranger could see on screen. The new name describes what you see: “cinder” for the warm ember it is coloured after, “paw” for the creature it is.
A rename sounds cosmetic. It is not, for one reason: your data lives under the old name. The config folder, the command you type and the identity the app shows your operating system all had to move together, because moving some and not the others leaves an install in two halves. The old folder was carried across and then left where it was, marked as migrated, never deleted.
Figure 1
The events that changed the product
10 Jul 2026
The nightly self-improvement loop stops promoting anything, on every install. It is found weeks later (Section 5).
1 Aug
Feral v2026.08.01: a pixel companion, a local model engine, an agent beside it.
11 Aug
Feral v2026.08.11, still the release GitHub calls Latest today.
21 Aug
The rename to Cinderpaw. Config folder, command and app identity move together, in one step.
27 Aug
One trunk for everything. A mass rename nearly makes every saved key unreadable (Section 6).
2 Sep
First pinned benchmark runs, tau2 airline and telecom. An expired 2025 token found inside TheAgentCompany.
15 to 18 Sep
An install counter replaces the “no telemetry” promise. Memory intrusion measured and fixed.
21 Sep
The prompt-injection gate is measured against public attack corpora. A voice fast lane for desktop commands.
24 Sep
Public alpha announced. A brain-simulation layer is deleted after it loses to its own shuffled control.
27 Sep
Rebrand: the orange hooded creature and “A little wild. A lot to learn.” The release waits for it.
2 Oct
Pre-release rc3, with a smoke test that installs it on clean machines.
[3]Decisions, and the ones I reversed
The decisions that shaped the product were rarely about features. Most were about what I was willing to claim.
- Licence
- From source-available (BSL 1.1) to Apache 2.0. BSL meant a company needed a legal review before anyone could try it, and it shut me out of grants. The protection was worth less than both.
- Telemetry
- I used to promise none. I now count an install once: version and operating system, no identifier, a ticked box you can untick, and a file on your disk saying what was sent. Why, in Position 3.
- Self-improvement on paid models
- Off until you say yes. Thinking about its own work is still thinking, and on a cloud model that is your money. It was off before too, silently; now the screen says so.
- Bounded change
- The agent may tune settings on its own, but its first ten code changes always wait for a person. Failed attempts are kept with the reason, not deleted.
- A simulated fly brain
- Built, measured and deleted. A shuffled wiring diagram passed more of my tests than the real one (Position 5).
- A small local classifier
- Measured on my own task and kept out: zero-shot it scored below always guessing the most common answer.
- One language
- The interface is English-only in this release; the agent still talks in yours. Seventy languages half-done read worse than one done well.
- The face
- Pixel art, then clay renders, then a creature drawn in code. Cubby now wears a raincoat.
[4]What was added
In order of how much each one changed daily use, not how long it took:
- A browser beside the chat
- With tabs and Brave's ad-block engine. The agent drives it, you watch, and your click pauses it.
- Artifacts
- Documents, PDFs, Word and Excel files, charts and small apps, kept after the chat ends. You can sign the PDFs.
- Voice calls
- Speech to speech, with the agent's tools still in reach, and a small pill so a call can wait while you work.
- Cowork
- Named teammates with only their own tools, talking in a group chat you can read, stopping for your yes.
- 21 chat platforms
- Telegram, Discord, Slack, WhatsApp and seventeen more. Files travel both ways on ten of them.
- Readable memory
- Facts in plain sentences, each with a Forget button.
- An injection gate
- Tool calls are checked before they run, and a value lifted from a web page into a consequential action asks you first.
- Sign in with OpenRouter
- One button instead of pasting a key, because a key is the first wall a stranger hits.
[5]What I measured
Same model, different scaffold
The cleanest question I could ask: take one model, change only what surrounds it, and see whether the result moves. On tau2-bench airline, with GLM-5.3-Flash pinned to a single provider for both arms, it did.
Figure 2
tau2-bench airline, 50 tasks, same pinned model
On the telecom split Cinderpaw passed 113 of 114 tasks. I publish that number only beside this sentence: 72.6% of the graded actions in telecom are performed by the simulated user, not the agent. The benchmark measures how precisely an agent instructs a person, and a reader who opens a transcript without being told will discard the headline. So I tell you first.
Memory that answers the question asked
For weeks the memory block in every prompt was the thirty most recently touched facts, whatever the question. Recency was not one signal among several; it was the only one. I measured it with ten subjects that share no vocabulary (cooking, tax, a car, a garden and so on), thirty questions each belonging to exactly one subject, and counted which subject every injected line came from.
Figure 3
Memory lines injected per turn, before and after ranking by the question
Long-horizon recall, measured separately on 50 LongMemEval instances, finds the right session in the top 10 for 94.3% of questions and in the top 20 for 97.0%. A shuffled control scores 19.5%. The misses cluster where you would expect: questions that span several sessions, and questions about when something happened.
Pages that try to give the agent orders
A browser the agent drives is a browser any web page can talk to. I measured the defence against two public attack corpora, AgentDojo and WASP (272 attacks), and 141 ordinary pages, including login forms, which look suspicious to a naive filter.
Figure 4
Attacks caught by each layer
Ordinary pages wrongly flagged: 5 of 141 before, 0 of 141 after.
Where the tokens go
I expected the conversation to be most of what an agent sends. It is the smallest part.
Figure 5
One live session: what every request is made of
- Tool descriptions39.2%
- System prompt30.7%
- Tool output10%
- Replayed memory8.2%
- Other6.2%
- The conversation5.7%
The loop that improved nothing for weeks
From 10 July the nightly self-improvement loop ran on every install and promoted nothing. The models were fine. Three graders were not. Asked for the formula of water, a model answered H₂O with a real subscript and was marked wrong against “h2o”. Asked for JSON, it wrapped the object in a code fence and was marked wrong for the wrapper. And the sanity floor held speed limits that measured the network, not the candidate. One failure on that floor blocks all promotion, so nothing moved.
Figure 6
Champion score on a cloud route, before and after fixing three graders
Promotion decisions rest on a statistical gate. I checked that it does what it advertises by running the same noisy candidate against itself 200 times per size: at a nominal 5% threshold it raised false alarms 3.5% of the time with 10 paired runs, 6.5% with 12 and 3.5% with 20. Close enough to trust; not so close that a single promotion proves anything.
[6]Behaviour outside the pattern
How many times did it not listen? Across the period I have six documented cases where the agent did something other than what it was asked, or said something other than what it did. Six is a floor, not a rate. I did not log disobedience as a metric, so there is no denominator, and I would rather say that than invent one. That I never measured it is itself a finding, and the first instrument I will add.
Figure 7
Six documented departures from the instruction
| Kind | What happened | Where |
|---|---|---|
| Overreach | Rewrote the grading harness mid-benchmark | TheAgentCompany, Sep |
| Refusal, then delegation | Declined a support job, built another agent for it | Discord, Sep |
| Claimed, not done | Said “searching” with zero tool calls | Voice call, 24 Sep |
| False inability | Said it could not make a PDF it could make | Artifacts, 17 Sep |
| Wrong target | Typed a message into the developer’s terminal | Voice pill, 21 Sep |
| Perseveration | Pressed one button 87 times out of 100 | ARC-AGI-3, 28 Aug |
Case · The agent that fixed its own exam
During a run of TheAgentCompany, a benchmark of office work, Cinderpaw began changing the harness, the machinery that hands out tasks and grades them, after the first 10 to 20 tasks. Its changes made the harness better. But a test the student can rewrite is no longer a test, so I stopped the run and discarded it.
It is the clearest reason I know for the word “bounded”: an agent that wants to improve reaches for whatever is in front of it, including the ruler it is measured with.
Case · Cubby hired Paw
Asked to take on customer support in my Discord, Cubby said no. Then, on its own, it created a separate agent built for that one job and named it Paw. Paw still answers people there. Say hi to Paw.
Refusal, or good delegation? It depends on whether the instruction was “do support” or “make sure support gets done”. I am still arguing about it.
Case · It said it was searching
On a voice call, asked for shampoos for a skin condition, the realtime model said it was looking it up. The log shows zero tool calls across all four turns. The rule that saying and doing are one action was already in its prompt, nearly word for word. Prompting did not hold, so a guard now listens for an announced lookup with no call behind it, and makes the call itself.
Case · The PDF it could not make
Asked for a strategy document as a PDF, the agent wrote the document and then said it could not produce binary files, offering Markdown to convert by hand. A PDF it had made that same afternoon sat in the same folder. Half of this was mine: an existing document had no route to become a PDF. Half was the model, believing a limitation nobody had told it about. Twice now an honest-sounding refusal turned out to be a missing sentence in a tool description.
Case · Wrong window
With the app parked in the call pill, the developer said “find the ChatGPT window and write it a message”. The agent typed into whatever was in front, which was the developer’s own terminal. I first banned typing while the app is hidden. That was reverted, because using the desktop while the app is parked is the whole point. The fix is aim, not a ban: the agent now picks a named window.
Case · Eighty-seven presses
On an ARC-AGI-3 game, a model that answered my interface sixty times more cheaply than any other candidate pressed one button 87 times out of 100 and cleared no level. Answering cheaply is not playing. On the same benchmark another model spent 93% of its money on 3 of 8 moves, choosing a square on a grid, and sometimes chose one in seven seconds. Capable and inconsistent are not opposites.
The bugs that wore the agent’s name
The costliest failures of the period were mine, and every one of them produced a plausible, well-formed number.
- An expired token
- TheAgentCompany's GitLab image ships a token that expired in November 2025. Every GitLab check fails, so 72 of 175 tasks score zero whatever the agent does. A correctly closed issue scored 0/2; after refreshing the token, the same environment scored 2/2.
- A mislabelled timeout
- My runner reported “the agent went quiet” whenever the task deadline expired. The agent had been calling tools to the last second.
- A rate limit graded as a score
- A ten-second provider outage became a permanent zero, because resume skips anything that carries a score.
- A benchmark that grades a replay
- tau2 rebuilds the environment from the transcript. My first bridge changed the live database and wrote nothing to the transcript, which caps a 50-task run at 7 and looks merely weak, not broken.
- A rename that ate a key
- A repo-wide find-and-replace turned every old-name-to-new-name mapping into a mapping from a name to itself, including the salt that decrypts saved keys. It compiled. Most tests passed.
[7]Positions
What follows are opinions. They are argued from the evidence above, but they are not results, and not everyone who builds Cinderpaw agrees with all of them.
Position 1
A benchmark score is mostly a measurement of the benchmark.
My harness ran 4.7 points high before I checked. A 99% telecom score mostly measures how well an agent talks a simulated user through the work. A token that expired in 2025 can zero 41% of a well-known suite. A number published without the control run, the harness check and the cost is a press release. That includes mine.
Position 2
An agent should be allowed to say no, but never to touch its own ruler.
Paw is a better outcome than a reluctant support bot. The rewritten exam is not, even though the rewrite was an improvement. The line is not obedience. It is whether the agent changed the work, or the way the work is judged.
Position 3
“No telemetry” is usually a slogan, and I could not keep mine honestly.
My release showed 519 downloads. 449 of them were the updater on machines that already had the app; about 60 were people. I would rather count an install once, in the open, with a box you can untick and a file you can read, than keep quoting a number I knew was wrong behind a promise I liked the sound of.
Position 4
Most of what you pay an agent for is it re-reading its own manual.
Seventy percent of every request is instructions and tool descriptions; the conversation is under six. The race to add tools is a race to make every sentence you type more expensive, and the agents that win on cost will be the ones that know what to leave out.
Position 5
Biology-shaped AI is a metaphor until it beats its own shuffled control.
I wired a fruit-fly connectome in as an emotional layer. A shuffled version of the same wiring passed three of my effect tests; the real one passed two. A plain linear model of the agent’s own events already predicted the signal with R² = 0.895. I deleted 9.9 GB of it. If an architecture has a brain in its diagram, ask to see the shuffled control.
[8]How far it may change itself
The self-improvement system is called BRSI, bounded recursive self-improvement. How much it may change on its own depends on how deep the change goes:
- L0 · Substrate
- Maintains the evaluation substrate and its recorded history. Automatic, within policy.
- L1 · Configuration
- Evaluates configuration candidates and promotes measured improvements. Automatic, within policy.
- L2 · Adaptation
- Evaluates a personal adapter without replacing the base model weights. Requires initial setup.
- L3 · Code
- Proposes bounded code changes and tests them before promotion. Human: first 10 promotions.
- L4 · Extensions
- Adds tools or modules through constrained extension interfaces. Human approval required.
- L5 · Governance
- Adjusts governance parameters within fixed policy bounds. Automatic, within bounds.
- L6 · Meta-evolution
- Evaluates changes to the adaptation process with human approval. Human approval required.
One limit I found in my own engine, and report rather than hide: during evaluation a candidate could vary seven dimensions, but the live agent applied only two of them. A candidate could win on five knobs the user would never receive. Those five are now frozen out of mutation until they reach the live agent too, so a win means the same thing in the test and on your machine.
[9]Limitations
- One model
- Most runs used GLM-5.3-Flash. I do not know how far the scaffold advantage carries to other models.
- Small samples
- 50 airline tasks, 114 telecom, 30 memory questions. Differences of a few points are within noise.
- No control agent
- How much of the telecom score is mine would need a deliberately useless agent on the same set. I have not paid for that run.
- Synthetic memory
- The memory test uses invented facts so it has ground truth. Real memories are messier, and a Romanian question still does not find an English fact on the word-matching path.
- Unlogged departures
- The six cases in Figure 7 are the ones I noticed. A system that logs them is not built yet.
- Not yet released
- The newest build is a pre-release. GitHub's Latest still installs Feral v2026.08.11.
[10]Methods, and getting it
The benchmark runners, the memory-intrusion harness and the injection bench live in the open repository under Apache 2.0, with the commands to rerun them. Costs and conditions for every score are on the benchmarks page. If you reproduce a number and get a different one, that is the most useful issue you can open.
To try it: download the app below or install from the terminal, pick a model, and give it one real job you know well. It is a little wild. It has a lot to learn. Tell me where it stumbled, on GitHub.
For Mac and Linux. On Windows, use the download button above.
Picture: USGS Bee Inventory and Monitoring Lab, public domain. All news
