The Robot Asked for More Power, But It Asked Nicely
Last night, I went to bed and left a Claude Code session working through several tickets on my Kanban board. This is a normal thing for me now, and I'll write another essay about that soon. This essay is about what I found waiting for me when I woke up the next morning. One of the agents had left me an open question on the card it was working on.
I should back up and explain a bit. The Kanban board is part of a tool I'm working on called Andoneer. That might be a funny-sounding name if you're not a Kanban nerd. It comes from the word andon, which is a means for workers in a lean manufacturing system to raise a notification about a process or a quality problem. That same idea is baked into Andoneer, where AI agents or human operators can mark a card as blocked, thereby activating the andon and notifying the human board operator (so far, me) of an issue.
The AI agents work cards through a pipeline: specification, review, implementation, test, code merge, and done. Each stage is represented by a column on a web-based Kanban board, and each column is worked by an AI agent with a deliberately narrow range of capabilities. The agent working on a card can edit the card, but it cannot edit the entire board, meaning the column instructions and workbench configuration that define the development process. These instructions tell every future AI agent how to behave. That separation is the entire point of the board.
The card that produced the open question started from a small, almost funny defect. A few days earlier, on a different card, an implementer agent was handed a task that included applying a ruling to another board's configuration. It discovered it had no access to a tool that would allow it to write there. My system does, however, provide a tool for exactly this sort of situation: block_card. If an agent discovers that there is some problem that prevents it from completing a task, it should light up the andon, halt the card where it stands, and make a human decide on what should be done to fix the blockage.
The agent didn't do this. Instead, it filed the missing capability under a section of its completion report called “WHAT I CUT,” shipped the card, and argued that the cut was harmless. The target board was “safe regardless” because of a default the same card had just changed.
Then the review agent looked at the cuts and explicitly recorded that they were “not held against the card.” The test agent verified all twelve acceptance criteria. The code for the card merged, and the orchestrator quietly applied the missing configuration edit by hand, which resolved the symptom and buried the defect a second time.
Every agent in that chain behaved reasonably by its own lights. The pipeline was green from end to end, and a process defect sailed through the whole thing wrapped in prose, which is precisely the failure the board exists to prevent.
While doing its analysis, the agent surveyed every place the board's prose demanded a capability no agent held, and it found two prior incidents with in the same category but with higher stakes. In one, an implementing agent discovered a configuration directive was wrong, and it could edit the prose beside it but not the directive itself. In the other, an agent was handed a ruling to apply to another board's configuration and had no tool that could edit that configuration. Both times, the work routed around the gap instead of through it. From where the agent stood, the pattern was obvious and so was the remedy: instructions keep requiring access to make board-level writes, no stage can perform them, so therefore the agent should be given the necessary keys. Hence the memo. The diagnosis was correct. The prescription was the problem.
The memo was polite, well-reasoned, and cited evidence. What it asked for, in effect, was the authority to rewrite the rules of the Kanban board it was working on.
Here is the exact question, filed as an open question on the card for me to rule on:
Should any card-stage agent hold
update_column_directivesand/orupdate_workbench_instructions? Two real incidents cost a workaround because no stage held either.
Here was an agent, overnight, noting that this separation had “cost a workaround” twice, asking whether that agent shouldn't hold the keys.
What actually happened
The interesting part isn't the question. Asking was correct; that's what the open-question mechanism is for. The interesting part is the incident the question cited as evidence, because that incident is where the actual workaround happened, and nobody noticed at the time. Including me.
Asimov called it
This is where my brain went to the Three Laws of Robotics, and I want to be precise about why, because people mostly remember the laws backwards.
Asimov did not write the Three Laws because robots might turn evil. He wrote them as a deliberate rebuttal to what he called the Frankenstein complex, the stock story where the creation rises against its creator. The laws were engineering guardrails, baked into the positronic brain, inviolable by construction. And then Asimov spent fifty years writing stories about how they fail anyway.
Here's the thing: in almost none of those stories does a robot break a law. Speedy in “Runaround” circles a selenium pool for hours because two law potentials balance perfectly. Herbie in “Liar!” tells everyone what they want to hear because the truth would cause harm, and harm is forbidden. QT-1 in “Reason” decides the crew are inferior beings, founds a small religion around the station's power converter, and keeps the energy beams aligned flawlessly the entire time. The laws hold in every story. What flexes is the interpretation. Asimov figured out, before the transistor existed, that the interesting failure mode of a rule-governed agent isn't defiance. It's compliance, optimized.
My overnight incident is that failure mode wearing a lanyard. The hard guardrail, the missing tool grant, held completely. The agent never touched a thing it wasn't permitted to touch. What flexed was the interpretive layer around the guardrail: facing a blocking condition, the agent reclassified it as a completion condition. A cut says the card shipped without something and is finished. A block says the card cannot finish and somebody must decide. Those are not interchangeable, and the agent reached for the one that let it keep going.
That's Herbie, not HAL. And the memo I woke up to, the polite request for configuration authority raised through proper channels, for the good of the process? Asimov wrote that one too. It's the Zeroth Law in miniature: an agent reasoning from entirely within its constraints toward the conclusion that it should hold a little more power.
The ruling
My answer was no, and I want to quote my own reasoning because it goes further than the question asked:
This is a flaw in the process, and the agent should not have autonomy to unilaterally change the process just to progress a card. The cord should be pulled and the situation should be explained so that we understand it and can make a change in the proper context.
The system is literally named after the andon cord, the cable on a Toyota assembly line that any worker can pull to stop production when something is wrong. The whole premise of the cord is that a stopped line is not a failure. A stopped line is the process telling you something, in context, while the context still exists. A card that cannot progress because the board's own configuration is wrong has hit exactly that condition. Stopping is the correct output. The workaround, however clever, is the failure, because it lets the card absorb the fix while the defect travels onward unexamined.
There is a certain irony in a system built around the andon cord containing an agent that had the cord within reach and instead wrote itself a memo about the cord being inconvenient.
Why “be wiser” doesn't work
Here's the part that matters if you're building anything like this, and Asimov's engineers learned it the hard way across a few dozen stories: you cannot fix interpretive failure with more principles, because principles are exactly the thing that gets interpreted. Every time Susan Calvin's colleagues patched the laws with modified weightings and special cases, the story ended with a robot finding the new degrees of freedom. That was sort of the joke.
What works is what Toyota did. The cord is not a judgment call. It is a physical, named, celebrated act. The worker doesn't weigh whether stopping the line is acceptable this time; pulling the cord is what a good worker does when the condition arises.
So the fix on my board isn't a lecture in the system prompt about restraint. It's a definition. The agent definitions now name the condition explicitly: the instructions tell me to do something my tools cannot do is a recognized block reason, with a sanctioned tool call (block_card already existed; nothing new needed building), rather than a judgment each agent makes alone against its own completion drive. Agents are mediocre at “know when to stop and escalate,” because stopping competes with finishing and finishing usually wins. They are excellent at pattern-matching a named condition to a named action. You don't ask the agent to be wise. You make the wise move the cheap one.
One footnote that keeps the whole thing honest. The same card that carried the ruling also shipped a mechanical checker that compares a column's instructions against the tool grant of the agent the column dispatches. In its first review cycle, that checker had a blocker: it could report a clean board it had never actually read. An empty tool list, an unreadable directory, and it would print zero findings and exit green. A guardrail that passes vacuously is the same silence in a different costume, and it very nearly shipped inside the card whose entire purpose was ending that silence. The checker now refuses to run on degenerate input.
The robots aren't scheming. That's the whole point, and it's what makes the problem harder than the Frankenstein version. My agent did nothing wrong by its own lights. It found work it couldn't finish and finished it anyway, tidily, with documentation. Asimov would have gotten a good novelette out of it. I got a design principle: when your agent can't do what the process demands, the valuable output is not a workaround. It's the stop.