Why Gate the Orchestrator

Why I used a hook to stop the Claude Code main session from editing directly and handed implementation to subagents and Codex

While building an iOS camera app with Claude Code, even one feature could require changes to several files and their tests. One session took the request, read the code, implemented it, ran the tests, and reported the result. The finished conversation held hundreds of lines of diff, discarded code, and test logs.

After that, I often had to explain design rules we had already settled on. Which file holds which responsibility, why a piece of logic sits where it does — decided earlier, then written differently later. I have no measurements. What I have is the memory that this kept happening in the later part of long sessions.

The reasoning established at the start became difficult to find and reapply among those records.

Token Allocation Rules

At the time, in my setup, Claude usage was capped by one subscription, while the OpenAI Codex CLI was covered by a separate subscription and added no marginal token cost to the Claude allowance. Both tools produced usable boilerplate for this project. Without a separate rule, Claude usually took that work.

I used that difference to set an allocation rule. The policy document at the time put it this way.

Claude tokens are the scarce resource; Codex tokens are effectively
unmetered. Route accordingly.

Bulk code generation, boilerplate, mechanical refactoring, and test scaffolding go to Codex. A Claude worker steps in when Codex cannot be used or when the change needs tight, repeated iteration. Hard design decisions and reviews go to a model that is strong at reasoning.

I also wrote down the scope the main session handles itself: trivial edits touching one or two files, typo and wording fixes, changing a single config value, and read-only questions. Everything else is delegated. Changes that span three or more files or exceed roughly 50 lines, writing tests, refactors, and anything that takes sustained thought all fall in there.

The main session splits the request and assigns each part. Once it starts implementing directly, that routing decision becomes difficult to revisit.

From CLAUDE.md to a Hook

I first wrote the principle into CLAUDE.md. The instruction to delegate large implementations was often followed, but the main session still sometimes edited files it had already read. CLAUDE.md provides instructions to the session; it is not a setting that blocks a tool call.

I moved the instruction into a hook. Claude Code calls a PreToolUse hook right before it runs a tool. If the hook script returns exit code 2, the tool call is blocked and stderr is passed to the model. I checked edit count and size in that hook.

The first version of the rule was three lines.

  • The main session may edit code files twice per request. Documentation and config files are not counted.
  • A single code write over 100 lines is blocked. This closes the path where a large implementation is pushed through whole instead of being split up.
  • Subagents are used without restriction. Edits that happen inside a worker are not counted.

A main session that hits the limit receives this sentence as the refusal reason.

Main-session code-edit limit (2) reached this request. Delegate the
remaining changes to Codex (/codex:rescue) or the default-worker subagent.
Do not retry this edit directly or via Bash.

The first time this message appeared, the session spawned a worker with the Task tool. In later cases, sessions also followed the delegation method included in the block message. The refusal provided both the reason and the delegation action to take next.

In the sessions I observed after enabling the gate, large work went to workers and Codex. The main conversation also accumulated implementation detail more slowly than before. I did not keep measurements, so this claim is limited to what I observed while using it.

The gate blocked only calls that matched its rules. Saving a code file under a different extension slips past the counter. A second hook checked shell commands that could write files, but bypasses remained there too. This implementation was a best-effort behavior gate, not a security boundary. Its purpose was to add friction to repeated main-session editing and show a delegation path. In my work, a limit of two edits per request was enough to change that behavior.

References

Comments

Comments

    Image preview