Search

Astra and Fable as My Advisors, Human as the Judge, Codex and Claude Code as the Executors

Tadashi Shigeoka · Thu, September 24, 2026

My working day has split, over the last six months, into three clearly separate layers: an advisor layer, a decision layer, and an execution layer. The advisor layer is OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1. The decision layer is me, the Human. The execution layer is OpenAI Codex and Claude Code subagents, together with the Hermes Agent skills I have been building.

This setup grew out of two earlier posts, using Fable 5 Low to plan and Codex 5.6 Sol Light Fast to implement and picking a coding model by opportunity cost. GPT-6 Astra extended that split one layer outward: Fable mostly plans inside the codebase I am editing, while Astra advises on things outside the codebase (new domains, external behavior).

This write-up captures, as of September 2026, why I keep the three layers separated, which decisions each layer owns, how plans get translated into specs for the executors, and how I keep widening what the agents can run autonomously.

The Three Layers at a Glance

The overall shape of the flow:

flowchart LR
  A["Advisor layer<br/>GPT-6 Astra / Claude Fable 5.1"] -->|"Plans, options, comparisons"| H["Decision layer<br/>Human (me)"]
  H -->|"Chosen plan + spec"| E["Execution layer<br/>Codex / Claude Code / Hermes Agent skills"]
  E -->|"Diffs, findings"| H
  H -->|"Follow-up questions"| A

The advisor layer puts options into words. After reading code or specs, it enumerates design forks, trade-offs, and open questions. It does not decide.

The decision layer decides which option to adopt, what to do and what to skip, and how far to let the executors go. I keep this one.

The execution layer produces the diffs for a confirmed spec, pulls together research findings, and runs long-running background tasks.

Advisor Layer: When I Reach for Astra Versus Fable

I use the two advisor models for different kinds of questions.

  • GPT-6 Astra: scouting new domains, reading and comparing English primary sources, verifying external behavior through the Computer Use API, and gathering judgment material outside the current codebase
  • Claude Fable 5.1: reading and planning inside an existing codebase, maintaining consistency across long conversation contexts, and planning inside Claude Code while reading code directly

Astra lives in the ChatGPT desktop app and in Computer Use sessions as an outside-the-codebase advisor. Fable lives inside Claude Code’s conversation as the in-codebase planner. I often throw the same question at both and compare the differences before I pick a direction, so I design around using them side by side.

I do not ask the advisor layer to decide for me. That is a lesson from bringing Loop Engineering into an existing product and paying for it: when the judgment infrastructure is thin, handing decisions to the model lets debt accumulate under gates that keep passing.

Decision Layer: Decisions I Still Hold

I have deliberately kept these decisions on my side:

  • What to spend time on: priority and sequencing
  • Final design calls: once the trade-offs are laid out, which direction to take
  • Product value judgments: UX, pricing, privacy, and brand consistency
  • Security, permissions, and billing operations: issuing tokens, granting permissions, sending anything externally, approving anything that costs money
  • Anything that goes to a human reader: final text for customers, candidates, or community posts

As of September 2026, I still consult the advisor layer about all of these, but I do not delegate the decision. The failure mode when a model gets one of these wrong lands somewhere a single round of review cannot pull back.

That said, the outline of “decisions the Human holds” has narrowed slowly over the past six months. Six months ago I was manually comparing libraries side by side; now I ask Astra for a comparison table, ask Fable how well each one fits from the implementation side, lay the two answers next to each other, and pick. The Human still picks, but the Human no longer gathers the comparison.

Execution Layer: Codex and Claude Code Subagents Doing the Work

The execution layer turns a confirmed spec into a mechanical diff. I keep Codex on 5.6 Sol at reasoning level Light and speed level Fast, and Claude Code on Claude Opus 5.5, as the defaults. I only switch Claude Code to Fable 5.1 when I am consulting it on a plan. Plans and decisions are already settled upstream in the advisor and decision layers, so I lean on Opus 5.5 on the Claude Code side for diff quality per round, while Codex stays on the lighter setup for throughput.

The execution layer handles:

  • Implementation: producing the diff described by the plan
  • Reading and research: cross-reading docs, code, and logs
  • Long-running background tasks: Codex background agents running while I work on something else
  • Routine execution via Hermes Agent skills: patterned work that combines browser UI operations and external API calls

I do not let the executors decide across the boundary of the spec. If a design change outside the plan becomes necessary, I pull it back to the Human instead of letting the agent choose. That boundary is written into each skill’s SKILL.md as “operations this skill does not perform.”

Handing a Plan Across to the Execution Layer

When a plan moves from the advisor layer to the execution layer, I rewrite it into a spec. Instructions that stay only in the conversation give the executor room to re-interpret, with nothing concrete to catch a drift.

A spec is roughly three sections:

  • Scope: what the executor may touch, what it may not touch, the boundaries of callers and callees
  • Acceptance criteria: tests that decide completion, expected post-change behavior, artifacts that must not remain
  • Forbidden operations: do not finish via an alternative path, do not proceed without approval, do not add scope mid-stream

Writing all three out before I hand the task to Codex or Claude Code raises the chance the diff lands inside the intended scope. Diffs that go outside bounce in review and return to the plan to be rewritten.

The advisor layer stays in the loop after hand-off, too. When the execution layer brings back a diff, I ask Astra or Fable things like “there is a spot that looks off from the plan: is this an intended deviation, a gap in the plan, or a decision I need to make?” and align their reads before I decide.

The Loop That Widens the Autonomous Range

While this setup is running, whenever I notice I am making the same judgment call by hand repeatedly, I fix the operation as a Hermes Agent skill. Once fixed, that operation reduces to: advisor layer plans → Human approves → skill executes.

Three questions, in order, decide whether something gets turned into a skill:

  • Repetition count: have I made the same judgment manually at least three times
  • Expressibility: can the branching logic be written in SKILL.md and scripts rather than kept as natural-language intent
  • Approval automation: is this operation free of permissions, billing, external sends, and product value judgments

The third one gets the most care. Each skill’s SKILL.md also spells out forbidden operations such as not proceeding without approval and not finishing via an alternative path, so the skill cannot route around the Human’s judgment.

As a result, the number of operations the Human decides on every single time has gone down compared with a month ago.

The amount of planning I ask the advisor layer for has gone up over the same month. As more operations become scriptable, the decisions sitting just outside the scripts (which skill to build next, how far to widen the forbidden operations in a given skill) get more attention from Astra and Fable.

Where This Setup Breaks

There are shapes where this does not work. From what I bring back to our daily standup and team meetings, the recurring failure modes are:

  • The advisor layer produces an over-confident plan the Human cannot keep up with on review: I narrow the scope of each consultation so a single plan stays small enough to actually read
  • The execution layer re-interprets the spec and ends up outside scope: I write the forbidden operations into both SKILL.md and the spec prompt, and bounce out-of-scope diffs in review
  • Approval automation widens too far: I keep operations touching permissions, billing, or external sends off the auto-approval list as an operating rule

Each of these starts with the Human pushing judgment work down into the advisor or execution layer. The recovery is the same in each case: pull the Human’s decision range back inward, and let the advisor and execution layers focus on advising and executing respectively.

That’s all from laying out the three-layer team setup (Astra and Fable as advisors, Human as the judge, Codex and Claude Code as the executors) and the loop I use to keep widening what the agents can run autonomously, from the Gemba.

References