Skip to main content
← /writing
  • #agentic-sdlc
  • #platform-engineering
  • #ai-coding

The Agent Is Not the Platform

Building a portable software-development workflow across Claude Code, Codex, Goose, Pi and OpenCode.

Vinny Carpenter17 min read3.4k words

never the stack · audio edition

The Agent Is Not the Platform

23:45

In July I wrote about a software-delivery pipeline built around three Claude agents and two human approval gates. It worked, and that was the problem. Every part of it assumed Claude Code. Since August I've been rebuilding the idea as AgentMachinist, a controller that runs the same workflow through Claude Code, Codex, Pi, or OpenCode behind thin adapters. Twenty-three releases later, the adapters have become the most interesting part of the project, because there has never been a better time to choose the wrong abstraction.

Claude Code, Codex, OpenCode, Pi, Goose, Amp, Gemini CLI, Cline: pick a week and one of them ships a release, a benchmark, or a workflow that suddenly seems indispensable. I've been running several of them, and I keep coming back to a question that matters more than which one is best. What happens if the agent isn't the thing we standardize on?

The normal approach is to choose an agent and build the development workflow around it: configure the instructions, add skills, set up hooks, connect tools, define permissions, and teach the team. Then another agent gets better, or a different model gets better at your kind of work, or the pricing changes, or the product changes direction. Suddenly a surprising amount of your development process is coupled to a tool that didn't exist two years ago. That feels backward to me.

I want the engineering system to be durable and the agent to be replaceable.

Hold everything constant except the agent

The experiment is simple to state. Take one opinionated software-development workflow and run it across Claude Code, Codex, Goose, Pi, and OpenCode. Use the same repository, the same task, the same starting commit, the same acceptance criteria, the same tests, and the same engineering expectations. Only the agent changes. GitHub Copilot CLI is the obvious sixth, and probably the one already on your company's approved list. It ships its own agents for planning, task execution, and code review, which makes it the clearest example of the opposite bet, and it joins the second round once its adapter exists.

The repository is Tycho, my Rust CLI for usage analytics on local AI transcripts, starting from the v0.9.0 tag. The change is to add OpenCode as a provider. Tycho reads Claude Code, Codex, Pi, and OpenAI transcripts today, so the work touches the parser, the reports, the pricing tables, and the tests. It also carries a design question the workflow has to surface. OpenCode moved its session storage from per-message JSON files to a SQLite database, so the spec has to decide whether to read the database, the legacy files, or both. Tycho's CI already runs cargo fmt, cargo clippy, and cargo test on Linux and macOS, which makes the gates deterministic and independent of any one harness.

Most comparisons of coding agents aren't comparisons of agents. One gets a carefully written prompt and another gets an AGENTS.md. One has MCP servers installed and another has repository-specific rules. One gets a better model and another gets a simpler task. Then we compare the output and declare a winner. I want to hold more of the system constant, because the question I care about isn't which agent writes the best code. It's how much of a good software-development process can live outside the agent.

Portable means the contract survives

Portable doesn't mean identical. I don't expect five agents to expose the same tools, permission system, context strategy, execution model, or native capabilities. I mean something narrower. The intent, workflow, acceptance criteria, verification, and evidence should survive a change of agent without redesigning the delivery system around it. The adapter can change as long as the contract holds.

That is also my answer to the objection I expect from senior engineers. A portable workflow converges on the lowest common denominator, they'll say. The reason to choose Claude Code is exactly the hooks, subagents, and permission model you lose the moment you abstract over them. The goal isn't to flatten every agent into the least capable one. If Claude Code has a stronger permission model, use it. If Goose has a useful workflow primitive, use it. If Pi makes extension easier, take advantage of that. The common layer defines the engineering expectations, and the adapter translates those expectations into whatever the harness understands. By harness I mean everything an agent wraps around the model: the loop, the tools, the permission model, the sandbox, and the session.

Separate the harness from the model

There is a complication. Switching from Claude Code to Codex changes the harness and the model at the same time, so a different result could come from either one. That means I need two experiments. The first is a whole-system comparison of all five agents, and it answers a practical question: what experience does a developer get from each system?

The second is a harness-controlled comparison. Pi, OpenCode, and Goose are provider-independent, so I can run all three on the same model as at least one of the other agents.

Same model
   │
   ├── Pi
   ├── OpenCode
   └── Goose

That answers a different question: what does the harness contribute when the underlying model stays constant? Both comparisons are useful, and they measure different things. Being explicit about that keeps the experiment honest about what it can claim.

Write the workflow down before comparing agents

Before comparing agents, I need a definition of the work I expect them to perform. My workflow has eight stages, and there's nothing novel about any of them.

UNDERSTAND
    ↓
FRAME
    ↓
PLAN
    ↓
IMPLEMENT
    ↓
VERIFY
    ↓
REVIEW
    ↓
REFINE
    ↓
DOCUMENT

That's part of the point. Good software engineering didn't become obsolete because the person typing the code might now be a model. If anything, agents make the process around the code more important, so the workflow should make each stage explicit rather than trusting the agent to infer it.

Understand before changing

The agent should begin by understanding the repository: the architecture, the relevant components, the existing conventions, where the tests live, what depends on what, and the likely blast radius of the requested change. I don't want the first action to be an edit. I want the first output to be an explanation of what the agent believes the system does, where the requested behavior lives, what assumptions it's making, and what the change might affect. Models are very good at producing plausible implementations from incomplete understanding, and that's useful right up until the implementation is confidently wrong.

Frame the problem

The next step is to define what success means: the desired outcome, the constraints, the assumptions, the non-goals, the acceptance criteria, and the significant risks. I argued in The Frame Is the Bottleneck that the constraint is moving from production to framing, and the agent workflow is where that shows up first. A powerful coding agent pointed at an ambiguous problem does not remove the ambiguity. It implements it faster.

In AgentMachinist, understanding and framing are the first half of the Spec phase. The prompt tells the agent to explore the repository before proposing anything. Then it prints a document with numbered, testable requirements, the interpretation it chose where the issue was ambiguous, the risks, a testing plan, an out-of-scope section, and open questions for a human to settle. I approve that document, pinned to an exact commit, before any file changes. It's the checkpoint I trust most.

Plan before implementation

Only after understanding and framing should the agent propose an implementation plan: the files involved, the design, the tests that should change or be added, and any architectural tradeoffs. The plan is a durable checkpoint between reasoning about the system and modifying it. Small changes don't need ceremony, but I want the intent to exist outside the agent's conversation history before implementation begins. That checkpoint turns out to be one of the things that makes work transferable between agents, and I'll come back to it.

Implementation needs a contract

Once I've approved the plan, the agent gets permission to make the change, and the rules should be boring:

  • make the smallest coherent change
  • follow the architecture and conventions already present
  • avoid unrelated refactoring
  • preserve compatibility unless the task says otherwise
  • update tests alongside behavior
  • don't silently broaden scope

Those are the expectations I'd have for an engineer working in the repository. The interesting question is where they should live. In a Claude-specific configuration, a Codex skill, a Goose recipe, a Pi extension, or an OpenCode agent definition? Or in a portable description of the engineering contract, with thin adapters translating it into whatever each agent expects?

AgentMachinist takes the second path, and the contract splits in two. The prompt carries the rules a model can follow. Implement exactly what the spec requires and nothing more, follow the conventions in neighboring code, write the tests the plan calls for, and report any deviation instead of deciding alone. The controller enforces the rules a model might not follow. It rejects an implementation that touches more files than the configured limit, refuses test deletions unless the repository allows them, and treats file contents, issue text, and diffs as input rather than instructions. The prompt is portable across all four adapters, and the enforcement is portable because it never lived in the agent.

Verification cannot be optional

An agent should not be able to declare success because the code looks right. It needs evidence. At minimum the build, the tests, static analysis, type checks, the relevant security checks, and the acceptance criteria all have to run. Not every repository has every one of those controls, so the rule is broader: claims about correctness should be backed by executable evidence whenever executable evidence is available. I made the longer case in Trust the Gate, Not the Actor.

The more code we allow agents to create, the more valuable deterministic verification becomes. Tests, type systems, static analysis, architecture rules, and observable acceptance criteria all matter more than they did, because the model is probabilistic and the verification system doesn't have to be. In practice, AgentMachinist runs the gates itself after the agent finishes, whatever the agent reported. The agent may run the same gates while it works and iterate until they pass, but the controller's run is the one that counts.

The implementing agent should not grade its own homework

The agent that implemented the change shouldn't be the only agent deciding whether the implementation is good. In my design, an independent reviewer gets the original task, the acceptance criteria, and the resulting diff, and it checks correctness, architecture, security, maintainability, test coverage, and edge cases. It reads the work without the implementing agent's explanation of why everything was done correctly, because I want an independent reading. The reviewer produces findings, and the implementing side then has to fix the findings it accepts, explain the ones it rejects, and rerun verification.

AgentMachinist runs Review as a separate, read-only run, even when it uses the same adapter as the implementation. The reviewer receives the task, the approved spec, the list of files that changed, the verification report, and the diff. It returns findings that each carry a severity, a confidence, a file, a line, and a remediation. A reviewer that edits a file, or returns output the controller can't parse, blocks integration. That separation between generation and verification is much closer to how I think agentic workflows will operate in serious software environments. Several forms of intelligence, deterministic controls, and clear responsibilities around the change will do the work, rather than one very smart agent doing everything.

The adapter boundary is the experiment

Each of these agents approaches configuration differently. Claude Code has its instructions, skills, hooks, subagents, and tooling model. Codex has its own approach to repository instructions, skills, sandboxing, and execution controls. Goose has recipes, extensions, MCP, and subagents. Pi deliberately keeps the core small and pushes richer workflow behavior into skills and extensions. OpenCode has its own agents, commands, tools, permissions, and provider abstraction. They're similar enough that a common workflow feels possible, and different enough that I'm not convinced it is.

The engineering system, with context, standards, acceptance criteria, workflow, verification, and tools, sits above a thin adapter, with Claude Code, Codex, Goose, Pi, and OpenCode below it

The experiment is about finding that adapter boundary: what can remain identical, what needs translation, and what is fundamentally agent-specific. The four adapters I've shipped already show where the line falls in practice. To keep Claude Code read-only during Spec and Review, the adapter runs it in plan permission mode with only the Read, Grep, and Glob tools. Codex gets a read-only sandbox and an ephemeral session. Pi needs four flags to turn off extensions, skills, prompt templates, and sessions, plus a tool allowlist of read, grep, find, and ls. OpenCode runs its built-in plan agent, which can't edit files, but the adapter still treats that as advisory and checks repository custody afterward instead of trusting the flag. The same phase and the same prompt produce four different ways to say "don't write anything," and one of them the harness can't fully enforce.

Agent-specific capabilities can stay, as long as the engineering system doesn't have to trust them blindly. Where a harness has a stronger native control, the adapter uses it, and where it doesn't, the controller enforces or verifies what it can independently. I may lose some of the harness's richest interactive features, and the experiment should tell me whether that loss shows up in the results.

Most of the system should sit above the agent

My hypothesis is that most of the development environment transfers. Repository context, architecture documentation, acceptance criteria, and tests should be portable. MCP gives us the beginnings of portable tool interfaces, and AGENTS.md is moving repository-level instructions toward a common format. The Agent Skills format gives reusable procedures a path toward portability even where harnesses differ in how they load or execute them. The workflow itself should be portable. In Your Agent Workflow Needs a Release Process I argued that a shared workflow deserves a version, a scope, and an owner. A workflow with a release process is one you can carry to another harness.

Three columns: what belongs to the agent, what needs translation, and what stays identical

Other things will be harder: permission systems, sandboxing, context management, subagent semantics, hooks, parallel execution, and session state. Those differences are real, and the harder objection is that many of them are exactly where safety lives. If sandboxing and permissions are agent-specific, the durable engineering system can't pretend they don't matter. I think the right answer is a layered one. Some controls remain native to the harness, and others belong to the controller. The engineering system then independently verifies the outcomes it can observe: that the workspace stayed unchanged during a read-only phase, that the required gates ran, that the candidate came from the approved spec, that the diff stayed within policy, and that the exact artifact being integrated is the one that was reviewed. A custody check isn't an OS sandbox, a test isn't a permission boundary, and controller-level verification doesn't make every agent control interchangeable, but it does reduce how much trust has to live inside one agent. If most of the engineering system sits above that line, switching agents becomes a much smaller problem, and that's the hypothesis the experiment may disprove. Mechanism can vary. Evidence should travel.

Handoffs may be more interesting than comparisons

After each agent completes the workflow on its own, I want to deliberately break the workflow across agents. Claude Code understands, frames, and plans, Codex implements, Goose verifies, and Pi reviews. Then change the order: let Pi understand the repository, Goose develop the plan, OpenCode implement it, and Claude Code review it.

Four handoffs: Claude Code produces the plan, Codex produces the diff, Goose produces the evidence, and Pi produces the findings

This is a much harder test than asking whether each agent can complete a ticket. It asks whether one agent can understand the artifacts produced by another agent well enough to continue the work. AgentMachinist can already assign a different adapter to each phase, so the plumbing exists. What I haven't done is run the scramble on purpose and study what breaks. If it breaks, that's informative. Maybe too much context exists only inside the conversation history of the agent that created it. Maybe the plan isn't explicit enough, or architectural understanding needs to become an artifact, or decisions need to be recorded differently, or the handoff format itself is underspecified. Those are engineering-system problems, and they're the kind I know how to fix.

Measure the human, too

I also want to track every time I have to intervene, and why. Was the intervention needed because the problem was poorly framed? Because the agent misunderstood the architecture, chose a bad implementation, needed permission, failed verification, or couldn't recover? Or because I noticed something neither the implementer nor the reviewer caught?

Maximum autonomy is the wrong measure of agentic development. I'm not particularly interested in celebrating that an agent completed 97 steps without asking me a question. I care whether human attention went where human judgment had the most leverage. If the agent handles implementation but repeatedly needs help framing the problem, that tells me something. If it plans well but can't reliably verify its work, that tells me something else. If independent review catches defects consistently, that's evidence for a different workflow architecture. The interventions may reveal more than the benchmark scores.

Platform engineering already solved this shape

The more I think about this problem, the more it resembles a lesson platform engineering learned the hard way. Giving developers infrastructure wasn't enough, giving them Terraform wasn't enough, and giving them Kubernetes was definitely not enough. The improvement came when platforms encoded a coherent path: good defaults, automation, guardrails, observability, self-service, and escape hatches. The goal was to make the common path safe, understandable, and efficient without eliminating choice.

Now developers are getting increasingly capable coding agents, and we could repeat the old pattern: "Here are five agents. Good luck." Or we could treat the agent as another user of the engineering platform, which I argued for in August. Give it the repository context it needs, tools with clear interfaces, executable verification, architecture constraints, a defined workflow, and enough freedom to solve the problem without rediscovering the development process every time. That makes it a platform problem, and platform problems are the kind we already know how to solve.

The agent should be replaceable

I don't know yet how portable this workflow will be, which is why I want to run the experiment instead of asserting the answer. Claude Code may execute parts of it dramatically better than everything else. Codex may expose a workflow primitive that's hard to reproduce elsewhere. Goose recipes may turn out to be an unusually good abstraction for this kind of orchestration, and its adapter is the one I'm writing now. Pi's minimalism may reveal that I've added far more machinery than the model needs. OpenCode's provider independence may make portability easier than I expect. I want the tools to challenge the thesis.

I'm still convinced that choosing the best agent isn't a durable developer-platform strategy. I made the provider-level version of this argument in Everyone Is Commoditizing Someone: the durable assets are the evaluations, context, tool contracts, and audit evidence that keep model providers replaceable. The agent layer follows the same logic. Models, agents, interfaces, and pricing will all change, and the engineering practices we expect around production software should survive all of them. So that's the experiment: one repository, one meaningful change, one development workflow, five coding agents, and one question I care about more than which agent finishes first. Can the engineering system become durable enough that the agent is replaceable? I'll report back with what transfers, what breaks, and where the abstractions leak.

// found this useful? share it

Post on X Share to LinkedIn
Vinny Carpenter

Written by Vinny Carpenter

VP Engineering · 30+ years building software

I lead engineering teams building cloud-native platforms at a Fortune 100 company. I write about engineering leadership, AI-assisted development, platform strategy, and the hard lessons that come from shipping at scale.

keep reading