This is version 1 of our model rollout pipeline. Nothing clever: no orchestration framework, no operator, no CRDs. A state machine in a service, a single CI workflow, and SSH. Writing it down because the shape turned out to be the interesting part.
The trigger is simple. A new model recipe lands in the repo, a contract check runs against it, and if it passes the rollout starts on its own. Nobody clicks anything.
The two halves
The thing that took longest to get right was realising there are two separate control problems, and they want different machinery.
Traffic and node state is bookkeeping: who serves customer traffic, who is excluded, which capabilities a lane advertises. That lives behind a control-plane API and the rollout engine calls it.
The machines are an SSH problem: pull an image, start a container, run a health check, stop the old one. That lives in a CI runner.
flowchart LR
R["New recipe lands in repo"] --> K["Contract check"]
K --> E["Rollout Engine
owns the state machine"]
E <-->|"forward / backward steps"| API["Control-plane API
traffic, markers, capabilities"]
E -->|"repository_dispatch
event_type: release-step"| GH["release-step.yml
ONE job on ONE runner"]
GH --> T["scripts/release/attach_queue_lane.py"]
T -->|ssh| M["Target machine
deploy.py, docker up/down"]
GH -->|"POST /v1/releases/step-runs/:id/callback"| E
The engine never touches a machine. The runner never touches traffic state. Each side only knows the other through one interface: a dispatch going out, a callback coming back.
Every forward step has a backward step
This is the part I would keep in version 2 unchanged.
The engine’s step table is written in pairs. A rollout is a walk forward through the list; a failure is a walk backward through whatever has already been applied.
| # | Forward | What it does |
|---|---|---|
| 1 | recheck_preconditions |
conditions still hold since the contract check |
| 2 | mark_canary_nodes |
new nodes accept pipeline traffic only |
| 3 | exclude_current_from_canary |
current nodes stay out of pipeline traffic |
| 4 | assert_smoke_passed |
the pipeline suite went green on the new nodes |
| 5 | clear_canary_marks |
new nodes may now take customer traffic |
| 6 | assert_capability_coverage |
the new version covers everything the lane serves |
| 7 | open_capability |
add the new capability to the lane envelope |
| 8 | soak |
both versions serve; compare the new one’s output |
| 9 | drain_current_nodes |
current nodes stop claiming work |
| 10 | retire_current_nodes |
current nodes leave the lane |
| # | Backward | Undoes |
|---|---|---|
| 1 | release_canary_marks |
the exclusion from step 3 |
| 2 | close_capability |
step 7 |
| 3 | undrain_current_nodes |
step 9 |
| 4 | revoke_new_ticket |
takes the lane off the new version |
| 5 | finish_failed |
records why |
Two things worth saying about this table.
The backward list is shorter than the forward list, and that is correct, not sloppy. Several forward steps are assertions — assert_smoke_passed changes nothing, so there is nothing to undo. Only the steps that mutate state need an inverse.
And the order matters: compensation runs in reverse, and it only runs for steps that actually applied. The engine records the high-water mark, then walks down from it.
How one step runs
The engine does not execute anything itself. It fires a repository_dispatch with event_type: release-step and a payload naming the tool, then waits for a callback.
sequenceDiagram
participant E as Rollout Engine
participant W as release-step.yml
participant M as Machine
E->>W: repository_dispatch(tool, run_id, callback_url)
W->>W: seed status = not_run
W->>W: validate payload, resolve tool from allowlist
W->>W: verify the callback credential
W->>M: ssh + run the tool
M-->>W: status, overall, checks.json, log
W->>E: POST /v1/releases/step-runs/:id/callback
E->>E: pass → next forward step
E->>E: fail → walk the backward list
The whole workflow is one job on one runner, in this order:
1 | ONE job on ONE runner |
Three of those lines are there because of something that went wrong once.
Seeding status = not_run as the very first step. If the job dies anywhere later, the result file already exists and says “never ran” rather than being absent. Absent and failed look identical from the engine’s side, and they are not the same thing.
Verifying the callback credential before doing any work. Otherwise you SSH into a production box, change its state, and then discover you cannot report what you did. Fail on the cheap check first.
Resolving the tool through an allowlist. The payload arrives from a dispatch and names a tool. If that name is used to build a path, anyone who can dispatch can run anything in the repo. The allowlist maps a fixed set of names to a fixed set of files, and an unknown name fails the job.
The machine side, end to end
A step like “attach this lane to the new engine” ends up here:
1 | release-step.yml |
The tool writes four files and nothing else matters to the caller:
status— one word, the machine-readable outcomeoverall— pass or fail for the whole stepchecks.json— per-check detail, which is what you read when it failslog— the raw thing you grep at 2am
The callback step reads those four files and POSTs them. The engine stores checks.json verbatim. That means the debugging artifact survives the runner, which is gone minutes later.
Why a callback instead of polling
The engine could poll the CI API for job status. We went with a callback because job status answers the wrong question: it tells you the workflow exited zero, not that the health check passed. A tool can fail its checks and still exit zero, and a workflow can be cancelled after the tool succeeded.
The callback carries the result the tool itself produced, which is the thing the state machine needs.