[Version 1] A Hand-Written Model Rollout Pipeline

A new model recipe lands in the repo and the rollout starts itself. One engine owns the state machine and the traffic APIs; one CI job per step owns the machines over SSH. Every forward step has a backward step.

Posted by Jessie Jia on 2026-10-08

This is version 1 of our model rollout pipeline. Nothing clever: no orchestration framework, no operator, no CRDs. A state machine in a service, a single CI workflow, and SSH. Writing it down because the shape turned out to be the interesting part.

The trigger is simple. A new model recipe lands in the repo, a contract check runs against it, and if it passes the rollout starts on its own. Nobody clicks anything.

The two halves

The thing that took longest to get right was realising there are two separate control problems, and they want different machinery.

Traffic and node state is bookkeeping: who serves customer traffic, who is excluded, which capabilities a lane advertises. That lives behind a control-plane API and the rollout engine calls it.

The machines are an SSH problem: pull an image, start a container, run a health check, stop the old one. That lives in a CI runner.

flowchart LR
    R["New recipe lands in repo"] --> K["Contract check"]
    K --> E["Rollout Engine
owns the state machine"] E <-->|"forward / backward steps"| API["Control-plane API
traffic, markers, capabilities"] E -->|"repository_dispatch
event_type: release-step"| GH["release-step.yml
ONE job on ONE runner"] GH --> T["scripts/release/attach_queue_lane.py"] T -->|ssh| M["Target machine
deploy.py, docker up/down"] GH -->|"POST /v1/releases/step-runs/:id/callback"| E

The engine never touches a machine. The runner never touches traffic state. Each side only knows the other through one interface: a dispatch going out, a callback coming back.

Every forward step has a backward step

This is the part I would keep in version 2 unchanged.

The engine’s step table is written in pairs. A rollout is a walk forward through the list; a failure is a walk backward through whatever has already been applied.

# Forward What it does
1 recheck_preconditions conditions still hold since the contract check
2 mark_canary_nodes new nodes accept pipeline traffic only
3 exclude_current_from_canary current nodes stay out of pipeline traffic
4 assert_smoke_passed the pipeline suite went green on the new nodes
5 clear_canary_marks new nodes may now take customer traffic
6 assert_capability_coverage the new version covers everything the lane serves
7 open_capability add the new capability to the lane envelope
8 soak both versions serve; compare the new one’s output
9 drain_current_nodes current nodes stop claiming work
10 retire_current_nodes current nodes leave the lane
# Backward Undoes
1 release_canary_marks the exclusion from step 3
2 close_capability step 7
3 undrain_current_nodes step 9
4 revoke_new_ticket takes the lane off the new version
5 finish_failed records why

Two things worth saying about this table.

The backward list is shorter than the forward list, and that is correct, not sloppy. Several forward steps are assertions — assert_smoke_passed changes nothing, so there is nothing to undo. Only the steps that mutate state need an inverse.

And the order matters: compensation runs in reverse, and it only runs for steps that actually applied. The engine records the high-water mark, then walks down from it.

How one step runs

The engine does not execute anything itself. It fires a repository_dispatch with event_type: release-step and a payload naming the tool, then waits for a callback.

sequenceDiagram
    participant E as Rollout Engine
    participant W as release-step.yml
    participant M as Machine
    E->>W: repository_dispatch(tool, run_id, callback_url)
    W->>W: seed status = not_run
    W->>W: validate payload, resolve tool from allowlist
    W->>W: verify the callback credential
    W->>M: ssh + run the tool
    M-->>W: status, overall, checks.json, log
    W->>E: POST /v1/releases/step-runs/:id/callback
    E->>E: pass → next forward step
    E->>E: fail → walk the backward list

The whole workflow is one job on one runner, in this order:

1
2
3
4
5
6
7
8
9
10
11
12
ONE job on ONE runner
├─ Initialize result (not_run) ← seeds $RESULT_DIR/status
├─ Validate payload, resolve the tool ← allowlist: name → file
├─ Checkout infra-bootstrap / perf-tuning / ci-pipeline
├─ Verify the tool is actually present
├─ Verify the callback credential ← before doing any work
├─ Provision SSH key
├─ Install the tunnel client
├─ Run the tool ← attach_queue_lane.py → deploy.py
│ writes $RESULT_DIR/{status,overall,checks.json,log}
├─ Report result (callback) ← reads those files, POSTs
└─ Cleanup ← rm the SSH key

Three of those lines are there because of something that went wrong once.

Seeding status = not_run as the very first step. If the job dies anywhere later, the result file already exists and says “never ran” rather than being absent. Absent and failed look identical from the engine’s side, and they are not the same thing.

Verifying the callback credential before doing any work. Otherwise you SSH into a production box, change its state, and then discover you cannot report what you did. Fail on the cheap check first.

Resolving the tool through an allowlist. The payload arrives from a dispatch and names a tool. If that name is used to build a path, anyone who can dispatch can run anything in the repo. The allowlist maps a fixed set of names to a fixed set of files, and an unknown name fails the job.

The machine side, end to end

A step like “attach this lane to the new engine” ends up here:

1
2
3
4
release-step.yml
→ scripts/release/attach_queue_lane.py
→ ssh to the machine
→ deploy.py (pull image, start container, health check)

The tool writes four files and nothing else matters to the caller:

  • status — one word, the machine-readable outcome
  • overall — pass or fail for the whole step
  • checks.json — per-check detail, which is what you read when it fails
  • log — the raw thing you grep at 2am

The callback step reads those four files and POSTs them. The engine stores checks.json verbatim. That means the debugging artifact survives the runner, which is gone minutes later.

Why a callback instead of polling

The engine could poll the CI API for job status. We went with a callback because job status answers the wrong question: it tells you the workflow exited zero, not that the health check passed. A tool can fail its checks and still exit zero, and a workflow can be cancelled after the tool succeeded.

The callback carries the result the tool itself produced, which is the thing the state machine needs.