outerlayer runner start under a
service manager, as an account of its own. Set the host up first with
Set up your first host.
The account the runner runs as
The runner never prompts. Under the service’s account:- Claude’s credential is in
runner.envbeside the config file. Built-in hooks use only that file. A terminalclaudelogin is not used. See Give the runner Claude’s credential. - git must fetch with no credential prompt. The runner uses the
account’s own git login for
outerlayer runner imageand to read a governing control plane’s context. Your own hooks may use it too. Use a deploy key, a credential helper with a stored token, or an SSH agent the service can reach. - With your own hooks, the agent must already be logged in under this account. Claude Code keeps its login per home directory.
- The factory must be reachable with the key in the config file.
idle. Run outerlayer runner check, and your command once by hand under
that account, to catch this early.
Hooks and the command run with the runner’s own CLI first on PATH, so an
outerlayer call in a hook uses the runner’s version.
What the service sets
outerlayer runner init ends by setting up the service. You write no unit.
To set it up again, run:
/etc/systemd/system/outerlayer-runner.service, then enables and starts it.
The runner starts at boot. It asks for sudo and names each sudo command
before it runs. The unit runs as the account that ran the command. Pass
--user <account> to name another.
It refuses on a Linux host without systemd, and for hooks that are your own
executables. The refusal names the reason.
--print writes the unit to
stdout and sets up nothing, so you can edit it for those hosts.
The unit holds these lines:
ExecStart=starts the runner from its install layout, so it can update itself. See Keep the runner up to date.EnvironmentFile=-namesrunner.envbeside the config file, readable only by the runner’s account. It holds Claude’s credential,CLAUDE_CODE_OAUTH_TOKEN=<token>orANTHROPIC_API_KEY=<key>. It also holds every secret a destination’s header reads and every valuebuild.variablespasses in. Without a Claude credential the runner claims nothing. Restart the service after you edit the file, because a service reads it only when it starts.Restart=on-failureandRestartForceExitStatus=75restart the runner after a crash and after the exit with status 75 that a self-update, or a versionrunner installput in place, asks for.TimeoutStopSec=istimeLimitMinutesplus ten minutes, because a stop waits for running builds.Requires=docker.serviceis set for container and microVM hooks.OOMPolicy=continuekeeps one build’s death from stopping the service. Without it, systemd stops the whole service, and every other build, when the kernel kills any process in it.Delegate=yeslets the runner give each microVM build a cgroup of its own.
ManagedOOMMemoryPressure=. The OOM daemon kills a whole
cgroup, which here is every build.
It is a system service, not a user service, so the account’s groups are
read again at each start. An account added to kvm then needs only a
service restart.
On WSL2 without systemd, setup turns systemd on in /etc/wsl.conf and
stops. It asks you to run wsl --shutdown once from Windows. Run the
command again after that. If Windows interop is off, setup prints the
schtasks command to run in Windows.
The scheduled task names your distribution, which WSL2 gives in
WSL_DISTRO_NAME. sudo drops that variable, so run setup as yourself,
not under sudo. Without the name, setup makes no task. It prints the
schtasks command with a placeholder for the name wsl -l shows.
Re-run and remove the service
The runner reads the config file again on every poll, so most edits need nothing else. Re-runouterlayer runner install --service after a change to
timeLimitMinutes, to the hooks, or to the account. With a unit that still
matches, it changes nothing and says the service is current.
When the unit changes, it rewrites the unit and reloads systemd. It restarts
the service only while no build is running. Otherwise it says the new unit
applies at the next restart.
A unit at that path that the command did not write is left alone, and the
command says the unit differs. Pass --replace to replace it.
To stop the service and remove it:
Memory limits
Every build has a memory limit of its own, and the runner never runs more thanconcurrency builds. So the builds together use at most
concurrency times build.memory.
- A container build runs with Docker’s
--memoryatbuild.memory. - Image builds run in the runner’s own builder, whose container has the same limit. See Image builds.
- Under a service with
Delegate=yes, each microVM build runs in a cgroup of its own. Abovebuild.memoryplus 256 MiB (memory.high), the kernel reclaims the page cache of the microVM’s disks and slows its writes. Abovebuild.memoryplus 1 GiB (memory.max), it kills the build’s microVM.memory.swap.maxis 0. - A build killed at that limit is released
failed, with a reason that says it ran out of memory. The other builds and the runner carry on. - Without
Delegate=yes, for example under a hand-written unit, microVM builds have no limit of their own, andouterlayer doctorwarns. On cgroup v1 the same is true.
concurrency times
build.memory with the memory the runner can use, less 2 GiB. That memory
is the host’s on Linux, the VM’s on macOS and the distribution’s on WSL2.
When it does not fit, setup sets up nothing and names both figures. A
runner started with such a config claims nothing, and its host entry says
why.
Image builds
The runner builds a repository’s image before its first build. It builds it again when the recipe or the CLI changes. Those builds run in the runner’s own builder, a BuildKit container that Buildx makes for it. The container’s memory limit isrunner.build.memory. A recipe step that uses
more is stopped there, so it cannot starve running builds. Docker’s default
builder cannot be limited this way.
What the host needs:
- Docker’s Buildx plugin, version 0.18 or later. It is
docker-buildx-pluginin Docker’s apt repository anddocker-buildxin Ubuntu’s own packages. Lima’sdocker-rootfultemplate has it. - Access to Docker Hub, to pull the BuildKit image once.
- A proxy, if any, that the builder can reach. The builder does not inherit
the Docker daemon’s proxy settings, so the runner passes its own
HTTP_PROXY,HTTPS_PROXYandNO_PROXY. A proxy on the host’s loopback address cannot be reached from the builder. A password in a proxy address shows in the host’s process list, so use a proxy that needs none.
outerlayer doctor fails the “Image build limit” check. See
When the Image build limit check fails.
The runner makes the builder when it is missing. It replaces the builder
when runner.build.memory changes, after the image builds it is running
end. Replacing it clears its cache. outerlayer runner image never
replaces a builder; it stops and says so. A build that hits the limit is
released failed at provision. See
When an image build runs out of memory.
Each new image is copied from the builder into Docker’s image store, which
takes time in proportion to its size. A new CLI release rebuilds each
image’s tools layer once. A copy of each image is also kept under
runner/images/<owner>/<name>/layouts/, and removed with the image.
The outerlayer command
outerlayer runner install writes ~/.local/bin/outerlayer, a short shell
script that runs the CLI in <runner dir>/current. A self-update switches
current to the new version. So the command a shell finds always runs the
version the service runs, and outerlayer runner status, doctor and
logs match the runner. The command prints the path it wrote.
runner install --service and runner init write it too.
The runner never edits a shell startup file. When ~/.local/bin is not on
PATH, the command says so and prints the line to add:
outerlayer comes first on PATH, for example one installed with
npm, runner install and outerlayer doctor name both paths and give the same
line. doctor shows this as the “outerlayer on PATH” check.
A file at that path that the runner did not write is left alone, and the
command says to run outerlayer runner install --replace. A launcher the
runner wrote earlier is rewritten without --replace, so the last install
decides which runner directory it follows.
When the runner claims nothing
At start and on every poll, the runner checks itself, as its own account and never as the person who set it up. It checks in this order:- No queue’s command is still the placeholder
runner initwrites. - For container and microVM hooks,
concurrencytimesbuild.memoryfits in memory, as in Memory limits. - For
builtin:vmhooks,/dev/kvmexists and the runner can open it. - For container and microVM hooks, its image builder exists and starts, as in Image builds.
- For any built-in hooks, it has a Claude credential.
blocked with the reason, and GET /v1/hosts lists the host as not
taking work. Once the check passes, it takes work again on the next poll.
Right after a start, the runner says it is still preparing its image
builder. That is not a fault. It claims nothing until the builder is
ready, then takes work with no action from you.
An account added to the kvm group, or a credential added to runner.env,
needs a service restart first. The runner reads both only when it starts.
See When the runner is blocked.
outerlayer doctor reports what the running service has: the groups of its
main process and the variables in its runner.env. It does not report your
shell’s.
A runner short of disk also claims nothing; see Housekeeping.
Start and stop the runner
outerlayer runner stoptakes no more work, waits for running builds to finish, and exits.systemctl stopand Ctrl-C do the same.outerlayer runner stop --nowends running builds with outcomestopped. A second Ctrl-C does the same.
The host key
A runner key works only from the host that first used it. On first run, the runner makes a host key at~/.outerlayer/runner/host-key.pem,
readable only by its account. The private half never leaves the machine.
The first claim binds the host key to the runner key, and the org audit
log records it. From then on the gateway refuses that runner key unless
the host key signed the request. A leaked runner key is useless without
its host key. A runner too old to sign is refused: upgrade it.
Signatures are valid once, and only within five minutes of the gateway’s
clock, so keep the host’s clock in sync.
When the gateway refuses a signature, see
When the gateway refuses a signature.
Settings → API keys shows each bound key’s host key fingerprint and
when it was bound. outerlayer doctor says whether the runner key is bound
to this host.
Moving a runner to a new host
To move a runner key to a new machine, or after you deletehost-key.pem:
- A member who may update API keys opens Settings → API keys and
chooses Clear host key on the runner key. The org audit log records
runner_key_binding_cleared. - Start the runner on the new host. Its first claim binds the new key.
runner_key_signature_invalid and
the old host keeps working. Do step 2 right after step 1, so nobody else
holding the key can bind it first.
A claim belongs to the runner key that took it; the host name is only a
label. A different key sending the same host name is refused with
claim_not_held.
Several factories, one machine
One config file is one factory. To serve two, run two runners, each with a config file in a directory of its own, such as/etc/outerlayer/acme/config.json. The runner keeps its pid file, host key,
jobs and install layout in runner/ beside its config file. Two config
files in one directory would share them.
Give each config file its own runner key. Each directory has its own
host key, so a key the first runner bound is refused from the second.
- Set
runner.hostin each file, so the two runners have different names. - Shared hook scripts should read
OUTERLAYER_URLandOUTERLAYER_APP_IDrather than assume one factory. - Each runner checks its builds against all of the machine’s memory. Size
the two runners’
concurrencyandbuild.memorytogether. runner install --servicewrites one unit,outerlayer-runner. For the second runner, runouterlayer runner install --config <path>. Then save the output ofouterlayer runner install --service --print --config <path>as a unit with another name, and enable it.
Keep the runner up to date
A runner started from its install layout updates itself to the CLI version its gateway names. It never moves to a lower version. When a higher version is named, the runner stops taking new work and waits until no build is running. It then verifies the version against the npm registry, installs it beside the current one, checks it, and restarts on it. The install layout lives under<config dir>/runner/:
versions/<version>/holds each installed version. The current and the previous one stay on disk.currentlinks to the version the service starts.last-update.jsonholds the last update’s outcome.
GET /v1/hosts lists it as lastUpdate, with
one of these outcomes:
To turn updates off, set
runner.autoUpdate to false. When it is not set,
updates are on if both hooks are built-in, and off if you supply your own
hooks.
Upgrade by hand
A host with updates off, or one that cannot update itself, is upgraded by hand. A host below the gateway’s minimum runner protocol keeps polling but takes no work until you do.-
Install the new CLI with npm:
-
Put it in the install layout with the npm-installed CLI, not the
launcher in
~/.local/bin, which still runs the old version:A host with its own hooks runsrunner installwithout--service. -
Wait.
runner installpointscurrentat the new version and prints that the service restarts on it once its running builds finish. It restarts nothing itself, so nosudo systemctl restartis needed.
current sees the
switch. It takes no new work and waits for its running builds to finish.
Then it exits with status 75, and the service starts the new version. The
install can run as any account that can write the runner’s directory,
because the runner only reads the link.
The new version starts under the update marker, as after a self-update. It
runs the runner’s checks at its first start. If they fail, or it starts
three times without finishing its first poll, it puts the previous version
back as current. last-update.json then records rolled_back. Otherwise
it records updated once its first poll completes.
These runners do not restart:
- A runner started by hand with
outerlayer runner start. Nothing would start it again, so it keeps running and logs that a restart picks the new version up. Runouterlayer runner stopandouterlayer runner start. - A runner whose service starts a version directory instead of
current. A restart would start the same version, so move the host onto the layout first (below). - A runner whose
currentnames a version that is not installed. It logs an error and keeps running its own version.
runner install refuses to install a CLI older than the version current
names unless you pass --allow-downgrade, and a refused install changes
nothing.
Pausing or holding updates for a factory
runner.autoUpdate is one host’s switch. To stop or limit updates on every
host of a factory at once, an owner or admin chooses Runner updates in
the factory’s settings menu. The API takes the same setting as
runner_updates on PATCH /v1/apps/{appId}. It takes effect on each host’s
next heartbeat.
- Follow the gateway (
follow, the default): runners update to the version the gateway was built from. - Pause (
paused): no runner updates. The heartbeat names each runner’s own version back to it. - Hold at a version (
pinned, withversion): runners below it update to it, and none moves past it. The version is a plainmajor.minor.patchno higher than the gateway’s own. A higher one is refused with 422runner_version_above_gateway.
Move an existing host onto the install layout
A host installed without the layout, or whose service starts a version directory instead ofcurrent, cannot update itself. It logs once that it
cannot, naming outerlayer runner install, and keeps taking work. To move
it:
- Run
outerlayer runner install --service [--config <path>]. It puts the CLI in the layout and rewrites the unit to startcurrent. - If it says the unit differs, the unit was not written by this command.
Run it again with
--replace.
outerlayer runner init does this for a new host. runner install needs
the CLI installed with npm, not a source checkout. A host with its own
executables as hooks runs outerlayer runner install without --service,
which prints the ExecStart= line for its own unit.
Housekeeping
A host onbuiltin:container or builtin:vm hooks removes what its own
builds leave behind, so it needs no cleanup script. The runner does this
once when it starts and then once an hour, between polls. It never delays a
running build’s lease renewals. Each removal is one log line that names
what went and why.
The runner removes:
- Job directories under
runner/jobs/of finished builds, once the directory last changed more thanjobsKeepDaysago. A build whose release has not landed yet is kept. - Containers and volumes labelled
ai.outerlayer.buildand with this runner’s ownai.outerlayer.runnerid that no unfinished build names. A failed cleanup leaves these. Another runner’s, on the same Docker daemon, are never touched. Objects made by an earlier version carry no runner id. The runner removes one of those only when a finished job in its ownrunner/jobs/names its build id. - Root disks under
runner/vm/rootfs/whose image Docker no longer has. A file still being written (.partial-in its name) is kept. - Spool files under
runner/spool/older than 24 hours, when no build is running. - Dangling images that carry the runner’s
ai.outerlayer.kindlabel, except one an unfinished build’s job file names. - Build cache of the runner’s own image builder, down to
buildCacheMax.
docker builder prune once.
Set these under runner.housekeeping in the config file:
Free disk. Before it claims, every runner checks free space on the file
system that holds its directory. Below
minFreeDisk it claims nothing, logs
why once, and sends lowDisk with its free and needed bytes in its
heartbeat. GET /v1/hosts then shows the host as not taking work. It claims
again as soon as there is room, with no restart. See
When the runner holds off for disk.
outerlayer runner status shows when housekeeping last ran and what it
freed, and outerlayer doctor warns when free disk is below the minimum.
A host whose hooks are its own executables gets no cleanup pass: its hooks
own what they create. The free-disk minimum still applies to it.
Uninstall the runner
-
Stop the service and remove its unit. Run
outerlayer runner stop --nowfirst if you do not want to wait for running builds. -
As the runner’s account, delete these:
~/.outerlayer/runner, which holds the host key, builds, logs and caches.~/.outerlayer/runner.env, which holds Claude’s credential.~/.outerlayer/config.json, which holds the runner key. Or remove only itsrunnerandapiKeyentries.~/.local/bin/outerlayer, the launcher that ran the CLI in~/.outerlayer/runner/current.~/.outerlayer/cli, ifouterlayer initmade it.
-
Remove the images, volumes and image builder the runner made.
docker buildx lslists the builder, namedouterlayer-and an id. -
Run
npm uninstall -g @outerlayer/cli, if nothing else uses it. - Delete the runner key in Settings → API keys. The gateway refuses it within about five minutes.
outerlayer runner uninstall --service removes the service in the
VM and the launch agent. Then delete the outerlayer-runner VM and remove
runnerVm from the Mac’s ~/.outerlayer/config.json. The runner key and
host key exist only in the VM, so this removes both. Delete the runner key
in Settings → API keys as well.