> ## Documentation Index
> Fetch the complete documentation index at: https://docs.outerlayer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Run a host as a service

> Run the runner under an account and a service manager, move it between machines, upgrade it and remove it.

A host that builds unattended runs `outerlayer runner start` under a
service manager, as an account of its own. Set the host up first with
[Set up your first host](/set-up-a-host).

## The account the runner runs as

The runner never prompts. Under the service's account:

* **Claude's credential** is in `runner.env` beside the config file. Built-in
  hooks use only that file. A terminal `claude` login is not used. See
  [Give the runner Claude's credential](/set-up-a-host#give-the-runner-claudes-credential).
* **git** must fetch with no credential prompt. The runner uses the
  account's own git login for `outerlayer runner image` and to read a
  governing control plane's context. Your own hooks may use it too. Use a
  deploy key, a credential helper with a stored token, or an SSH agent the
  service can reach.
* **With your own hooks, the agent** must already be logged in under this
  account. Claude Code keeps its login per home directory.
* **The factory** must be reachable with the key in the config file.

A prompt makes the build wait until the idle limit ends it with outcome
`idle`. Run `outerlayer runner check`, and your command once by hand under
that account, to catch this early.

Hooks and the command run with the runner's own CLI first on `PATH`, so an
`outerlayer` call in a hook uses the runner's version.

## What the service sets

`outerlayer runner init` ends by setting up the service. You write no unit.
To set it up again, run:

```bash theme={"system"}
outerlayer runner install --service
```

On Linux, it writes one systemd unit at
`/etc/systemd/system/outerlayer-runner.service`, then enables and starts it.
The runner starts at boot. It asks for `sudo` and names each `sudo` command
before it runs. The unit runs as the account that ran the command. Pass
`--user <account>` to name another.

| Where the runner lives | What it sets up |
| - | - |
| Linux | The unit, enabled and started. |
| macOS | The same unit inside the Linux VM `runner init --vm` makes, and the launch agent that starts the VM at login. |
| Windows (WSL2) | The same unit inside WSL2, and a Windows scheduled task that starts the distribution at login and keeps it running. |

It refuses on a Linux host without systemd, and for hooks that are your own
executables. The refusal names the reason. `--print` writes the unit to
stdout and sets up nothing, so you can edit it for those hosts.

The unit holds these lines:

* `ExecStart=` starts the runner from its install layout, so it can update
  itself. See [Keep the runner up to date](#keep-the-runner-up-to-date).
* `EnvironmentFile=-` names `runner.env` beside the config file, readable
  only by the runner's account. It holds Claude's credential,
  `CLAUDE_CODE_OAUTH_TOKEN=<token>` or `ANTHROPIC_API_KEY=<key>`. It also
  holds every secret a destination's header reads and every value
  `build.variables` passes in. Without a Claude credential the runner claims
  nothing. Restart the service after you edit the file, because a service
  reads it only when it starts.
* `Restart=on-failure` and `RestartForceExitStatus=75` restart the runner
  after a crash and after the exit with status 75 that a self-update, or a
  version `runner install` put in place, asks for.
* `TimeoutStopSec=` is `timeLimitMinutes` plus ten minutes, because a stop
  waits for running builds.
* `Requires=docker.service` is set for container and microVM hooks.
* `OOMPolicy=continue` keeps one build's death from stopping the service.
  Without it, systemd stops the whole service, and every other build, when
  the kernel kills any process in it.
* `Delegate=yes` lets the runner give each microVM build a cgroup of its own.

The unit has no `ManagedOOMMemoryPressure=`. The OOM daemon kills a whole
cgroup, which here is every build.

It is a system service, not a user service, so the account's groups are
read again at each start. An account added to `kvm` then needs only a
service restart.

On WSL2 without systemd, setup turns systemd on in `/etc/wsl.conf` and
stops. It asks you to run `wsl --shutdown` once from Windows. Run the
command again after that. If Windows interop is off, setup prints the
`schtasks` command to run in Windows.

The scheduled task names your distribution, which WSL2 gives in
`WSL_DISTRO_NAME`. `sudo` drops that variable, so run setup as yourself,
not under `sudo`. Without the name, setup makes no task. It prints the
`schtasks` command with a placeholder for the name `wsl -l` shows.

## Re-run and remove the service

The runner reads the config file again on every poll, so most edits need
nothing else. Re-run `outerlayer runner install --service` after a change to
`timeLimitMinutes`, to the hooks, or to the account. With a unit that still
matches, it changes nothing and says the service is current.

When the unit changes, it rewrites the unit and reloads systemd. It restarts
the service only while no build is running. Otherwise it says the new unit
applies at the next restart.

A unit at that path that the command did not write is left alone, and the
command says the unit differs. Pass `--replace` to replace it.

To stop the service and remove it:

```bash theme={"system"}
outerlayer runner uninstall --service
```

It removes the unit, and the launch agent on a Mac or the scheduled task on
Windows.

## Memory limits

Every build has a memory limit of its own, and the runner never runs more
than `concurrency` builds. So the builds together use at most
`concurrency` times `build.memory`.

```mermaid theme={"system"}
flowchart TB
    H["Memory the runner can use"] --> R["2 GiB for the runner and the system"]
    H --> B["concurrency × build.memory"]
    B --> V["Each microVM build: build.memory<br/>+256 MiB before writes slow<br/>+1 GiB before it is killed"]
```

* A container build runs with Docker's `--memory` at `build.memory`.
* Image builds run in the runner's own builder, whose container has the same
  limit. See [Image builds](#image-builds).
* Under a service with `Delegate=yes`, each microVM build runs in a cgroup
  of its own. Above `build.memory` plus 256 MiB (`memory.high`), the kernel
  reclaims the page cache of the microVM's disks and slows its writes. Above
  `build.memory` plus 1 GiB (`memory.max`), it kills the build's microVM.
  `memory.swap.max` is 0.
* A build killed at that limit is released `failed`, with a reason that
  says it ran out of memory. The other builds and the runner carry on.
* Without `Delegate=yes`, for example under a hand-written unit, microVM
  builds have no limit of their own, and `outerlayer doctor` warns. On
  cgroup v1 the same is true.

Setup refuses builds that cannot fit. It compares `concurrency` times
`build.memory` with the memory the runner can use, less 2 GiB. That memory
is the host's on Linux, the VM's on macOS and the distribution's on WSL2.
When it does not fit, setup sets up nothing and names both figures. A
runner started with such a config claims nothing, and its host entry says
why.

## Image builds

The runner builds a repository's image before its first build. It builds it
again when the recipe or the CLI changes. Those builds run in the
runner's own builder, a BuildKit container that Buildx makes for it. The
container's memory limit is `runner.build.memory`. A recipe step that uses
more is stopped there, so it cannot starve running builds. Docker's default
builder cannot be limited this way.

What the host needs:

* Docker's Buildx plugin, version 0.18 or later. It is `docker-buildx-plugin`
  in Docker's apt repository and `docker-buildx` in Ubuntu's own packages.
  Lima's `docker-rootful` template has it.
* Access to Docker Hub, to pull the BuildKit image once.
* A proxy, if any, that the builder can reach. The builder does not inherit
  the Docker daemon's proxy settings, so the runner passes its own
  `HTTP_PROXY`, `HTTPS_PROXY` and `NO_PROXY`. A proxy on the host's loopback
  address cannot be reached from the builder. A password in a proxy address
  shows in the host's process list, so use a proxy that needs none.

A host without a usable Buildx claims no work. Its heartbeat says why, and
`outerlayer doctor` fails the "Image build limit" check. See
[When the Image build limit check fails](/troubleshoot-builds#when-the-image-build-limit-check-fails).

The runner makes the builder when it is missing. It replaces the builder
when `runner.build.memory` changes, after the image builds it is running
end. Replacing it clears its cache. `outerlayer runner image` never
replaces a builder; it stops and says so. A build that hits the limit is
released `failed` at `provision`. See
[When an image build runs out of memory](/troubleshoot-builds#when-an-image-build-runs-out-of-memory).

Each new image is copied from the builder into Docker's image store, which
takes time in proportion to its size. A new CLI release rebuilds each
image's tools layer once. A copy of each image is also kept under
`runner/images/<owner>/<name>/layouts/`, and removed with the image.

## The outerlayer command

`outerlayer runner install` writes `~/.local/bin/outerlayer`, a short shell
script that runs the CLI in `<runner dir>/current`. A self-update switches
`current` to the new version. So the command a shell finds always runs the
version the service runs, and `outerlayer runner status`, `doctor` and
`logs` match the runner. The command prints the path it wrote.
`runner install --service` and `runner init` write it too.

The runner never edits a shell startup file. When `~/.local/bin` is not on
`PATH`, the command says so and prints the line to add:

```bash theme={"system"}
export PATH="$HOME/.local/bin:$PATH"
```

When another `outerlayer` comes first on `PATH`, for example one installed with
npm, `runner install` and `outerlayer doctor` name both paths and give the same
line. `doctor` shows this as the "outerlayer on PATH" check.

A file at that path that the runner did not write is left alone, and the
command says to run `outerlayer runner install --replace`. A launcher the
runner wrote earlier is rewritten without `--replace`, so the last install
decides which runner directory it follows.

## When the runner claims nothing

At start and on every poll, the runner checks itself, as its own account and
never as the person who set it up. It checks in this order:

1. No queue's command is still the placeholder `runner init` writes.
2. For container and microVM hooks, `concurrency` times `build.memory` fits
   in memory, as in [Memory limits](#memory-limits).
3. For `builtin:vm` hooks, `/dev/kvm` exists and the runner can open it.
4. For container and microVM hooks, its image builder exists and starts, as
   in [Image builds](#image-builds).
5. For any built-in hooks, it has a Claude credential.

When one fails, the runner claims nothing and logs why once. Its heartbeat
carries `blocked` with the reason, and `GET /v1/hosts` lists the host as not
taking work. Once the check passes, it takes work again on the next poll.

Right after a start, the runner says it is still preparing its image
builder. That is not a fault. It claims nothing until the builder is
ready, then takes work with no action from you.

An account added to the `kvm` group, or a credential added to `runner.env`,
needs a service restart first. The runner reads both only when it starts.
See [When the runner is blocked](/troubleshoot-builds#when-the-runner-is-blocked).

`outerlayer doctor` reports what the running service has: the groups of its
main process and the variables in its `runner.env`. It does not report your
shell's.

A runner short of disk also claims nothing; see [Housekeeping](#housekeeping).

## Start and stop the runner

* `outerlayer runner stop` takes no more work, waits for running builds to
  finish, and exits. `systemctl stop` and Ctrl-C do the same.
* `outerlayer runner stop --now` ends running builds with outcome
  `stopped`. A second Ctrl-C does the same.

## The host key

A runner key works only from the host that first used it. On first run,
the runner makes a host key at `~/.outerlayer/runner/host-key.pem`,
readable only by its account. The private half never leaves the machine.

The first claim binds the host key to the runner key, and the org audit
log records it. From then on the gateway refuses that runner key unless
the host key signed the request. A leaked runner key is useless without
its host key. A runner too old to sign is refused: upgrade it.

Signatures are valid once, and only within five minutes of the gateway's
clock, so keep the host's clock in sync.

When the gateway refuses a signature, see
[When the gateway refuses a signature](/troubleshoot-builds#when-the-gateway-refuses-a-signature).

**Settings → API keys** shows each bound key's host key fingerprint and
when it was bound. `outerlayer doctor` says whether the runner key is bound
to this host.

### Moving a runner to a new host

To move a runner key to a new machine, or after you delete `host-key.pem`:

1. A member who may update API keys opens **Settings → API keys** and
   chooses **Clear host key** on the runner key. The org audit log records
   `runner_key_binding_cleared`.
2. Start the runner on the new host. Its first claim binds the new key.

Until step 1, the new host is refused as `runner_key_signature_invalid` and
the old host keeps working. Do step 2 right after step 1, so nobody else
holding the key can bind it first.

A claim belongs to the runner key that took it; the host name is only a
label. A different key sending the same host name is refused with
`claim_not_held`.

## Several factories, one machine

One config file is one factory. To serve two, run two runners, each with a
config file in a directory of its own, such as
`/etc/outerlayer/acme/config.json`. The runner keeps its pid file, host key,
jobs and install layout in `runner/` beside its config file. Two config
files in one directory would share them.

**Give each config file its own runner key.** Each directory has its own
host key, so a key the first runner bound is refused from the second.

* Set `runner.host` in each file, so the two runners have different names.
* Shared hook scripts should read `OUTERLAYER_URL` and `OUTERLAYER_APP_ID`
  rather than assume one factory.
* Each runner checks its builds against all of the machine's memory. Size
  the two runners' `concurrency` and `build.memory` together.
* `runner install --service` writes one unit, `outerlayer-runner`. For the
  second runner, run `outerlayer runner install --config <path>`. Then save
  the output of `outerlayer runner install --service --print --config <path>`
  as a unit with another name, and enable it.

## Keep the runner up to date

A runner started from its install layout updates itself to the CLI version
its gateway names. It never moves to a lower version. When a higher version
is named, the runner stops taking new work and waits until no build is
running. It then verifies the version against the npm registry, installs it
beside the current one, checks it, and restarts on it.

The install layout lives under `<config dir>/runner/`:

* `versions/<version>/` holds each installed version. The current and the
  previous one stay on disk.
* `current` links to the version the service starts.
* `last-update.json` holds the last update's outcome.

The runner's log says what each update did. The host's heartbeat also
reports the last one, and `GET /v1/hosts` lists it as `lastUpdate`, with
one of these outcomes:

| Outcome | Meaning | What to do |
| - | - | - |
| `updated` | The new version runs and passed its checks. | Nothing. |
| `rolled_back` | The new version failed its checks once it ran, or started three times without finishing its first poll. The runner went back to the previous version. | Read the runner's log for the failing check. The runner does not try that version again until the gateway names another. |
| `verification_failed` | The version's checksum, provenance or signatures did not check out. Nothing was installed. | Nothing on the host. The runner waits for the gateway to name a different version. |
| `not_published` | The registry does not have the version yet. | Nothing. The runner asks again on its next poll. |
| `retrying` | Something on the host stopped the update: a Node version outside the CLI's range, npm older than 9.5, an unreachable registry, a failed `npm install`, or the new version's `runner check` failing. | Fix the cause in the log. The runner retries after 5 minutes, then waits twice as long each time, up to 6 hours. |

To turn updates off, set `runner.autoUpdate` to `false`. When it is not set,
updates are on if both hooks are built-in, and off if you supply your own
hooks.

### Upgrade by hand

A host with updates off, or one that cannot update itself, is upgraded by
hand. A host below the gateway's minimum
[runner protocol](/reference/cli-runner#runner-protocol) keeps polling but
takes no work until you do.

1. Install the new CLI with npm:

   ```bash theme={"system"}
   npm install -g @outerlayer/cli
   ```
2. Put it in the install layout with the npm-installed CLI, not the
   launcher in `~/.local/bin`, which still runs the old version:

   ```bash theme={"system"}
   "$(npm prefix -g)/bin/outerlayer" runner install --service
   ```

   A host with its own hooks runs `runner install` without `--service`.
3. Wait. `runner install` points `current` at the new version and prints
   that the service restarts on it once its running builds finish. It
   restarts nothing itself, so no `sudo systemctl restart` is needed.

On its next poll, a runner its service started through `current` sees the
switch. It takes no new work and waits for its running builds to finish.
Then it exits with status 75, and the service starts the new version. The
install can run as any account that can write the runner's directory,
because the runner only reads the link.

The new version starts under the update marker, as after a self-update. It
runs the runner's checks at its first start. If they fail, or it starts
three times without finishing its first poll, it puts the previous version
back as `current`. `last-update.json` then records `rolled_back`. Otherwise
it records `updated` once its first poll completes.

These runners do not restart:

* A runner started by hand with `outerlayer runner start`. Nothing would
  start it again, so it keeps running and logs that a restart picks the new
  version up. Run `outerlayer runner stop` and `outerlayer runner start`.
* A runner whose service starts a version directory instead of `current`.
  A restart would start the same version, so move the host onto the layout
  first (below).
* A runner whose `current` names a version that is not installed. It logs an
  error and keeps running its own version.

`runner install` refuses to install a CLI older than the version `current`
names unless you pass `--allow-downgrade`, and a refused install changes
nothing.

### Pausing or holding updates for a factory

`runner.autoUpdate` is one host's switch. To stop or limit updates on every
host of a factory at once, an owner or admin chooses **Runner updates** in
the factory's settings menu. The API takes the same setting as
`runner_updates` on `PATCH /v1/apps/{appId}`. It takes effect on each host's
next heartbeat.

* **Follow the gateway** (`follow`, the default): runners update to the
  version the gateway was built from.
* **Pause** (`paused`): no runner updates. The heartbeat names each runner's
  own version back to it.
* **Hold at a version** (`pinned`, with `version`): runners below it update to
  it, and none moves past it. The version is a plain `major.minor.patch` no
  higher than the gateway's own. A higher one is refused with 422
  `runner_version_above_gateway`.

A runner never moves down. Holding a factory at a version lower than a host
runs leaves that host where it is. To resume, choose Follow the gateway again.

### Move an existing host onto the install layout

A host installed without the layout, or whose service starts a version
directory instead of `current`, cannot update itself. It logs once that it
cannot, naming `outerlayer runner install`, and keeps taking work. To move
it:

1. Run `outerlayer runner install --service [--config <path>]`. It puts the
   CLI in the layout and rewrites the unit to start `current`.
2. If it says the unit differs, the unit was not written by this command.
   Run it again with `--replace`.

`outerlayer runner init` does this for a new host. `runner install` needs
the CLI installed with npm, not a source checkout. A host with its own
executables as hooks runs `outerlayer runner install` without `--service`,
which prints the `ExecStart=` line for its own unit.

## Housekeeping

A host on `builtin:container` or `builtin:vm` hooks removes what its own
builds leave behind, so it needs no cleanup script. The runner does this
once when it starts and then once an hour, between polls. It never delays a
running build's lease renewals. Each removal is one log line that names
what went and why.

The runner removes:

* **Job directories** under `runner/jobs/` of finished builds, once the
  directory last changed more than `jobsKeepDays` ago. A build whose
  release has not landed yet is kept.
* **Containers and volumes** labelled `ai.outerlayer.build` and with this
  runner's own `ai.outerlayer.runner` id that no unfinished build names. A
  failed cleanup leaves these. Another runner's, on the same Docker daemon,
  are never touched. Objects made by an earlier version carry no runner id.
  The runner removes one of those only when a finished job in its own
  `runner/jobs/` names its build id.
* **Root disks** under `runner/vm/rootfs/` whose image Docker no longer has.
  A file still being written (`.partial-` in its name) is kept.
* **Spool files** under `runner/spool/` older than 24 hours, when no build is
  running.
* **Dangling images** that carry the runner's `ai.outerlayer.kind` label,
  except one an unfinished build's job file names.
* **Build cache** of the runner's own image builder, down to `buildCacheMax`.

The runner only removes what carries its labels or lives under its own
directory. It never touches anything that belongs to an unfinished build,
however old. When any job file does not parse, or Docker does not answer,
the runner removes nothing from Docker on that pass.

The cap applies to the cache of the runner's own builder only. Docker's
default builder is not touched. Cache that earlier runner versions left
there stays until you run `docker builder prune` once.

Set these under `runner.housekeeping` in the config file:

| Setting | Default | Meaning |
| - | - | - |
| `runner.housekeeping.jobsKeepDays` | `30` | Days to keep a finished build's job directory. |
| `runner.housekeeping.buildCacheMax` | `20g` | The most the runner's image builder keeps in its cache. `"off"` leaves the cache alone. |
| `runner.housekeeping.minFreeDisk` | `runner.build.disk` | Free space the runner needs before it claims work. |

**Free disk.** Before it claims, every runner checks free space on the file
system that holds its directory. Below `minFreeDisk` it claims nothing, logs
why once, and sends `lowDisk` with its free and needed bytes in its
heartbeat. `GET /v1/hosts` then shows the host as not taking work. It claims
again as soon as there is room, with no restart. See
[When the runner holds off for disk](/troubleshoot-builds#when-the-runner-holds-off-for-disk).

`outerlayer runner status` shows when housekeeping last ran and what it
freed, and `outerlayer doctor` warns when free disk is below the minimum.
A host whose hooks are its own executables gets no cleanup pass: its hooks
own what they create. The free-disk minimum still applies to it.

## Uninstall the runner

1. Stop the service and remove its unit. Run `outerlayer runner stop --now`
   first if you do not want to wait for running builds.

   ```bash theme={"system"}
   outerlayer runner uninstall --service
   ```
2. As the runner's account, delete these:
   * `~/.outerlayer/runner`, which holds the host key, builds, logs and
     caches.
   * `~/.outerlayer/runner.env`, which holds Claude's credential.
   * `~/.outerlayer/config.json`, which holds the runner key. Or remove only
     its `runner` and `apiKey` entries.
   * `~/.local/bin/outerlayer`, the launcher that ran the CLI in
     `~/.outerlayer/runner/current`.
   * `~/.outerlayer/cli`, if `outerlayer init` made it.
3. Remove the images, volumes and image builder the runner made. `docker
   buildx ls` lists the builder, named `outerlayer-` and an id.

   ```bash theme={"system"}
   docker image rm --force $(docker image ls -q --filter label=ai.outerlayer.kind)
   docker volume rm $(docker volume ls -q --filter label=ai.outerlayer.build)
   docker buildx rm <the outerlayer- builder>
   ```
4. Run `npm uninstall -g @outerlayer/cli`, if nothing else uses it.
5. Delete the runner key in **Settings → API keys**. The gateway refuses it
   within about five minutes.

On a Mac, `outerlayer runner uninstall --service` removes the service in the
VM and the launch agent. Then delete the `outerlayer-runner` VM and remove
`runnerVm` from the Mac's `~/.outerlayer/config.json`. The runner key and
host key exist only in the VM, so this removes both. Delete the runner key
in **Settings → API keys** as well.

```bash theme={"system"}
limactl delete --force outerlayer-runner
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.