Skip to main content
Use this page when a host is running but a build did not go as expected.

Read the runner’s log

The runner prints one line per event, timestamp first. Failure lines also go to standard error.
#3 is the item number. A line ending continuing or retrying next tick is a failure the runner absorbed. unexpected error: <message> is one it did not expect; it keeps going. A repeating failure is logged once, then every tenth tick as still failing: <message> (N ticks), then recovered: when it stops. A drain says what it is waiting on, and repeats it every thirty seconds. To stop taking work, run outerlayer runner stop. It waits for running builds to finish. outerlayer runner stop --now ends them as stopped.

Outcomes and reasons

Every build ends with an outcome. A failed build often also records a reason naming the cause. outerlayer runner logs <item> prints the newest build’s log. Other reasons are a sentence, or a line a hook wrote; see The cleanup hook.

When the gateway refuses a signature

Each refusal is a 401, logged to the org audit log as runner_key_signature_refused.

When a call to the platform does not land

The runner retries the claim, renew and release for up to two minutes when the gateway does not answer. Refusals are never retried. A release that still fails is logged without released:
runner status --recent marks it (release failed), and the next runner start sends it again. The claim lapses fifteen minutes after the last renewal, and the item returns to the queue. If someone released the claim first, the runner logs lease was already released, with the outcome the platform stored and the outcome this attempt ended with. The platform keeps the outcome it stored first.

When the runner holds off for disk

A host that takes no work and logs this has less free disk than it needs:
The runner logs it once, keeps polling and keeps renewing the leases of builds that already run. outerlayer runner status shows the same line, and outerlayer doctor warns about it. GET /v1/hosts lists the host with lowDisk, the free and needed bytes. Its status is not changed by this. The runner measures the file system that holds its own directory, <config dir>/runner. Free space there, or lower runner.housekeeping.minFreeDisk. When there is room the runner logs free disk is back above the minimum · taking work again and claims on its next poll, with no restart. By default the minimum equals runner.build.disk: room for one more build. Housekeeping removes what builds leave behind, which is usually enough. If the disk is full of something else, such as Docker images a person pulled, the runner does not remove it. See Housekeeping.

When the runner is blocked

A host that takes no work and logs claiming nothing: followed by a reason has failed one of its own checks. What the runner checks, and how it reports a failed check, is in When the runner claims nothing. A typical line:
Once the check passes, the runner logs the runner can run its work again · taking work and claims on its next poll. Two fixes take effect only after a service restart: /dev/kvm access and the Claude credential. The others take effect on the next poll. outerlayer doctor reports the KVM group and the credential of the running service, not of your shell.

When the Image build limit check fails

The runner builds every image in its own builder, a container with a memory limit of runner.build.memory, which needs Docker’s Buildx plugin. When Buildx or its docker-container driver is missing, or older than 0.18, outerlayer doctor fails the “Image build limit” check, and the runner claims nothing. It logs claiming nothing: with the reason, and its heartbeat carries the same reason. Install the plugin: docker-buildx-plugin from Docker’s apt repository, or docker-buildx from Ubuntu’s own packages. Then run outerlayer doctor again. The runner takes work again by itself, with no restart. A runner that has Buildx but is still starting its builder logs the runner is still preparing its image builder. The first start pulls the BuildKit image, which takes a moment. If the reason does not change, it names why the builder did not start, for example that Docker Hub cannot be reached. The builder container does not use a proxy on the host’s loopback address, so a daemon that reaches the network only through one needs a proxy the container can reach.

When an image build runs out of memory

A recipe step that uses more than runner.build.memory is stopped by the kernel, in the runner’s image builder. The build is released failed at the provision stage with this reason:
Only that build ends. The builder and the other builds carry on. The log named in the reason shows the step. Raise runner.build.memory, and check that concurrency times build.memory still fits in the host’s memory less 2 GiB. When the limit changes, the runner replaces its builder once the image builds it is running have ended, and the next image build starts with an empty cache.

When a build runs out of memory

In a service that sets Delegate=yes, each microVM build’s Firecracker process runs in a cgroup of its own. When the kernel kills it at the cgroup’s limit, the build is released failed with this reason:
Only that build ends. The runner and the other builds carry on. Raise runner.build.memory, and check that concurrency times build.memory still fits in the host’s memory less 2 GiB. A build killed below its limit means this host needs more than the 1 GiB reserved for Firecracker or the 2 GiB reserved for the system. Lower runner.concurrency to leave more room. A service without Delegate=yes gives microVM builds no limit of their own, and outerlayer doctor warns. Run outerlayer runner install --service --replace to write the unit this CLI generates. A container build’s limit is Docker’s --memory at build.memory.

When an item comes back

A build request is used up by any claim after it: a host build released ok, incomplete, failed, timed_out or idle, a claim that lapsed, or a local session. So a failed build is not claimed again. To retry, run outerlayer work build, or select Build again on a host in the Last build block of the item page. The control is offered only when the item has no open pull request, no running build and no unused request. work build --local does not queue the item. interrupted, stopped, lease_lost and exit 75 leave the request unused, so the item stays available. An item that still needs work comes back through the amend queue when a person sends notes or fails it, not through the implement queue. A failed amend build is not retried until a person fails the item again. While no session is linked, the item’s Sessions list names who holds a live claim: a host by name, or a person’s name followed by “(local session)”. When nothing is queued it names outerlayer work build. Run it yourself gives the command that starts a session by hand.