> ## Documentation Index
> Fetch the complete documentation index at: https://docs.outerlayer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshoot builds

> Read the runner's log, look up how a build ended, and fix refused signatures, releases that do not land and items that come back.

Use this page when a host is running but a build did not go as expected.

## Read the runner's log

The runner prints one line per event, timestamp first. Failure lines also
go to standard error.

```text theme={"system"}
2026-09-22T14:46:22Z claimed implement #3 acme/api#5 (claim c39386da…) → jobs/implement-3-c39386da…
2026-09-22T14:46:25Z #3 provision ok (3s) · workdir jobs/implement-3-c39386da…/work
2026-09-22T14:58:40Z #3 command exited 0 (12m 15s)
2026-09-22T14:59:12Z #3 sync failed: gateway answered 503 · continuing
2026-09-22T14:59:13Z #3 released · outcome ok · 12m 51s
2026-09-22T15:02:00Z list failed: not authorized (401) · retrying next tick
```

`#3` is the item number. A line ending `continuing` or `retrying next tick`
is a failure the runner absorbed. `unexpected error: <message>` is one it
did not expect; it keeps going.

A repeating failure is logged once, then every tenth tick as
`still failing: <message> (N ticks)`, then `recovered:` when it stops. A drain says what it is waiting on, and
repeats it every thirty seconds.

To stop taking work, run `outerlayer runner stop`. It waits for running
builds to finish. `outerlayer runner stop --now` ends them as `stopped`.

## Outcomes and reasons

Every build ends with an outcome. A failed build often also records a
`reason` naming the cause. `outerlayer runner logs <item>` prints the newest
build's log.

| Outcome or reason | Meaning and what to do |
| - | - |
| `ok` | The command exited zero and the item's checks pass. Judge the work by the pull request. |
| `incomplete` | The command exited zero, but the item has no pull request, or its checks still fail or have no result. The reason says which. If the build opened no pull request, it needs to open one with `outerlayer work open-pr`. Otherwise fix the checks on the pull request. Either way, you can run `outerlayer work build` again, or select **Build again on a host** on the item page when no pull request is open. |
| `failed` | The command exited non-zero, or provision failed: the hook exited non-zero, wrote no `workdir=` line, or named a missing directory. Read the log and reason, fix it, then run `outerlayer work build` again, or select **Build again on a host** on the item page. |
| `timed_out` | Ran past `runner.timeLimitMinutes` (480). Raise it or split the item. |
| `idle` | No log output for `runner.idleLimitMinutes` (30) and, with a control group, under two seconds of CPU a minute. Usually a prompt. Make the command non-interactive. |
| `lease_lost` | The runner lost its claim. Nothing to do; the item stays available. |
| `interrupted` | The runner stopped mid-build. Nothing to do; the item stays available. |
| `stopped` | An operator ran `runner stop --now` or pressed Ctrl-C twice. The item stays available. |
| No outcome (exit 75) | The provision hook declined. The item stays queued. |
| `agent_credential_missing` | No Claude credential; [give the runner one](/set-up-a-host#give-the-runner-claudes-credential) and restart. |
| `disk_limit_exceeded` | The build outgrew [`runner.build.disk`](/container-builds#the-build-block). Raise it. |
| `workflow_permission_required` | GitHub refused a push that changed a workflow; see [build tokens](/container-builds#the-git-remote-and-the-builds-tokens). |
| A gateway code, such as `repository_not_in_installation` | The read token was refused; see [build tokens](/container-builds#the-git-remote-and-the-builds-tokens). |
| `context_unavailable`, `context_conflict`, `context_adoption_failed`, `context_adoption_conflict` | Governed context failed; see [governed context](/container-builds#a-governed-repositorys-context). |
| `recipe_ambiguous` | Only named dev container configs. Add `.devcontainer/devcontainer.json`. |
| `recipe_needs_docker` | The recipe uses Compose; see [why it is refused](/build-image-from-devcontainer#why-compose-recipes-are-refused). |
| `recipe_context_outside` | The recipe reads outside `.devcontainer/`, or it is a symlink. Commit a real directory and move the build files into it. |
| `repository_unreadable` | Give the host's git login read access. |
| `default_branch_unknown` | Set the repository's default branch on the git host. |
| `invalid_repository` | The item's repository is not `owner/name`. Fix it. |
| `build_in_progress` | Another build of the same image ran too long. Retry when it finishes. |
| `build_failed` | The image build failed. Read the image build log the runner names. |
| `devcontainer_cli_missing` | Reinstall `@outerlayer/cli`. |
| A claim refusal code, such as `pull_request_from_fork` | The gateway refused the claim; see [Refusals](/reference/cli-runner#refusals). |

Other reasons are a sentence, or a line a hook wrote; see
[The cleanup hook](/write-your-own-hooks#the-cleanup-hook).

## When the gateway refuses a signature

Each refusal is a `401`, logged to the org audit log as
`runner_key_signature_refused`.

| Code | Meaning and what to do |
| - | - |
| `runner_key_signature_required` | No signature. Upgrade the CLI, and check that `<config dir>/runner/host-key.pem` exists. |
| `runner_key_signature_invalid` | Usually a runner key used from a second machine, or a body changed after signing. If the key moved on purpose, [clear its binding](/run-a-host-as-a-service#moving-a-runner-to-a-new-host); otherwise treat it as leaked: revoke it and make a new one. |
| `runner_key_signature_stale` | The host's clock is over five minutes off. Fix and synchronise it. |
| `runner_key_signature_replayed` | Something is resending the runner's requests. Restart the runner and find it. |

## When a call to the platform does not land

The runner retries the claim, renew and release for up to two minutes when
the gateway does not answer. Refusals are never retried.

A release that still fails is logged without `released`:

```text theme={"system"}
2026-09-22T15:02:41Z #16 release failed after 6 attempt(s): the store did not answer (503); the lease expires at 15:14 UTC
```

`runner status --recent` marks it `(release failed)`, and the next
`runner start` sends it again. The claim lapses fifteen minutes after the
last renewal, and the item returns to the queue.

If someone released the claim first, the runner logs
`lease was already released`, with the outcome the platform stored and the
outcome this attempt ended with. The platform keeps the outcome it stored
first.

## When the runner holds off for disk

A host that takes no work and logs this has less free disk than it needs:

```text theme={"system"}
holding off new work: 2.0 GB free on the runner's file system, 40.0 GB needed (runner.housekeeping.minFreeDisk) · claiming again once there is room
```

The runner logs it once, keeps polling and keeps renewing the leases of
builds that already run. `outerlayer runner status` shows the same line, and
`outerlayer doctor` warns about it. `GET /v1/hosts` lists the host with
`lowDisk`, the free and needed bytes. Its status is not changed by this.

The runner measures the file system that holds its own directory,
`<config dir>/runner`. Free space there, or lower
`runner.housekeeping.minFreeDisk`. When there is room the runner logs
`free disk is back above the minimum · taking work again` and claims on its
next poll, with no restart. By default the minimum equals
`runner.build.disk`: room for one more build.

Housekeeping removes what builds leave behind, which is usually enough. If
the disk is full of something else, such as Docker images a person pulled,
the runner does not remove it. See [Housekeeping](/run-a-host-as-a-service#housekeeping).

## When the runner is blocked

A host that takes no work and logs `claiming nothing:` followed by a reason
has failed one of its own checks. What the runner checks, and how it reports
a failed check, is in
[When the runner claims nothing](/run-a-host-as-a-service#when-the-runner-claims-nothing).
A typical line:

```text theme={"system"}
claiming nothing: the runner cannot open /dev/kvm, so builtin:vm cannot start a microVM. Add the runner's account to the kvm group, then restart the service, because a service reads its groups only when it starts
```

Once the check passes, the runner logs
`the runner can run its work again · taking work` and claims on its next
poll. Two fixes take effect only after a service restart: `/dev/kvm` access
and the Claude credential. The others take effect on the next poll.

| The reason names | What to do |
| - | - |
| `/dev/kvm does not exist` | The machine has no KVM. Run microVM builds on a host with KVM enabled, or set `runner.hooks.provision` and `runner.hooks.cleanup` to `builtin:container`. |
| `/dev/kvm`, which the runner cannot open | Add the runner's account to the `kvm` group, then restart the service: `sudo usermod -aG kvm <account>` and `sudo systemctl restart outerlayer-runner`. A service reads its groups only when it starts, so a shell that shows the group says nothing about the service. Or set `runner.hooks.provision` and `runner.hooks.cleanup` to `builtin:container`. |
| Claude credential | Put `CLAUDE_CODE_OAUTH_TOKEN` or `ANTHROPIC_API_KEY` in the `runner.env` beside the config file, then restart the service. The runner reads the file only when it starts. |
| `runner.commands.<queue>` is still the placeholder | Edit the command in the config file to start your agent. The runner reads the file again on its next poll. |
| The runner is still preparing its image builder | Nothing, at first: the first start pulls the BuildKit image. If the reason does not change, see [When the Image build limit check fails](#when-the-image-build-limit-check-fails). |
| Buildx, or an image builder that is not usable | Install Docker's Buildx plugin. See [When the Image build limit check fails](#when-the-image-build-limit-check-fails). |
| `concurrency` times `build.memory` | Lower `runner.concurrency` or `runner.build.memory`, or run the runner on a host with more memory. The reason names the memory the builds need and the memory available, which is the runner's memory less 2 GiB. |

`outerlayer doctor` reports the KVM group and the credential of the running
service, not of your shell.

## When the Image build limit check fails

The runner builds every image in its own builder, a container with a memory
limit of `runner.build.memory`, which needs Docker's Buildx plugin. When
Buildx or its `docker-container` driver is missing, or older than 0.18,
`outerlayer doctor` fails the "Image build limit" check, and the runner
claims nothing. It logs `claiming nothing:` with the reason, and its heartbeat
carries the same reason.

Install the plugin: `docker-buildx-plugin` from Docker's apt repository, or
`docker-buildx` from Ubuntu's own packages. Then run `outerlayer doctor`
again. The runner takes work again by itself, with no restart.

A runner that has Buildx but is still starting its builder logs
`the runner is still preparing its image builder`. The first start pulls the
BuildKit image, which takes a moment. If the reason does not change, it
names why the builder did not start, for example that Docker Hub cannot be
reached. The builder container does not use a proxy on the host's loopback
address, so a daemon that reaches the network only through one needs a proxy
the container can reach.

## When an image build runs out of memory

A recipe step that uses more than `runner.build.memory` is stopped by the
kernel, in the runner's image builder. The build is released `failed` at the
`provision` stage with this reason:

```text theme={"system"}
the image build ran out of memory: the kernel stopped a build step at the limit of runner.build.memory (8g). Raise runner.build.memory, or lower what the recipe's build steps use; the build log is at <path>
```

Only that build ends. The builder and the other builds carry on. The log
named in the reason shows the step. Raise `runner.build.memory`, and check
that `concurrency` times `build.memory` still fits in the host's memory less
2 GiB. When the limit changes, the runner replaces its builder once the image
builds it is running have ended, and the next image build starts with an empty
cache.

## When a build runs out of memory

In a service that sets `Delegate=yes`, each microVM build's Firecracker
process runs in a cgroup of its own. When the kernel kills it at the cgroup's
limit, the build is released `failed` with this reason:

```text theme={"system"}
the build ran out of memory: the kernel killed its microVM at the limit of runner.build.memory (8g) plus 1 GiB for Firecracker and its disks' page cache. Raise runner.build.memory, or lower runner.concurrency
```

Only that build ends. The runner and the other builds carry on. Raise
`runner.build.memory`, and check that `concurrency` times `build.memory` still
fits in the host's memory less 2 GiB. A build killed below its limit means
this host needs more than the 1 GiB reserved for Firecracker or the 2 GiB
reserved for the system. Lower `runner.concurrency` to leave more room.

A service without `Delegate=yes` gives microVM builds no limit of their own,
and `outerlayer doctor` warns. Run `outerlayer runner install --service --replace` to write the unit this CLI generates. A container build's limit is
Docker's `--memory` at `build.memory`.

## When an item comes back

A build request is used up by any claim after it: a host build released
`ok`, `incomplete`, `failed`, `timed_out` or `idle`, a claim that lapsed, or a local
session. So a failed build is not claimed again. To retry, run
`outerlayer work build`, or select **Build again on a host** in the **Last
build** block of the item page. The control is offered only when the item has
no open pull request, no running build and no unused request. `work build --local` does not queue the item.

`interrupted`, `stopped`, `lease_lost` and exit 75 leave the request
unused, so the item stays available.

An item that still needs work comes back through the amend queue when a
person sends notes or fails it, not through the implement queue. A failed
amend build is not retried until a person fails the item again.

While no session is linked, the item's **Sessions** list names who holds a
live claim: a host by name, or a person's name followed by "(local session)". When nothing is
queued it names `outerlayer work build`. **Run it yourself** gives the
command that starts a session by hand.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.