> ## Documentation Index
> Fetch the complete documentation index at: https://docs.outerlayer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Policy and validators

> Declare what evidence a change must carry, and how strict the verdict is, from files on the base branch.

Every pull request is evaluated against a set of checks. The policy decides which checks run and how much each counts.

The checks appear on the work item page's **Checks** tab. On GitHub, they appear as one evidence comment. Each row on it is a check. GitHub also gets one `OuterLayer evidence` check run. It reads two things side by side: whether OuterLayer's own checks pass, and where a person's review stands. OuterLayer writes both only when **Post session summaries on pull requests** is on, in **Settings → General**. It is off by default. See [Connect a repository](/connect-repository#what-the-connection-is-used-for).

The verdict is one of these:

* **pass**: every check shown passed.
* **flagged**: something needs attention.
* **waiting**: nothing failed, but a result has not arrived. The comment says whether it waits on a required result or session evidence.

A person's review is not one of the checks, so it never makes the verdict flagged or waiting. The comment shows the review in its own block above the checks.

Two places hold a repository's policy: the policy file, `.outerlayer/policy.yaml`, and the validator files in `.outerlayer/validators/`. Both are read from the pull request's base branch, so a change takes effect when it merges. A few keys below are read from the default branch instead, and say so. When a factory names a context source, the files come from that repository, and a change applies at the next evaluation. See [Share instructions across repositories](/share-instructions-across-repositories).

## The policy file

`.outerlayer/policy.yaml`:

```yaml theme={"system"}
validators:
  code-review: off
merge_gate: on-flag
review: required
build:
  workflows: allow
criteria:
  replace: people
  missing: warn
  test_proof: any
```

| Key | Values |
| - | - |
| `validators` | Per-check level overrides. `warn` counts toward the verdict. `info` shows the row but never counts. `off` removes the row. |
| `merge_gate` | How strict the `OuterLayer evidence` check run on GitHub is. `none` (default): the check is always neutral, so it never blocks a merge. `on-flag`: the check fails when one of OuterLayer's checks fails, and when a person has requested changes. A missing review does not fail it. Any other value is a policy error, and the check behaves as `none`. See [The OuterLayer evidence check](#the-outerlayer-evidence-check). |
| `review` | `required` (default, also with no policy file) means the check's title says when a person's review is still pending. `optional` says no review is required. Neither value holds the verdict or the check open, and a person's request for changes fails the check under `merge_gate: on-flag` with either value. See [Review and merge](/review-the-work). |
| `build.workflows` | `deny` (default) or `allow`. `allow` lets a build's push token change files under `.github/workflows/`. Read from the default branch. See [GitHub App permissions and build tokens](/github-app-and-build-tokens). |
| `criteria.replace` | Who may replace an item's recorded acceptance criteria once it has a list. `people` (default): only a person; an agent session, or a key bound to no member, gets `409`. `anyone`: those callers may replace it too. Read from the default branch. See [`outerlayer emit criteria`](/reference/cli-emit#outerlayer-emit-criteria). |
| `criteria.test_proof` | Where test results must come from to prove a criterion that declares `test` proof. `any` (default) takes results from any source. `ci` takes only results recorded from CI, so an agent session's own run proves nothing. See [Attach evidence](/attach-evidence#prove-a-criterion-with-test-results). |
| `criteria.missing` | What a pull request does when its item has no recorded criteria: `warn` (default), `info` or `off`. `warn` flags the verdict. `info` shows the row without flagging. `off` drops it. A repository with no policy file is never flagged for it. See [Recorded criteria in the verdict](#recorded-criteria-in-the-verdict). |

To let agent sessions replace a recorded criteria list, set `criteria: replace: anyone`.

Every key is optional. An empty file sets nothing, and every key takes its default.

An unknown key or an invalid value voids the whole file for that evaluation. The comment shows a policy error row. Your overrides drop out, and every key takes its default. The validator files still apply. Fix the file on the base branch.

`extends` is not a key. OuterLayer has no preset to extend: every check comes from a validator file. A file that sets `extends`, including `extends: outerlayer:recommended@v1`, gets a policy error row that names the fix. Run `outerlayer init --template default`, which writes `.outerlayer/validators/code-review.yaml` if it is missing, then delete the `extends` line.

## The OuterLayer evidence check

The title of the `OuterLayer evidence` check says what decided it. With `merge_gate: on-flag`:

| Title | Result |
| - | - |
| Waiting on OuterLayer checks: *names* | In progress. The checks are still running. |
| *n* OuterLayer check failing: *name* | Fails. |
| Changes requested by *name* · OuterLayer checks pass | Fails. A person asked for changes in OuterLayer, whatever `review` says. |
| OuterLayer checks pass · review pending in OuterLayer | Passes. Nobody has reviewed yet. |
| OuterLayer checks pass · approved by *name* | Passes. |
| OuterLayer checks pass · review pending: new commits since *name*'s approval | Passes. The approval predates the latest commits. |
| OuterLayer checks pass · review not required | Passes. The policy has `review: optional`. |

A person's "no" blocks a GitHub merge under `merge_gate: on-flag`. A missing review does not. With `merge_gate: none`, every completed result is neutral and blocks nothing, but the titles are the same.

A pull request nobody ran through OuterLayer is never evaluated and gets no comment. It can still get a neutral `OuterLayer evidence` check titled "Not run through OuterLayer", which says nothing was checked and never blocks a merge. It gets that check when the base branch's policy has `review: required` (the default) or a `merge_gate` other than `none`, and **Post session summaries on pull requests** is on. Otherwise OuterLayer writes nothing to it.

Every check comes from a file in `.outerlayer/validators/`. OuterLayer defines none of its own, so a repository with no validator files has nothing checked, with or without a policy file.

## The code review check

`outerlayer init` writes a policy file and one check, `.outerlayer/validators/code-review.yaml`. So does `outerlayer init --template default`. Each file is written only when it is missing, and is yours to edit. Without its comments, it reads:

```yaml theme={"system"}
id: code-review
kind: validation
row: "A code review result is recorded for this change"
level: warn
run:
  where: ci
  emit: code-review
```

The check passes when a pass is recorded under `code-review` for the work item. After a review, emit its report as an artifact, run `outerlayer sync`, and copy the report's deep link from the evidence comment. Then record the verdict with that link as its proof:

```bash theme={"system"}
outerlayer emit artifact review.html --caption "Review of the change and what it found"
outerlayer sync
outerlayer emit code-review --item 412 --result pass --link <the report's deep link>
```

Record `--result fail` when the review left something unfixed. A session, CI or a person can record it. See [Record a check](/record-a-check).

The row says a result is recorded, not what the review found. While a host still holds the item's claim, a missing result reads as waiting. Afterwards it reads as not proven.

Level it under `validators:`, for example `code-review: off`, or delete the file.

A level for `code-review-ran`, the id this check had when the product defined it, is ignored while no validator file declares that id.

A policy file written for the retired `acceptance-criteria` check keeps loading. Its level is ignored, and a validator file of yours may take that id. A report bound with `--for acceptance-criteria` is not a check any more: the item page links it at the top of its **Criteria** tab.

No check reads a CI result unless you declare one; the item page shows GitHub's own status instead. To put a CI result in the verdict, see [Checks that run in CI](#checks-that-run-in-ci).

Some rows have no level and cannot be turned off: policy errors, an unreadable policy, a person's pass or fail on the item, rows a linked issue asks for, and `criterion-unproven`.

## Recorded criteria in the verdict

An item's recorded criteria count toward the verdict of its open and merged pull requests. A pull request closed without merging is not flagged.

* **An unproven criterion** is a recorded criterion whose declared proof is not attached. The row is `criterion-unproven` and it always counts toward the verdict; no policy key lowers it. A team that does not want a proof kind does not declare it. A criterion that declares no proof never produces one.
* **No recorded list** produces `criteria-missing`, which `criteria.missing` levels. `warn` flags the verdict. `info` records the fact without flagging it. `off` drops it. Any other value is a policy error.

```yaml theme={"system"}
criteria:
  missing: info
```

While a host holds the item's claim, both read as pending, so a build in progress is not flagged. They flag once the claim ends. A merged pull request keeps the flag on its final record.

A flagged row has the effects any flagged row has:

* **Approve** stays disabled on the item page, though the item still reads **Ready for review**.
* A runner releases the build as incomplete.
* The GitHub check fails when `merge_gate: on-flag` is set.

The Checks tab adds no row for either one.

A repository with a policy file whose builds do not record criteria has its open pull requests flagged. Set `criteria.missing` to `info` or `off`, or record criteria with [`outerlayer emit criteria`](/reference/cli-emit#outerlayer-emit-criteria).

## Custom validators

Each file in `.outerlayer/validators/` declares one check. Only `.yaml` and `.yml` files are read, and only the first twenty by name. More files raise a policy error row.

```yaml theme={"system"}
id: e2e-ran
kind: validation
row: "The end-to-end suite ran and passed in this session"
level: warn
when:
  paths: ["apps/web/**"]
require:
  session.ran:
    command: "playwright test"
    status: ok
```

| Key | Required | Meaning |
| - | - | - |
| `id` | yes | A lowercase letter, then lowercase letters, digits, `.`, `-` or `_`, up to 64 characters. Unique in the directory, and not a [reserved id](#reserved-ids). |
| `kind` | yes | `validation` renders a row. `signal` reserves the id and renders nothing. |
| `row` | for `validation` | The sentence the row shows. One line, at most 200 characters. Claim only what the requirement proves. |
| `name` | no | A short label for the row, at most 40 characters. |
| `level` | no | `warn` (default), `info` or `off`. |
| `when` | no | Scope. When it does not match, the check renders no row. |
| `require` or `run` | exactly one | What satisfies the check. |
| `needs` | no | `commands` or `edits`: the recorded detail the check reads. A session recorded without it reads as not checkable, not failed. |

### Scoping with `when`

| Key | Matches when |
| - | - |
| `paths` | Any changed file matches a glob (at most 20). `**` matches any depth, `*` one segment. |
| `issue.type` | The linked issue's type matches, case-insensitive. |
| `issue.labels` | The linked issue carries any of the labels. |

A file scoped on the issue renders no row when no issue is linked.

### Requirements

`require` names one requirement. To accept any of several, list them under `any`:

```yaml theme={"system"}
require:
  any:
    - session.ran:
        command: "yarn test"
    - validator: e2e-ran
    - emitted: e2e
```

| Alternative | Satisfied when |
| - | - |
| `session.ran` | A recorded session ran a command starting with those words, with the given `status` (`ok` by default, or `error`). Matching is whole-word prefix: `playwright` matches `playwright test e2e`. A run with `--help`, `--dry-run` or `--version` never counts. |
| `validator` | Another custom validator in the same directory passed on this pull request. |
| `artifact` | An artifact of `kind` (`report`, `screenshot`, `video`, `log`, `test` or `file`), emitted with `--for` set to `for`, is attached to the pull request. Both keys are required. Nothing attached yet reads as waiting while a host still holds the work item's claim, and as failed afterwards. See [Attach evidence](/attach-evidence). |
| `emitted` | A check under that name was recorded as a pass with `outerlayer emit <name>`. A file in the same directory must declare the name in its `run:` block, with `where: ci` or `where: host`. Only that declaring check requires a host's result; an `emitted:` requirement accepts a pass from any recorder. Nothing recorded yet reads as waiting while a host still holds the work item's claim, and as failed once that claim is released or expires. |

### Checks that run in CI

`run` declares a check whose result comes from `outerlayer emit`:

```yaml theme={"system"}
id: e2e-in-ci
kind: validation
row: "The end-to-end suite passed in CI for this change"
run:
  where: ci
  emit: e2e
```

A `where: ci` check takes no `command`. OuterLayer never runs it. Record the result with `outerlayer emit e2e --result pass --item 412`, from CI or a shell. See [Record a check](/record-a-check).

The emit name `criteria` is reserved, because `outerlayer emit criteria` records acceptance criteria. Choose another name.

### Checks that run on your host

`where: host` declares a command your own [host](/reference/cli-runner) runs after a build:

```yaml theme={"system"}
id: unit-tests
kind: validation
row: "The unit tests pass on a clean checkout of the branch"
run:
  where: host
  emit: unit-tests
  command: "yarn test --ci"
```

`command` is required for `where: host` and not allowed for `where: ci`. It is a shell line. Exit status 0 is a pass and any other status is a fail.

After a build succeeds, the host runs each command in a new environment that holds no OuterLayer key. Only a result recorded by a host satisfies the check: a pass recorded under the same name by an item key, the build session or a person is ignored. The runner reference describes [the steps, the time limit and each failure reason](/reference/cli-runner#host-checks).

#### What a host check can read

Before a command runs, the host writes a folder and sets `OUTERLAYER_EVAL_INPUT` to its path. The folder is read-only: a command cannot write into it or change a file in it. Read the files and call nothing. The environment holds no OuterLayer key, so a command cannot ask OuterLayer for this data.

| File | What it holds |
| - | - |
| `work-item.json` | The work item as OuterLayer returned it to the host, as one JSON object. It is absent when the host could not read the item. |
| `diff.patch` | The branch's diff, in git's patch format, from the merge base of the default-branch commit the build started from and the branch's head, to the head. A branch the build continued from before the default branch moved on shows only its own changes. It is empty when the host could not read the commits. |
| `sessions/<trace-id>.jsonl` | One transcript per session, subagents included, as one JSON object per line. The trace id matches the `traceId` of the session in `work-item.json`. |
| `activity.json` | The commands and edits of those sessions, and what could not be read. |

The transcripts come from the build environment, where the agent ran. The host copies them out before it stops that environment, and removes secrets from them before the command can read them. The check environment is a new one and never held them.

`activity.json` is one JSON object. Its commands and edits come from the same code the evidence rules use, so a script and a `session.ran` alternative see the same commands, in the same order. The host reads the transcripts in full, whatever capture tier your tenant stores. When your tenant stores less than full, a rule sees fewer commands than the script does.

| Key | Value |
| - | - |
| `sessions` | One entry per session: `traceId`, `subagent`, `status` (`copied` or `missing`), `file` for a copied one, and `reason` for a missing one. When further transcripts hold the same session id, the host copies the first and adds `duplicates`, the number it left out, to that entry. The session is still `copied`, never `missing`. A transcript the host could not copy is `missing`, and the command still runs. The host copies at most 500 transcripts, 7 MiB each and 64 MiB in all, within two minutes. A transcript past those limits is `missing`, and its `reason` names the limit. |
| `sessionsUnavailable` | Present only when the host could not list the transcripts at all. It holds the reason, and `sessions` is empty. |
| `commands` | The commands that ran, in order. Each has `seq`, `traceId`, `turn`, `command`, `kind` (`test`, `lint`, `build`, `vcs`, `migration` or `other`) and `status` (`ok`, `error` or `rejected`). `command` is the command with wrappers such as `cd ... &&`, `yarn` and `npx` removed. A command run by a subagent carries the `traceId` of the session that launched it. |
| `edits` | The file edits, in order. Each has `seq`, `traceId`, `turn`, `file` and `status`. `seq` orders commands and edits together. |
| `recorded` | `{ "commands": boolean, "edits": boolean }`. When `recorded.commands` is `false`, command detail was not recorded, and `commands` is `null`. Do not read it as "no command ran": a session that ran no command looks the same. |
| `diff` | `status` (`written` or `unavailable`), the `base` commit the build started from (the patch itself starts at the merge base of `base` and `head`) and the `head` commit, `baseSource` (`build`, or `recipe` for a job an older runner started), and a `reason` when unavailable. |
| `workItem` | `status` (`written` or `unavailable`), and a `reason` when unavailable. |

The folder is present for every `where: host` command, at the same path.

## Reserved ids

`commits-from-sessions`, `red-then-green`, `no-test-tampering` and `tests-after-last-edit` are reserved. A `validators:` entry naming one is ignored. A custom validator file cannot use one as its `id` or name it in `require`.

## When a file fails to load

A validator file loads whole or not at all. These drop it:

* an unknown key, a bad level, or both `require` and `run`
* a `where: host` check with no `command`, or a `command` on a `where: ci` check
* a `require.validator` naming no file in the directory, or a reserved id
* an `emitted:` name no file declares with `run`
* a reserved emit name or id

The evidence comment shows a policy error row naming the file and the problem. The rest of the comment still renders. Fix the file on the base branch.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.