onspec documentation

Everything here describes v1 behavior exactly. If the tool and this page disagree, that's a bug: tell us.

Getting started

Install the CLI, write one spec, and verify a change against it.

$ npm install -g onspec

# Have existing code? Draft specs from it instead of starting blank:
$ onspec reverse

# Grade your specs for verifiability:
$ onspec lint

# Check the current change against its governing specs:
$ onspec verify --base origin/main --test-results test-results.xml

onspec expects to run inside a git repository. Specs live in a specs/ directory at the repo root (configurable), one spec per *.spec.md file.

Not an engineer, or drafting specs with a chat assistant? Start with Writing specs with AI: the copy-paste prompt and the browser-only path from draft to merged.

Spec format

A spec is a markdown file with YAML frontmatter. The frontmatter is the machine-checkable contract; everything below it is free-form context for humans and for LLM assessment.

# specs/csv-export.spec.md
id: SPEC-0042
title: CSV export includes archived records
status: approved
refs:
  - PROJ-123
covers:
  - src/export/**
criteria:
  - id: C1
    text: Archived records appear when include_archived=true
    verify: test
    evidence: tests/export.test.ts::includes archived records
  - id: C2
    text: Export format constant stays RFC 4180
    verify: assertion
    evidence: src/export/csv.ts#FORMAT = "RFC4180"
invariants:
  - Export format remains RFC 4180 compliant
non_goals:
  - Bulk archive/unarchive operations

Field reference

FieldRequiredRules
idyesMust match SPEC-NNNN (four digits). Unique across the repo; duplicates are reported and only the first file loads.
titleyesNon-empty string.
statusyesdraft, approved, or superseded. Only approved specs satisfy drift detection; superseded specs are excluded from verification.
coversyesAt least one glob, matched against repo-relative paths (dotfiles included). Defines which changes this spec governs.
criteriayesAt least one. Each has id (C1, C2, ...; unique within the spec), text, verify (test | assertion | manual), and optional evidence.
refsnoExternal trace links: Jira keys, issue URLs, RFC links. Surfaced in reports; URLs render as links in the PR comment.
invariantsnoStatements that must never break. Context for reviewers and LLM assessment.
non_goalsnoWhat this spec deliberately excludes. Keeps generators from over-building and assessors from over-reaching.

A spec file that fails validation is collected as a parse error and reported; it never aborts loading of the healthy specs.

Evidence pointers

Each criterion names its own proof. The pointer format depends on verify:

verify: test

Evidence is path/to/test.file::test name. The test name should match what your runner reports; matching is deliberately lenient (name containment plus file matching when the report provides one), so describe-block prefixes like suite > test name resolve correctly.

Results come from a JUnit XML report passed via --test-results. Any runner that emits JUnit works: vitest (--reporter=junit), jest (jest-junit), pytest (--junitxml), go test (gotestsum), JUnit itself, and most others.

verify: assertion

Evidence is path/to/file#literal snippet. The criterion is met exactly when the file contains the snippet. A deterministic grep-anchor for facts a test doesn't naturally cover: exported constants, config values, flags.

verify: manual

Evidence is free text describing who checks and how. Manual criteria are never auto-passed; every report surfaces them, and lint warns about them.

No evidence

A test or assertion criterion without an evidence pointer falls through to LLM assessment of the diff (when a key is available). It is a visible gap, not an error.

Verdict semantics

Deterministic evidence always wins over LLM opinion. The exact resolution rules:

ConditionVerdict
Named test passed in the JUnit reportmet
Named test failed in the JUnit reportunmet
Named test was skippeduncertain
Evidence test file does not exist, or the test name is not found in itunmet
Test exists but no --test-results report was provideduncertain
Test exists but is absent from the report (e.g. its file crashed)uncertain
Assertion file contains the snippetmet
Assertion file missing, or snippet not foundunmet
verify: manualmanual
No evidence pointer, no API keyuncertain
No evidence pointer, key present: LLM judges the diffmet / unmet / uncertain

The citation rule. An LLM met or unmet verdict must cite at least one file that actually exists in the repo, or it is downgraded to uncertain. Falsifiable claims only; a bare confident verdict is never accepted.

Exit codes. All commands exit 0 by default (advisory). With --strict: verify exits 1 when any criterion is unmet, drift exits 1 on any finding, lint exits 1 on any error-severity finding. Exit 2 means the command itself failed (bad ref, unreadable input).

Commands

onspec verify

Finds the specs whose covers globs match the changed files, then produces a verdict per criterion.

FlagMeaning
-b, --base <ref>Base ref to diff against. Default: config base, then HEAD~1.
--head <ref>Head ref. Default: the working tree, including untracked files.
-t, --test-results <file>JUnit XML report anchoring test-evidence verdicts.
--allVerify every non-superseded spec, not just those governing the diff.
--no-llmSkip LLM assessment; evidence-less criteria stay uncertain.
-m, --model <model>Model for LLM assessment. Default: claude-opus-5.
-f, --format <fmt>terminal (default), markdown (for PR comments), or json (machine-readable, for dispatchers and tooling).
-o, --output <file>Write the report to a file instead of stdout.
--strictExit 1 when any criterion is unmet.

onspec drift

Flags changed code that no approved spec governs. Three finding kinds: unspecced-change (no spec covers the file), stale-approval (only a draft spec covers it), and superseded-coverage (only a superseded spec covers it). Fully deterministic; no LLM involved. Takes --base, --head, --format, --output, --strict.

onspec lint

Grades every spec A to F for verifiability. Errors: evidence pointing at a missing file, parse failures. Warnings: dead covers globs, manual criteria, missing evidence pointers. Info: ambiguous wording (can a machine check "quickly"?), missing non-goals or invariants.

onspec reverse

Reverse-generates draft specs from existing code and tests: the brownfield on-ramp. The LLM only drafts; everything that must be true is enforced deterministically afterwards. Evidence pointers that don't resolve against the real repo are stripped and reported, ids are assigned in sequence after the highest existing spec, and output is always status: draft. Approval is a human act, performed by editing the file in a reviewed PR.

FlagMeaning
--code <globs...>Source globs to spec. Default: config code globs.
--tests <globs...>Test globs to mine for evidence anchors. Default: test/**, tests/**, src/**/*.test.*, src/**/*.spec.*.
--prompt-onlyPrint the drafting prompt and exit; drive any agent you already trust.
--from-json <file>Skip the built-in LLM call; ingest drafts an agent produced. Validation still runs locally.
-m, --model <model>Model for the built-in drafting call. Default: claude-opus-5.

Configuration

Optional onspec.config.json at the repo root. Every field has a default; the file can be omitted entirely.

{
  "specDir": "specs",      // where *.spec.md files live
  "code": ["src/**"],     // what counts as code that must be specced (drift scope)
  "base": "HEAD~1"        // default base ref when --base is not passed
}

The API key is read from the ANTHROPIC_API_KEY environment variable. onspec does not load .env files; source them yourself or export the variable.

CI setup

GitHub Action

Run your tests with JUnit output, then let the Action post a single self-updating conformance comment on the PR:

# .github/workflows/onspec.yml
name: onspec
on: [pull_request]
jobs:
  onspec:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v5
        with: { fetch-depth: 0 }
      - run: npm ci && npx vitest run --reporter=junit --outputFile=test-results.xml
      - uses: Avant-Concepts-LLC/onspec@v1
        with:
          test-results: test-results.xml
          anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}  # optional
          strict: "false"  # advisory first; earn the right to block
InputMeaning
base / headRefs to compare. Default: the PR's base and head SHAs; on push events, falls back to the previous commit.
test-resultsPath to a JUnit XML report produced by an earlier step.
anthropic-api-keyEnables LLM assessment for evidence-less criteria. Omit it and the Action runs fully offline.
strict"true" fails the job on unmet criteria or drift. Default "false".

fetch-depth: 0 matters: the diff needs enough history to reach the base ref.

Closing the loop: dispatch

An approved spec with unmet criteria is a machine-readable work order. The reference dispatch workflow watches merges to specs/, runs verify --all --format json, and opens a work-order issue for each newly approved spec whose criteria aren't yet delivered. Point any coding agent (or human) at the onspec-dispatch label and the loop closes: spec approved, agent implements, onspec verifies the PR, humans merge.

Any other CI

The Action is a thin wrapper. Anywhere else, run the CLI directly:

$ npx onspec verify --base origin/main --test-results results.xml --format markdown --output report.md
$ npx onspec drift --base origin/main

GitLab CI

Include the maintained template; it runs on merge requests and posts a self-updating MR note when a token is provided:

# .gitlab-ci.yml
include:
  - remote: https://raw.githubusercontent.com/Avant-Concepts-LLC/onspec/v1/templates/onspec.gitlab-ci.yml

Configure via CI/CD variables: ONSPEC_TEST_RESULTS (path to a JUnit artifact from an earlier job), ANTHROPIC_API_KEY (optional, masked), ONSPEC_GITLAB_TOKEN (project access token with api scope, enables the MR comment; without it the report lands in the job log and as an artifact), and ONSPEC_STRICT ("true" to block).

Data handling

Short version: nothing leaves your machine unless you provide an API key, and then only the minimum, under your own account.

  • No key: onspec makes no network calls. All verdicts derive from your repo and your test results.
  • Key set: for each criterion lacking deterministic evidence, onspec sends Anthropic the git diff and the governing spec's text, and nothing else, directly from your machine or CI runner under your key and your account's data terms. onspec has no server.
  • Never sent: your full repository, test results, environment variables, or anything for criteria that resolved deterministically. A fully test-anchored spec suite verifies with zero LLM calls.
  • reverse sends the source and test files matched by your globs when invoked with a key, since drafting requires reading the code. Prefer not to? --prompt-only plus --from-json keeps drafting inside an agent you already trust, with validation done locally.