Contributing
pnpm install && pnpm bootstrap && pnpm devRuns offline by default — a mock model and an embedded Postgres, so no API key or database is needed to work on it.
Changes that need no spec and no ADR
Open a pull request directly for any of these:
- Documentation fixes: a wrong command, a stale path, an unclear sentence.
- New or stronger tests, including ones that pin behaviour a spec already describes.
- New eval cases in
evals/golden/, written against the fixture tenant. - Bug fixes that restore specified behaviour: the spec says one thing, the code does another, and the fix makes the code agree.
A spec is required when a change adds behaviour or changes what existing behaviour is: write it in specs/ first. An ADR is required when a change picks between defensible options that a later reader would question: record it in docs/adr/. Either way, the pull request template and CI apply to every pull request; this section changes what you write first, not which rules apply.
To report a bug instead of fixing it, the issue form asks for a reproduction with pnpm simulate against the fixture tenant (CONFIG_DIR=test/fixtures/config) and the mock model, never a transcript from a real deployment.
Before opening a PR
pnpm typecheck && pnpm lint && pnpm format:check && pnpm test && pnpm eval:mockCI runs exactly these, plus a build, secret scanning and CodeQL.
The PR template opens with ## Why. Write that part first: the problem or the decision, not a summary of the diff. See specs/006-pull-requests.md.
How this project is organized
Behavior is specified in specs/ before it is implemented, and decisions are recorded in docs/adr/.
specs/README.md is the index: which specs describe code that exists, and which describe code that does not. It is generated - run pnpm spec:index after changing any spec's frontmatter, or CI will fail.
specs/000-constitution.mdholds non-negotiables. A change there needs an ADR that supersedes the clause.specs/008-spec-metadata.mddefines that frontmatter.status: implementedis only permitted where a test cites the spec, so the index cannot quietly claim more than the suite proves.specs/004-testing.mddefines what must be tested, what deliberately is not, and the rules a test here follows.specs/005-language.mdrequires English throughout, and that no customer-facing copy lives in source at all - it belongs to the tenant, in configuration.specs/009-tenant-eval-suites.mdexplains whyevals/golden/holds framework cases against the demo tenant only, and how a tenant points the runner at a suite of its own.specs/016-model-graded-evals.mddefines when a judge model may grade an eval reply: only after agreeing with a hand-labelled calibration set in the same run, and counted apart from asserted results.specs/006-pull-requests.mddefines what a pull request must say. The first heading is always## Why- the diff already says what changed.specs/010-release-workflow.mdexplains why a release is cut by merging a Release PR rather than by pushing tomain: the changelog is generated from commit subjects, so it gets read as a diff before it is published.specs/011-dependency-updates.mdexplains how dependency PRs are batched, and which ones merge without review. Patches auto-merge on green CI; minors and majors do not.specs/014-docs-site.mdexplains how these documents are published as a site: rendered where they live, from an allowlist, never copied into a separate docs tree.specs/015-test-service.mddefines the browser test page: off unlessDEMO_UI=true, mock model only, and it drives the real ManyChat route rather than a shortcut around it.specs/021-contributor-surface.mddefines how issues and questions are taken in: forms only, no blank issues, and the same no-tenant-data acknowledgement a pull request carries.- Tests cite the spec clause they enforce. If a spec change breaks a test, that is a real finding — not a test to update mechanically.
- New non-obvious decisions get an ADR. "Why not Redis" is more useful to the next reader than the code that avoided it.
Repository skills
.claude/skills/ holds instructions for the workflows this repo has conventions about, so they do not depend on remembering them.
| Skill | Use it to |
|---|---|
adr | Pressure-test a decision, then record it in docs/adr/. |
spec | Settle what a spec must say, then write it in specs/. |
new-skill | Add another one of these. |
eval-review | Grade pnpm eval replies against review criteria and the 016 rubric, quoting the reply. |
implement-spec | Build a specified spec in its own worktree, tests citing each Verification item. |
open-pr | Open a PR whose body follows the template and 006, after a leak check. |
They are committed on purpose - a skill encoding this repo's format is repo tooling. Anything that would work unchanged in an unrelated project is a personal skill and does not belong here.
adr and spec both open by asking questions, and will stop rather than write if the answers are not there - an ADR whose rejected alternative nobody can argue, or a spec claiming behaviour no test covers, is worse than the missing document, because both read as authoritative.
Things worth knowing
Zod schemas are the source of truth. Types are inferred and the OpenAPI document is generated from them. Do not hand-write a parallel interface.
Only src/agent/registry.ts may import a provider package. Everything else depends on the AgentRunner port. This is what keeps the model swappable.
Channel differences are data. Add a row to CHANNEL_CAPABILITIES, do not add a branch to the renderer.
Timeouts have an ordering invariant. MODEL_ABORT_MS must exceed RACE_DEADLINE_MS; losing the race must not cancel the model call, or the deferred path can never run. EnvSchema enforces this and a regression test covers it.
Local models run via Ollama. docker compose --profile local-model up starts the container; set AGENT_MODEL=ollama:llama3.1:8b. Ollama models price at zero so the budget cap does not fire on free turns. The token cap still applies.
Mock at the provider boundary carefully. doGenerate returns the nested provider-facing usage shape ({ total, noCache, cacheRead, cacheWrite }), which the SDK flattens for callers. Using the flattened shape in a mock silently yields undefined token counts and makes the budget cap a no-op. Use the helpers in test/helpers/model.ts.
The README's images are made from the fixture tenant. pnpm demo:record regenerates docs/assets/demo.svg by running the simulator against the fixture tenant and the mock model; it cannot be pointed at anything else. docs/assets/social-preview.svg is the source of the repository's social preview, uploaded by hand as a 1280x640 PNG. No test can read an image, so a reviewer checks both for real tenant copy (specs/021).
Adding a channel
Implement ChannelAdapter in src/channels/<name>/, add a capability row, and write the renderer tests first — the capability profile is the specification.