# Running TaxPilot AI

```bash
python -m app             # the daemon
python -m app check       # verify this deployment and exit
python -m app migrate     # apply pending schema changes and exit
```

## Exit codes

| | |
| --- | --- |
| `0` | Fine |
| `1` | A check failed, or something broke |
| `78` | Misconfigured — `EX_CONFIG` |

78 rather than 1 for configuration, so a deployment script can tell "you set this
up wrong" from "it broke" without parsing logs.

## `check` before switching traffic

It makes exactly the same assertions the daemon makes at boot, so its answer
means something. Three checks, and the split between them is the point:

| Check | Missing means | |
| --- | --- | --- |
| **CMS** | Cannot reach it, or the agent lacks `clients.read` / `proposals.submit` | **Fatal** |
| **Database** | No `TAXPILOT_DATABASE_URL`, or PostgreSQL unreachable | **Fatal** |
| **OCR** | No engine installed | **Degraded** |

Missing OCR does not stop the process. Every document still reaches a reviewer —
unread, and marked high risk because nothing could be read. Refusing to start
would turn a degraded service into no service, and the documents would arrive
anyway.

Missing `proposals.submit` **is** fatal: without it every document the AI reads
has nowhere to go, so the process would run, read, and silently discard.

All checks run even after one fails. Reporting the first and stopping means an
operator fixes it, restarts, discovers the second — several round trips on
somebody else's infrastructure that one log could have saved.

## Migrations run before the daemon starts

`python -m app` applies pending migrations first. A process that starts against a
schema it does not match fails on the first write — halfway through handling a
real document, rather than at boot where it belongs.

## Shutting down

SIGTERM (`docker stop`) and SIGINT (Ctrl-C) both stop it cleanly: the current job
finishes, nothing new starts, the process exits.

The loop waits on an Event rather than calling `time.sleep`. That is not a
stylistic preference — `time.sleep(30)` leaves a SIGTERM unanswered for up to
thirty seconds, past Docker's ten-second grace period, at which point it sends
SIGKILL and whatever was running dies mid-write. Measured: **0.40s to stop inside
a 5-second wait window.**

A `stopping` flag is checked between jobs as well as between passes, so a stop
arriving during a slow pass is not held up by everything after it.

## Jobs

| Job | Interval | |
| --- | --- | --- |
| `decisions` | `TAXPILOT_POLL_SECONDS` (30s) | Ask the CMS what reviewers decided; resume the affected runs |
| `archive` | daily | Move finished runs out of the working set |

`decisions` holds **every** workflow that can pause for a reviewer — document
intake and bank statement intake. A decision comes back naming a proposal and
nothing else, so which definition to resume against is a property of the paused
run; a poller holding one workflow would step a forwarding run through the
reading workflow's sequence. A new workflow that submits proposals must be added
to `Container.poller` at the moment it is defined.

A directive naming a step the workflow does not have — "read the document again"
sent to a workflow that reads nothing — rewinds to its first step rather than
failing the run. A reviewer asking a reasonable question must not cost a
document.

A failing job never stops the loop. The CMS restarting, a network blip, one
malformed document — each is an ordinary Tuesday, and a supervisor that exits on
the first turns thirty seconds of trouble into an outage lasting until somebody
notices. Failures are counted per job with the last error kept.

Intervals are monotonic, so an NTP correction or a daylight-saving change cannot
make a job run twice or stall for an hour.

## Health endpoints

| | | On failure the orchestrator should |
| --- | --- | --- |
| `GET /health/live` | Is the process wedged? | **Restart** it |
| `GET /health/ready` | Can it do useful work? | **Withhold traffic**, do not restart |

`200` when healthy, `503` when not. Loopback by default — inside a container that
suffices for Docker's own probe, and an accidental port publish then cannot put an
infrastructure status page on the internet.

### Liveness deliberately ignores everything external

This is the whole design, and getting it wrong is expensive. If liveness checked
the CMS, then a CMS taken down for ten minutes of maintenance would have the
orchestrator restart the AI over and over — discarding in-flight work, achieving
nothing, and slowing recovery once the CMS returned.

Liveness therefore considers exactly one thing: **is the loop still coming
round?** A job hung on a socket with no timeout leaves a process that exists,
answers signals, and does nothing. That is the one condition a restart fixes.

It tolerates several missed passes before declaring death, because a restart
throws away in-flight work and one slow iteration is not a reason to kill
anything. A daemon that has been *asked* to stop still reports alive — it is
exiting on purpose, and reporting otherwise would have an orchestrator restart
something that was told to go away.

### Degraded is ready

Missing OCR, unresponsive memory and an unconfigured WhatsApp inbox are all
**degraded**, not failing. Documents still reach a reviewer; workflows already in
flight still finish. Failing readiness over an optional component would take a
working service out of rotation.

Only these fail readiness: no database, an unreachable database, an unreachable
CMS, or a CMS that has not granted `clients.read` and `proposals.submit` — the
last because a reachable CMS that refuses proposals means every document read has
nowhere to go.

### The response says little on purpose

Component names and coarse states only. Never a hostname, a DSN, a driver message
or an allowed number — this is the endpoint most likely to be exposed by
accident, and `connection to 10.0.0.4:5432 failed` is a map of the infrastructure
for whoever reads it. Watched numbers appear masked (`…567`).

Every component is reported even after one fails, because an operator wants the
whole picture rather than the first thing that broke.

### Docker

```dockerfile
HEALTHCHECK --interval=30s --timeout=5s --start-period=60s --retries=3 \
  CMD python -c "import urllib.request,sys; \
      sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8080/health/live', timeout=4).status==200 else 1)"
```

Python rather than `curl`, which a slim image may not carry. A generous
`start-period` matters: PaddleOCR loads its models on first use, and a probe that
starts too early kills the container while it is still coming up.

Docker has only one health concept, so it gets **liveness**. Readiness belongs in
a proxy or orchestrator that can withhold traffic without killing anything —
pointing Docker's check at `/health/ready` would restart the container every time
the CMS blinked.

## Composition

`app/runtime/container.py` is the only file that names concrete implementations
together. Everything else takes its collaborators as arguments, which is what
makes the rest of the platform testable without a CMS, a database or an OCR
engine.

Construction is lazy, one component at a time, so `check` and `migrate` do not
pay for models they never load.

**`UploadDocumentTool` is deliberately not registered.** It files without a human
— ADR-0004's promoted path — and a deployment gets it only by deciding to.

**No OCR tool reaches the bank statement workflow**, and that is enforced by the
workflow's steps rather than by the registry, which holds one for the reading
path. Its own tests register no OCR tool at all, so a step that tried to read a
statement fails with "no tool named ocr_document" rather than quietly working.

**The classification model is off unless configured**, and refused unless it
runs on this host. `TAXPILOT_MODEL_URL` plus `TAXPILOT_MODEL_NAME` turn it on;
anything that is not a loopback or private address disables the stage and says
so at boot rather than stopping the deployment. A document's text is Level 3 and
does not leave the installation — see ADR-0009 in the CMS repository for the
rule and for what would have to change to revisit it.

    TAXPILOT_MODEL_URL=http://127.0.0.1:11434    # Ollama, llama.cpp, LM Studio
    TAXPILOT_MODEL_NAME=llama3.2
    TAXPILOT_MODEL_TIMEOUT=20

The endpoint is spoken to as an OpenAI-compatible chat API, which all three of
those serve, so the runtime is a deployment choice rather than a code one.

## Known gaps

- **The WhatsApp listener is not here yet.** The HTTP server it needs now exists
  and takes registered routes, so the webhook goes into that one rather than
  needing a second on another port. Phase 3.
- **Signal *delivery* is unverified on Windows**, which cannot deliver SIGTERM
  the way Linux does. The handler and the interruptible wait are both tested;
  the OS half needs a Linux host, which is where this deploys.
- **Memory is still in-process.** Recall works and does not survive a restart.
