# TaxPilot AI

An autonomous digital employee for TaxPilot CMS. Not a chatbot: it receives
documents, understands them, and carries out work through tools — asking a human
when it is uncertain, and whenever the rules say it must.

**237 tests.** One deployment serves one TaxPilot installation.

## What it does today

A client sends a photograph of a document on WhatsApp. TaxPilot AI reads it,
works out what kind of document it is, extracts the identifiers, and looks for
the client it belongs to. Then it asks.

```
receive → read → classify → extract → identify → propose → (a human decides) → remember
```

**It cannot file anything itself.** The proposal goes to the CMS Approval Queue;
a person reviews it there; the CMS performs the filing. That is ADR-0008, and it
is structural rather than a setting — this process has no path to a client's
record except through a decision somebody made.

## Status

| Area | State |
| --- | --- |
| Request signing, CMS client | Done — verified against PHP byte for byte |
| Tool contract and registry | Done |
| OCR, extraction, classification | Done — Pakistani identity and banking documents |
| Redaction and the data policy | Done |
| Memory (records, embedding, recall) | Done — in-process store; pgvector schema written |
| Workflow engine | Done — resumable, pauses for external decisions |
| Document intake, end to end | Done — proposal submission and decision polling |
| WhatsApp transport | Parsing and session rules done; **live sending is not wired up** |
| Durable run storage | Done — PostgreSQL, with optimistic locking and archival |
| Health, metrics, monitoring | Done — `/health/live`, `/health/ready`, `/metrics` |
| Distribution | Done — signed packages, health-gated install, rollback |
| Version negotiation | Done — refuses a CMS contract it does not understand |

**Not yet possible:** running against a live installation. ADR-0001's deployment
topology is undecided, and the first customer's shared hosting cannot run Python
or PostgreSQL. Everything above is proven against the CMS in tests and on
captured bytes, but no real document has travelled the whole path.

## Architecture

Decision records live in the CMS repository, under `docs/architecture`. The ones
that shape this codebase most:

- **ADR-0001** — one deployment per installation. **There is no tenant concept
  here**, and adding one means multi-tenancy has crept in.
- **ADR-0002** — three data levels. CNICs, bank statements and original
  documents never leave the installation.
- **ADR-0004** — every Tool declares an execution policy. Writes need a human.
- **ADR-0005** — the Agent API is the *only* way to reach the CMS. No database
  driver, no browser automation, no file access.
- **ADR-0008** — the CMS owns proposals and performs approved writes. This
  platform asks; it does not file.

## Layout

```
app/
  api/        the signed CMS client — the only channel out
  security/   request signing and redaction
  tools/      capabilities; the unit everything is built from
  ocr/        local reading, classification, field extraction
  memory/     what earlier runs learned
  workflow/   the engine, the intake definition, decision polling
  whatsapp/   transport, provider abstraction, session rules
  release/    building, verifying, installing and rolling back a release
  runtime/    the daemon, health, metrics, the HTTP server
  config/     one deployment's settings
scripts/      fixtures for the CMS's cross-language tests
tests/
```

## Running

```bash
python -m pip install -e ".[dev]"
python -m pytest
```

## Configuration

```bash
TAXPILOT_CMS_URL=https://my.example.com   # the installation this AI serves
TAXPILOT_AGENT_KEY=tpa_...                # from `php artisan agent:issue`
TAXPILOT_AGENT_SECRET=...                 # shown once, at issue
```

Missing any of these and the process refuses to start, rather than running and
failing on every call — the second is far harder to diagnose from a log of 400s.

The agent needs `proposals.submit` and `clients.read` in the CMS. It does **not**
need `documents.write`: that is ADR-0004's promoted path, granted per Tool once
measured accuracy has earned it, and held by no agent by default.

## The signature scheme

```
HMAC-SHA256( timestamp \n nonce \n METHOD \n path \n rawBody , secret )
```

Method, path and body are all covered, so none can be altered in flight. The
timestamp bounds the replay window; the nonce is single-use inside it.

`path` is the routed path **without** the query string, because that is what
PHP's `$request->path()` returns. `rawBody` is the bytes actually sent — the CMS
verifies against those and never re-encodes.

Two implementations of one sentence in two languages is exactly how integrations
drift apart silently, so this is pinned twice:

- `tests/test_signing.py` checks vectors produced by PHP's `hash_hmac`.
- `scripts/emit_upload.py` and `scripts/emit_proposal.py` capture real signed
  requests, which the CMS replays byte for byte. Those fixtures deliberately
  carry non-UTF-8 file bytes and non-ASCII text: Python escapes non-ASCII as
  `\uXXXX` and PHP does not, so a CMS that re-encoded before verifying would
  pass every same-language test and fail these.

## Tools

Every capability is a Tool. The AI chooses which to run; a Tool contains no AI
logic and never calls another — that would build a hidden call graph the workflow
engine cannot see, resume or audit.

Each declares an execution policy, and **the default is `DISABLED`**: a Tool
whose author forgot to declare one cannot run.

| Policy | Meaning |
| --- | --- |
| `AUTOMATIC` | Reads, or work that writes nothing. Being wrong wastes a call |
| `REQUIRES_APPROVAL` | Returns a proposal instead of acting |
| `DISABLED` | Not executable. Deletion and permission changes live here permanently |

`submit_filing_proposal` is `AUTOMATIC`, which looks wrong for something that
ends in a client's record changing and is not: submitting writes nothing. The
gate is a person, not the policy.

## The workflow engine

Steps are persisted as they complete, so an interrupted run resumes rather than
starting over — and never repeats a step that already did its work. For a
workflow that files documents, doing something a second time is the expensive
mistake.

A step can end three ways, not two. `ToolResult.pending()` says the work was
handed to somebody else and has not come back: reporting that as failure would
retry a submission that already landed, and reporting it as success would carry
on as though a document had been filed.

`DecisionPoller` brings answers back. When a reviewer sends work back it reads
their directive and rewinds to the step named — "read the document again" reruns
the OCR and everything downstream, because a different reading changes the
classification, the fields and possibly the client. Resubmission uses a fresh
idempotency key: the CMS returns the same proposal for a repeated one, which
would hand the reviewer back the answer they just rejected.

## What memory keeps

A summary, never a transcript. Memory is embedded and may be sent to a hosted
model to reason over, so raw OCR text would leave the installation the first time
anyone asked a question about that client. Identifiers are masked before storage,
and restricted keys are refused outright.

## Known gaps

- **Live WhatsApp sending.** Evolution records rather than transmits without a
  base URL; `MetaBusinessProvider` awaits business verification and template
  approval. Both are stated gaps rather than unverified HTTP calls.
- **Schema rollback.** `python -m app.release rollback` restores code and the
  pointer, never the database — a fast rollback cannot also restore a dump
  without discarding every workflow run since the install. Migrations must
  therefore be additive; see [docs/distribution.md](docs/distribution.md).
- **Pull-based updates.** Releases are pushed by an operator, which is right for
  the vendor-hosted model. A customer hosting their own deployment would need a
  pull transport — the verification and install machinery already exists, so it
  is a transport, not a redesign.
