# Monitoring

`GET /metrics` — Prometheus text format, on the same loopback server as the
health probes.

## Two questions, two places

This answers **"how is the process doing"**: latency, failures, retries,
durations. Operations.

It deliberately does not answer **"how is the AI doing at its job"**. Approval
time, correction rate, ambiguity and rejection reasons live in the **CMS**,
computed from proposals — because that is the business question, and the CMS is
where the people asking it already are. An accountant should never be shown API
latency percentiles, and duplicating those figures here would give two systems
the same number and two chances to disagree about it.

## What is exposed

| Series | Type | |
| --- | --- | --- |
| `taxpilot_cms_requests_total{endpoint,outcome}` | counter | ok / refused / unreachable |
| `taxpilot_cms_request_seconds{endpoint}` | histogram | API latency |
| `taxpilot_ocr_documents_total{outcome}` | counter | read / unreadable |
| `taxpilot_ocr_confidence` | histogram | distribution, not an average |
| `taxpilot_ocr_seconds` | histogram | |
| `taxpilot_whatsapp_sent_total{outcome}` | counter | ok / failed |
| `taxpilot_memory_recalls_total{outcome}` | counter | hit / empty |
| `taxpilot_memory_recall_seconds` | histogram | |
| `taxpilot_jobs_total{job,outcome}` | counter | per scheduled job |
| `taxpilot_job_seconds{job}` | histogram | |
| `taxpilot_workflows{state}` | gauge | last 7 days, from the database |
| `taxpilot_workflows_retried` | gauge | runs where a step was attempted twice |
| `taxpilot_workflow_median_seconds` | gauge | typical time start to finish |

### Confidence is a distribution, not an average

A mean hides the shape that matters. Clean scans at 0.98 and phone photographs
at 0.35 average to "mediocre", which describes neither population and suggests
tuning a threshold that is already correct for both. The histogram shows two
humps; the mean shows one lie.

### Counters reset on restart; history comes from the database

Which is correct for this format — a scraper samples every few seconds and
computes rates itself, and `rate()` handles a counter resetting.

The `taxpilot_workflows*` gauges are different: they are **derived** from
`workflow_runs` and `workflow_steps` at scrape time. Those tables already record
what happened and when, so a parallel metrics table would be a second copy of
the same facts, kept in step by hand, and eventually disagreeing with the first —
at which point nobody knows which to believe.

### Absent, not zero

`taxpilot_workflow_median_seconds` is omitted entirely when nothing has finished.
Emitting `0` would report instant workflows rather than no workflows, and a
dashboard renders those identically.

Success rate follows the same rule: computed over runs that **finished**, not
runs that started. Counting proposals still awaiting a reviewer as failures would
blame the AI for a queue nobody has worked through.

## Two things that keep the monitoring from becoming the problem

**Bounded label values.** CMS calls are labelled by *endpoint*, not by path:
`clients/42` and `clients/77` are the same operation, and one series per client
id is unbounded cardinality — the classic way to bring down a monitoring system
with the telemetry meant to protect it.

**A scrape cannot take the endpoint down.** If the database read for the derived
gauges fails, those series are dropped and the rest still serves. Losing one
series beats losing the monitoring of everything else — including the alert that
would have said the database was down.

Label values are escaped, because they come from things like error reasons, and
one stray quote produces output no scraper can read — which loses every series,
not just that one.

## Scraping it

```yaml
scrape_configs:
  - job_name: taxpilot-ai
    static_configs:
      - targets: ["127.0.0.1:8080"]
```

Loopback by default, like the health probes. This is operational detail about a
system handling client documents; it is not a public page, and an accidental
port publish should not make it one.

## Worth alerting on

| | Why |
| --- | --- |
| `taxpilot_ocr_documents_total{outcome="unreadable"}` rising | Documents arriving that nothing can read — a camera, a format, or an OCR failure |
| `taxpilot_whatsapp_sent_total{outcome="failed"}` above zero | A message that fails to send is invisible to whoever was waiting for it |
| `taxpilot_cms_requests_total{outcome="unreachable"}` sustained | The CMS is down; readiness will already be failing |
| `taxpilot_jobs_total{outcome="failed"}` sustained | Something is wrong that the loop is politely surviving |
| `taxpilot_workflows{state="awaiting"}` climbing | Proposals are arriving faster than anybody reviews them |

The last one is the most useful and the least technical: it means the queue is
growing, which is a staffing question rather than a software fault.
