# Reading documents

Local, always. ADR-0002 puts original documents at Level 3, which rules out
Google Vision, Azure and Textract regardless of their accuracy — each requires
sending the image off-box, and a photograph of a CNIC is exactly what must not
go.

PaddleOCR over Tesseract because the input is phone photographs sent over
WhatsApp: rotated, shadowed, unevenly lit. That is precisely where Tesseract's
accuracy falls away. The cost is a ~2 GB image and a 4 GB memory floor per
deployment.

## ⚠️ oneDNN must stay disabled on PaddlePaddle 3.3.1

With `enable_mkldnn=True`, inference raises:

```
NotImplementedError: (Unimplemented) ConvertPirAttribute2RuntimeAttribute
not support [pir::ArrayAttribute<pir::DoubleAttribute>]
```

and **no text is read at all**. This is a fault in the released version, not a
configuration mistake — verified on this platform, with the same image reading
cleanly the moment oneDNN is off.

`OcrConfig.enable_mkldnn` therefore defaults to `False`. Working beats fast. A
deployment whose Paddle build handles it can set `TAXPILOT_OCR_MKLDNN=1` and take
the speed-up.

**This must be carried into the container image.** It is the single most likely
cause of "the AI reads nothing" after a deploy, and the symptom — every document
arriving unreadable and high-risk — looks exactly like bad scans.

## Configuration

| Variable | Default | |
| --- | --- | --- |
| `TAXPILOT_OCR_LANGUAGE` | `ur` | The **Arabic-script recogniser, which reads Urdu and Latin** — not an Urdu-only setting. See below |
| `TAXPILOT_OCR_FALLBACK_LANGUAGE` | `en` | Read again in this alphabet when the first pass identifies nobody. Empty to disable |
| `TAXPILOT_OCR_DETECTION_MODEL` | `PP-OCRv5_mobile_det` | Pinned: choosing `ur` otherwise selects a *server* detector, at 103.7s a document against 10.4s |
| `TAXPILOT_OCR_ANGLE_CLASSIFIER` | on | Rotated text lines |
| `TAXPILOT_OCR_PAGE_ORIENTATION` | on | A page photographed sideways reads as nothing without it |
| `TAXPILOT_OCR_UNWARP` | off | Flattens curved pages; another model, and aimed at bound books |
| `TAXPILOT_OCR_MKLDNN` | off | See above |
| `TAXPILOT_OCR_MIN_CONFIDENCE` | `0.30` | Below this the engine emits noise from edges and shadows |
| `TAXPILOT_OCR_MODEL_DIR` | unset | Bake models into the image; see below |
| `TAXPILOT_OCR_THREADS` | `2` | Otherwise Paddle takes every core and starves the poller |

### Why the Arabic-script model, on documents that look English

`ur` selects PaddleOCR's Arabic-script recogniser. It reads **Urdu and Latin**,
and it is the default because it reads these documents better — including their
English.

This was `en`, on the reasoning that Pakistani tax documents are overwhelmingly
English and every extracted field (CNIC, IBAN, mobile, amounts) is Latin digits.
The second half is true. The first was wrong about the documents that matter
most.

Measured on a real CNIC, both sides:

| | English model | Arabic-script model |
| --- | --- | --- |
| Front | 248 chars | **316 chars** |
| Back | 91 chars | **120 chars** |

And the difference is not in the Urdu. The English model rendered the card's
English headline as `A at ar` / `LAN`. The Arabic-script model read
`PAKISTAN National Identity Card` and `ISLAMIC REPUBLIC OF PAKISTAN` correctly.
**Surrounding Urdu was making the English model misread the English** — and
"national identity card" is the strongest marker the classifier has.

The risk in switching was English-only documents, so that was measured too, on a
rendered bank statement: **identical output** — 271 characters from both, 0.997
against 0.996 confidence, all six expected markers found by each, both
classified `bank_statement`.

So this **replaces** the English model rather than joining it. One engine, no
second pass, no extra memory. (608 MB peak was with both resident during the
comparison; only one is loaded in service.)

A deployment whose documents really are English-only can set
`TAXPILOT_OCR_LANGUAGE=en` and have the dedicated model back.

### The second pass

No single model reads every Pakistani identity document. Measured on four real
cards a customer photographed:

| | Arabic-script | English |
| --- | --- | --- |
| old all-Urdu card | 95 chars, **no CNIC** | 70 chars, CNIC |
| old all-Urdu card | 153 chars, CNIC | 55 chars, CNIC |
| new card, back | 91 chars, CNIC | 86 chars, CNIC |
| new card, front | 317 chars, CNIC | 298 chars, CNIC |

The Arabic-script model stays the primary — it reads far more, and it reads the
English that the English model garbles on a bilingual card. But on the first
card it missed the identity number the English model found, and that number is
the difference between filing against a named client and handing a reviewer a
stranger's card.

So a document that the first read **cannot identify anybody from** is read
again with `TAXPILOT_OCR_FALLBACK_LANGUAGE`, and the two reads are **combined**.
Combining was never worse than the better single read on any of the four, and
it changed no classification — no type gained by luck, none lost to noise.

The trigger is deliberately *not* "did we classify it". A type nobody
recognised is a real outcome that a second alphabet rarely fixes; the second
card above is unknown in both and carries its CNIC in both, so reading it again
would spend ten seconds to learn nothing.

Live, after wiring:

```
doc1 old all-urdu   39.1s  paddleocr:ur+paddleocr:en  166 chars  cnic FOUND
doc2 old all-urdu   10.0s  paddleocr:ur               153 chars  cnic found
doc3 new back        7.1s  paddleocr:ur                91 chars  cnic found
doc4 new front      15.8s  paddleocr:ur               317 chars  cnic found
```

One document paid for the second pass. Three did not.

Set `TAXPILOT_OCR_FALLBACK_LANGUAGE=` (empty) to switch it off.

### Models belong in the image

Left unset, PaddleOCR downloads several hundred megabytes on first use. A service
that fetches its models at start-up fails the day the mirror is slow — which is
after a deploy, when nobody is watching. Bake them in and point
`TAXPILOT_OCR_MODEL_DIR` at them.

## Reading order is not detection order

The engine returns regions roughly top-to-bottom, which is not the same as
reading order. That matters more than it sounds, because everything downstream —
classification and every extraction pattern — runs over `OcrResult.text`, the
blocks joined by newlines.

Get it wrong and a CNIC printed as `Identity Number:` on one line and
`35202-1234567-1` on the next arrives with a paragraph between them. The pattern
that would have matched does not, and the document is "unreadable" for a reason
that has nothing to do with the scan.

`app/ocr/layout.py` groups regions into lines and orders them. Two decisions in
there are worth knowing:

- **Line tolerance is a fraction of typical character height, not a pixel
  count.** A phone photograph and a flatbed scan of the same document differ by
  an order of magnitude in resolution; a pixel threshold works for exactly one.
- **Typical height is a median, not a mean.** One full-width heading would drag
  a mean upward and merge genuinely separate body lines into one.

## Failure is an answer, never an exception

`read()` never raises. A missing file, a missing library, a corrupt image and a
page of noise all return an empty `OcrResult` — which every caller already
handles by sending the document to a human. An exception here would end a
workflow that should have reached the Approval Queue.

A deployment without PaddleOCR installed still starts, still reaches its CMS, and
puts every document in front of a reviewer unread and marked high risk. That is a
bad day, not an outage — and it is logged loudly, because the two engines are
indistinguishable from every caller's point of view and differ entirely in
whether the product works.

## Testing

```bash
python -m pytest -m "not slow"   # 2.5s, while iterating
python -m pytest                 # 40s, includes the real engine — CI runs this
```

`TestAgainstTheRealEngine` loads PaddleOCR and reads a rendered document whose
contents are known exactly. Everything else in that file uses result shapes taken
from the documentation, which is fine for testing translation logic and proves
nothing about the library.

Those three tests earned their runtime immediately: they are what caught the
oneDNN failure above. No amount of fake-based testing would have found it, and it
would otherwise have surfaced on the VPS, after deployment, as "the AI cannot
read anything".
