AI · AUTOMATION · CLOUD · Cork AI Consulting · professional services firm
IntakeLens IllustrativeDocument Intake & Triage Service
A professional services firm received client documents through four channels and keyed the same two dozen fields out of them by hand. I built an intake service that classifies each document, extracts its fields, and routes it — but only when it is confident. Below the floor it goes to a person, because in this business a misfiled document is a compliance problem, not an inconvenience.
This is an anonymised, illustrative scenario — representative of the shape of engagement and the way I work, not a named client or a delivered set of results. The real, named work on this site is labelled as such.
- −95% median intake turnaround, 60 hours to 3
- 94.5% field-extraction accuracy above the confidence floor
- 100% of low-confidence documents routed to a human
Value returned
€21k
Annualised administrative cost returned
12 hours a week at a conservative €34/hour loaded cost, redeployed to client-facing work.
Time returned
12h
Administrative hours returned each week
Time previously spent opening, classifying and re-keying inbound documents.
Effort invested
7 weeks
Effort to build
Cork AI Consulting — channel mapping, extraction model, routing service and staff handover.
Quality
94.5%
Field-extraction accuracy
On documents above the confidence floor. Everything below it reaches a person on first touch.
01 — Context
The problem.
Documents arrived by email, through a web form, by post as scans, and occasionally by hand. Whoever opened one classified it, keyed roughly two dozen fields into the practice management system, and moved it to a queue. Nothing about that work needed a qualified person, but all of it needed a careful one.
The cost was turnaround as much as labour. A document that arrived on Friday afternoon might not be in a queue until Tuesday, and the firm had no way to see what was sitting unprocessed — the backlog only existed in individual inboxes.
02 — Method
The approach.
The build, in the order it happened.
-
Made the channels one channel
All four intake routes now land in a single queue with the source recorded. That alone made the backlog visible for the first time, and it is the change staff mentioned first.
-
Classify, then extract
A document type is settled before any field extraction is attempted, because the fields worth extracting depend on it. Classification is cheap and auditable; extraction is the expensive part and only runs against a known schema.
-
A confidence floor with a human path
Every extracted field carries a confidence. Below the floor, the document routes to a person with the uncertain fields highlighted rather than being filed on a guess. The floor was set from the firm’s own tolerance for a misfiling, not from a default.
-
Auditable by design
Every decision, its confidence, and any human correction is logged against the document. Corrections feed the weekly review, and the log is what let the firm sign off on the service at all.
03 — Outcome
What changed.
- Median intake turnaround fell from around 60 hours to 3, and a Friday document is now in the right queue on Friday.
- Field-extraction accuracy of 94.5% above the confidence floor, with every below-floor document reaching a person on first touch.
- Hand-keyed fields per document fell from 26 to 2 — the two that genuinely need a judgement call.
- The backlog became a number on a screen instead of an unknown quantity spread across four inboxes.
04 — Before & after
The same measures, either side of the work.
Each pair is scaled against its own larger value, so the comparison is honest rather than flattering.
Median intake turnaround
−95%- Before
- 60 hours
- After
- 3 hours
Administrative hours per week on intake
−80%- Before
- 15 hours
- After
- 3 hours
Fields keyed by hand per document
−92.3%- Before
- 26 fields
- After
- 2 fields
05 — The numbers
Median intake turnaround.
Tracked as hours, from Baseline through Wk 7 — low 3, high 60.
06 — Screens
What it looks like in use.
The review screen. Fields below the confidence floor are highlighted, so a person reads the two that are uncertain rather than all twenty-six.
One queue, four channels, with age on the face of it. Making the backlog visible changed behaviour before the model did.
07 — Stack & role
Built with.
- Python (spaCy, pdfplumber, scikit-learn)
- Claude API
- FastAPI
- PostgreSQL
- Cloud Run
- Role
- Cork AI Consulting — founder & lead consultant, sole delivery
- Duration
- 7 weeks
- Client
- Cork AI Consulting · professional services firm
- Period
- 2026
08 — Questions
The questions I get asked about this one.
Reviewing everything reproduces the original cost. The floor concentrates human attention on the documents where it changes the outcome, and it is set from the firm’s own tolerance for a misfiling rather than from a model default.
Documents stay inside the firm’s own cloud tenancy, access is logged per document, and retention follows their existing policy. That constraint shaped the architecture from the first week rather than being retrofitted.
It is measured on a held-out sample, not on the training data, and every human correction in the live log is a fresh test case. Accuracy is re-reported weekly rather than claimed once at launch.