Portfolio Document Intake & Triage Service

AI · AUTOMATION · CLOUD · Cork AI Consulting · professional services firm

IntakeLens Illustrative

Document Intake & Triage Service

2026 7 weeks 6 min read

A professional services firm received client documents through four channels and keyed the same two dozen fields out of them by hand. I built an intake service that classifies each document, extracts its fields, and routes it — but only when it is confident. Below the floor it goes to a person, because in this business a misfiled document is a compliance problem, not an inconvenience.

This is an anonymised, illustrative scenario — representative of the shape of engagement and the way I work, not a named client or a delivered set of results. The real, named work on this site is labelled as such.

  • −95% median intake turnaround, 60 hours to 3
  • 94.5% field-extraction accuracy above the confidence floor
  • 100% of low-confidence documents routed to a human

Value returned

€21k

Annualised administrative cost returned

12 hours a week at a conservative €34/hour loaded cost, redeployed to client-facing work.

Time returned

12h

Administrative hours returned each week

Time previously spent opening, classifying and re-keying inbound documents.

Effort invested

7 weeks

Effort to build

Cork AI Consulting — channel mapping, extraction model, routing service and staff handover.

Quality

94.5%

Field-extraction accuracy

On documents above the confidence floor. Everything below it reaches a person on first touch.

Screenshot of the Document Intake & Triage Service: three headline metric cards above a bar chart of median intake turnaround, 60 hours at Baseline down to 3 by Wk 7.
Inbound client documents read, classified, key fields extracted and routed to the right queue — with a confidence floor and a human path.

01 — Context

The problem.

Documents arrived by email, through a web form, by post as scans, and occasionally by hand. Whoever opened one classified it, keyed roughly two dozen fields into the practice management system, and moved it to a queue. Nothing about that work needed a qualified person, but all of it needed a careful one.

The cost was turnaround as much as labour. A document that arrived on Friday afternoon might not be in a queue until Tuesday, and the firm had no way to see what was sitting unprocessed — the backlog only existed in individual inboxes.

02 — Method

The approach.

The build, in the order it happened.

  1. Made the channels one channel

    All four intake routes now land in a single queue with the source recorded. That alone made the backlog visible for the first time, and it is the change staff mentioned first.

  2. Classify, then extract

    A document type is settled before any field extraction is attempted, because the fields worth extracting depend on it. Classification is cheap and auditable; extraction is the expensive part and only runs against a known schema.

  3. A confidence floor with a human path

    Every extracted field carries a confidence. Below the floor, the document routes to a person with the uncertain fields highlighted rather than being filed on a guess. The floor was set from the firm’s own tolerance for a misfiling, not from a default.

  4. Auditable by design

    Every decision, its confidence, and any human correction is logged against the document. Corrections feed the weekly review, and the log is what let the firm sign off on the service at all.

03 — Outcome

What changed.

  • Median intake turnaround fell from around 60 hours to 3, and a Friday document is now in the right queue on Friday.
  • Field-extraction accuracy of 94.5% above the confidence floor, with every below-floor document reaching a person on first touch.
  • Hand-keyed fields per document fell from 26 to 2 — the two that genuinely need a judgement call.
  • The backlog became a number on a screen instead of an unknown quantity spread across four inboxes.

04 — Before & after

The same measures, either side of the work.

Each pair is scaled against its own larger value, so the comparison is honest rather than flattering.

Median intake turnaround

−95%
Before
60 hours
After
3 hours

Administrative hours per week on intake

−80%
Before
15 hours
After
3 hours

Fields keyed by hand per document

−92.3%
Before
26 fields
After
2 fields

05 — The numbers

Median intake turnaround.

Tracked as hours, from Baseline through Wk 7 — low 3, high 60.

IntakeLens — Median intake turnaround, hours, 2026.

07 — Stack & role

Built with.

  • Python (spaCy, pdfplumber, scikit-learn)
  • Claude API
  • FastAPI
  • PostgreSQL
  • Cloud Run
Role
Cork AI Consulting — founder & lead consultant, sole delivery
Duration
7 weeks
Client
Cork AI Consulting · professional services firm
Period
2026

08 — Questions

The questions I get asked about this one.

Reviewing everything reproduces the original cost. The floor concentrates human attention on the documents where it changes the outcome, and it is set from the firm’s own tolerance for a misfiling rather than from a model default.

Documents stay inside the firm’s own cloud tenancy, access is logged per document, and retention follows their existing policy. That constraint shaped the architecture from the first week rather than being retrofitted.

It is measured on a held-out sample, not on the training data, and every human correction in the live log is a fresh test case. Accuracy is re-reported weekly rather than claimed once at launch.

Got a report that takes two days to assemble?

That is usually a one-week fix. Tell me what you are reconciling by hand and I will tell you what I would automate first.