XtractSol
A hand pulling one red binder from a row of identical binders
Document Classification
Document Intelligence

Sort the mail, so your team can do the work.

Our agents recognise what each incoming document is, tag it, split bundled files into their parts, and route each one to the right place. Anything unclear goes to a person instead of a guess.

Recognised by what the document says, not what the file is called
80%+
Of purchase transactions require reviewing at least eleven documents before closing.
ALTA, March 2026
21%
Involve more than fifty records tied to a single property's ownership history.
ALTA, March 2026
52%
Of title professionals spend eleven hours or more each month on fraud prevention.
ALTA, March 2026
15%
Spend more than fifty hours a month on it, including forged property documents.
ALTA, March 2026
KNOWING WHAT A DOCUMENT ACTUALLY IS, RATHER THAN WHAT ITS FILE NAME CLAIMS, IS THE FIRST STEP IN BOTH SORTING AND SCRUTINY.
SOURCE: AMERICAN LAND TITLE ASSOCIATION, MEASURING THE COMPLEXITY OF TITLE PRODUCTION, PUBLISHED 31 MARCH 2026, SURVEY OF 449 TITLE PROFESSIONALS ACROSS 47 STATES.

How much of your team's morning goes on working out what arrived?

Thirty minutes with the engineer who builds these agents. Bring a real batch.

Book the call
Document Classification
Before the work starts

Six ways the sorting slows you down.

Before anyone can act on a document, somebody has to work out what it is and where it belongs. At volume, that is a job on its own.

  • 01 A file named scan_0714 that could be any of nine things
  • 02 Twenty-two pages in one PDF, holding four separate documents
  • 03 A document that reaches the right team two days after it arrived
  • 04 Something filed under the wrong transaction and quietly lost
  • 05 A shared inbox where everyone assumes somebody else is triaging
  • 06 A search that fails because nobody tagged it when it landed
Document Classification
What we build

What goes into XtractSol classification

Each agent is built around your document types, your routing rules and the systems it feeds. Start with a handful of types or cover the whole intake.

01 / RECOGNITION

Identified by content, not file name

The agent reads each document to work out what it is, so contracts, statements, invoices and correspondence are identified correctly whatever they were saved as.

02 / TAGGING

Findable later, not just filed now

Beyond the type, each document is tagged with the details that matter for locating and routing it, so nobody hunts through folders a month from now.

03 / ROUTING

Straight to the right place

Once identified, a document goes to the right team, file or workflow on your rules, rather than waiting in a general inbox for somebody to redirect it.

04 / SPLITTING

Bundled scans come apart

When several documents arrive in one file, the agent finds where each begins and ends, classifies each part, and routes them separately.

05 / EXCEPTIONS

Unclear goes to a person

When the agent is not confident, it hands over rather than guessing. Automation on the clear cases, a clean handoff on the ambiguous ones.

06 / DELIVERY

Into the systems you use

Classified and tagged documents land where the work happens, so the sorting connects to the workflow instead of living in a separate tool.

Document Classification
A worked example

One scan, four documents, four destinations

A bundled scan is the case that costs the most time and gets talked about the least. Someone has to open it, work out where one document ends and the next begins, then file each part separately.

Incoming · scan_batch_0714.pdf22 pages, one file
1 to 67 to 1112 to 1415 to 1920 to 22
Page 1Page 22
1 to 6Warranty deedHigh confidenceRecording queue
7 to 11Deed of trustHigh confidenceRecording queue
12 to 14Tax certificateHigh confidenceFile, tax section
15 to 19CorrespondenceTagged to the transactionFile, general
20 to 22Not confidently identifiedHandwritten cover and two unclear pagesHeld for a person
FOUR DOCUMENTS ROUTED, ONE HELD. ILLUSTRATIVE VIEW OF A BUNDLED SCAN, NOT AN ACTUAL FILE.

The last three pages are the point. The agent could have forced them into the nearest matching category and the batch would have looked completely processed. Instead they sit in one place, clearly marked, waiting for somebody to spend a minute on them. A document held back is a minute of work. A document confidently filed in the wrong place is a problem that surfaces weeks later, usually when someone is looking for something else.

Document Classification
Human in the loop

Confident on the clear cases. Careful on the rest.

You set how sure the agent has to be before it routes a document on its own.

01

Read

Each document read for what it contains, not what it was named.

Agent
02

Split

Bundled files separated where one document ends and the next starts.

Agent
03

Classify

A type assigned, with the confidence behind it recorded.

Agent
04

Route

Clear cases sent to the right team, file or workflow on your rules.

Agent
05

Hold

Anything below your threshold stops rather than being filed.

Agent
06

Decide

A person sorts the held items, and the correction feeds back in.

Your team
Agent handlesPerson decidesEvery routing decision logged with the confidence behind it
Document Classification
The honest version

Why we would rather hand a document back

Held beats misfiled

A classifier that never hesitates is not more accurate, it is just quieter about being wrong

Any classifier can be tuned to route everything. It will look excellent on a dashboard, because a document sent somewhere is a document processed. The cost shows up later and somewhere else, when a file is missing a document that was filed against a different transaction, and nobody knows to look for it because the queue said it was handled.

So we tune for the opposite behaviour. The agent routes what it is confident about and holds what it is not, and you set where that line sits. Push the threshold up on document types where a mistake would be expensive and more items reach a person. Push it down where a mistake is cheap to fix. That trade belongs to you, and it should be a deliberate decision rather than a default nobody chose.

ConfidentClear cases routed automatically on your rules
CautiousAnything below your threshold held, never forced
LoggedEvery decision recorded with the confidence behind it
A held document costs a minute. A confidently misfiled one costs whatever it costs to notice it is missing.
You set the thresholdPer document type, so caution goes where the consequences are, rather than being applied evenly to everything.
Document Classification
Intake and delivery

Where documents arrive, and where they land

We connect to the channels your documents actually come in on, and deliver into the systems your team already works in.

EmailScanned batchesFaxVendor portalsWeb uploadsSharePointResWareQualiaSoftProRamQuestOutlookIn house systems

Accuracy varies by document type and condition. Clean typed documents that differ clearly from one another are easier than near identical forms or handwritten material. We set realistic expectations for your intake during the audit.

Document Classification
Control

Speed without misfiled documents

  • Confidence threshold set by you, per document type
  • Anything below it held for a person rather than routed
  • Every classification and routing decision logged and reversible
  • Your documents are never used to train any model
Document Classification
How we build it

You see the accuracy before it routes real work

We learn what lands in your inbox and where it all needs to go, then build in stages you can inspect.

STEP 01

Discovery call

What you receive, how much of it there is, and how long the sorting takes today.

STEP 02

Workflow audit

Your document types and routing rules mapped against real examples of what actually arrives.

STEP 03

Build and tune

Tuned on your real documents, including the near identical ones, with thresholds agreed per type.

STEP 04

Deploy and refine

Connected to your channels and systems with review in place, then tracked as new types appear.

Document Classification
Good to know

Questions we actually get

How many document types can it recognise?

As many as your intake actually includes, from a handful to a broad mix. What matters more than the count is how similar the types are to each other. Ten clearly different documents are easier than four that look almost identical, and we tell you which situation you are in during the workflow audit.

What happens to documents it cannot identify?

They go to a person rather than being guessed at. You get automation on the clear cases and a clean handoff on the ambiguous ones, so nothing is misfiled just to keep the queue moving. You also set how confident the agent has to be before it routes something on its own.

Can it handle files with several documents in one?

Yes, and this is usually the part that saves the most time. When several documents arrive bundled in a single scan, the agent works out where each one starts and ends, classifies each part separately, and routes them independently. A bundled file stops being a manual sorting job.

What happens if it routes something to the wrong place?

Every routing decision is logged with the confidence behind it, so a misroute is findable rather than mysterious. Corrections feed back into tuning, and you can raise the confidence threshold on any document type where a mistake would be expensive. More caution means more documents reaching a person, and that trade is yours to set.

What happens when a document type we have never seen arrives?

It is flagged as unrecognised rather than forced into the closest matching category. New types are added as they appear, so the same document is handled cleanly the next time it arrives.

What does it cost?

A build fee plus a rate per document that scales with volume. You get a firm number after the workflow audit, before you commit to anything.

Ready to get started?

Book a free discovery call and we will map how agentic AI can fit your workflows.