Cut LLM Document-Extraction Cost by 85% Without Losing Accuracy
Cut LLM Document-Extraction Cost by 85% Without Losing Accuracy
Most production LLM extraction pipelines route every field through a frontier model, every time. It works, and for a while nobody questions the bill. Then volume grows, or someone runs the per-document math, and the real question shows up. How much of this spend is buying accuracy, and how much is just habit. In practice, most of it is habit. The majority of fields in a typical extraction job do not need a frontier model at all. A large share of the model calls answer questions the document never asked. A surprising fraction of calls re-extract data that has not changed since the last run. None of that is a model problem. It is an architecture problem, and architecture problems are the good kind, because you can fix them without waiting on anyone's roadmap.
Set a Trusted Baseline First
Before optimizing anything, run the workload on a frontier model and treat those results as the baseline. Early on the goal is not low cost. It is output you trust and a confidence signal you believe. Once that baseline exists, it becomes the yardstick for everything cheaper.
Step-by-step process:
- Run the full pipeline on your current frontier model and capture the outputs as ground truth.
- Measure accuracy on a held-out sample to establish what "good" looks like.
- Use those results as the reference for all subsequent optimizations-only replace a field with a cheaper model if it matches the baseline within acceptable tolerance.
TL;DR: Set the baseline with the expensive model first, then prove a cheaper model can match it.
Adaptive Field Selection
The cheapest extraction call is the one you never make. Even though a contract schema may define 100+ fields, a given document-such as a simple service agreement-might only contain 40 to 50 of them. Extracting the rest produces "not found" responses that cost tokens and return nothing.
How it works: Classify the document type first, then extract only the fields that can actually appear in it. This one move eliminates 30 to 40 percent of calls per document.
Impact: By reducing the number of actual extractions, you shrink the entire downstream cost surface. Every later optimization then operates on a smaller base.
Right-Size the Model to the Field (Tiered Routing + Distillation)
Not every field needs the same horsepower. In a typical extraction job:
- ~65 to 70 percent of fields are structurally simple (dates, amounts, party names, reference numbers, categorical flags). These have predictable formats and consistent locations. A lightweight model handles these at a fraction of frontier cost with equivalent accuracy.
- ~20 percent need moderate reasoning (interpreting indirect phrasing, resolving ambiguous references, synthesizing across sections). A mid-tier model (for example a Haiku or Nova Lite class model) is plenty.
- ~10 percent genuinely need a frontier model (obligation summaries, risk assessment, termination-condition analysis, cross-document context).
Routing logic: Route each field to the cheapest tier that can do the job. The routing decision is nearly free-it's a field-type to tier lookup table plus a confidence-based fallback-so a tier that returns low confidence escalates to the next one. No ML is needed for routing because the mapping is stable for a given document corpus.
Result: Blended per-document cost drops 60 to 75 percent, and quality holds because the hard fields still get frontier attention.
Distillation: Earn Your Way to a Cheaper Tier 1
Tiered routing gets better when Tier 1 is a model you distilled for your own documents. Here's the process:
- Train a teacher (a frontier model) on 200 to 500 representative documents to produce gold-standard extractions.
- Distill a student model (smaller, cheaper) to reproduce that behavior on your specific fields.
- Managed services such as Amazon Bedrock Model Distillation handle the training without requiring ML expertise.
Why it works: The student model learns your narrow task-the dates, amounts, party names, and term lengths you see every day-and runs much cheaper and faster with a small accuracy gap (about 2 percent) on the tasks it was trained for.
Progression path:
- Months 1-2: Run routing-only on general-purpose models (40 to 50 percent savings) while storing every validated result as future training data.
- Month 3 onward: Once enough examples accumulate, run distillation so Tier 1 becomes your student model-a further 30 to 40 percent off.
- Quarterly: Re-distill with more data, and the share of fields handled by the cheapest tier climbs from about 70 percent toward 85 percent.
Stop Paying Twice: Field-Level Caching
Documents get reprocessed when schemas are updated, prompts improve, or a quality review flags systematic errors. Without caching, reprocessing 200 documents costs exactly what processing them the first time did.
Caching strategy: Key each result by the document-section SHA-256, the field name, the prompt version, and the model version. If all four match, return the cached value with zero inference.
Impact: When you update a prompt for one field type, only that field's cache entries invalidate, and everything else serves from cache. A prompt update typically touches three to five field types, so 85 to 90 percent of fields serve from cache on reprocessing.
Example: For a 200-document corpus reprocessed after a four-field prompt change, that's 800 extractions instead of 7,000-a 88 percent reduction on that cycle.
Stop Paying for Idle: Serverless Scale-to-Zero
A 200-document-per-month workload is idle most of the time. Documents land in object storage, a lightweight function writes metadata that triggers a stream event, and a containerized extraction agent runs only when work exists. During idle periods the system costs nothing beyond storage.
Implementation: KEDA provides a clean way to achieve scale-to-zero and event-driven scaling for this on Kubernetes.
Why These Compound Instead of Add
A common mistake is to add the savings together. Distillation saves 75 percent, adaptive selection 35 percent, caching 30 percent, so the total must be 140 percent-that is impossible. The techniques do not each operate on the original cost. Each operates on the residual the previous stage left behind. They address different sources of waste, so they multiply:
| Stage | Technique | Operates On |
|---|---|---|
| 1 | Adaptive field selection | Total field count (60 to 70 percent of baseline) |
| 2 | Distillation + tiered routing | Remaining fields (20 to 25 percent of baseline) |
| 3 | Field-level caching | Reprocessed documents (10 to 15 percent on reprocessing) |
| 4 | Serverless scale-to-zero | Infrastructure (zero idle cost) |
End-to-end result: Roughly ((1 - 0.35) \times (1 - 0.65) \times (1 - 0.25)), landing around 0.07 to 0.15 of the original-an 85 to 93 percent reduction.
Each technique matters most at a different stage. Adaptive selection eliminates unnecessary work, distillation reduces the cost of necessary work, caching eliminates repeated work, and serverless eliminates idle overhead. No one of them gets you there alone.
Making It Safe: The Quality-Floor Monitor
Aggressive cost optimization earns an obvious objection. If 70 percent of fields go to a cheaper model, you need a way to know it is not silently getting worse, or drifting as document patterns change.
Monitoring approach: Each week, randomly re-run 5 to 10 percent of Tier-1 fields through the frontier model as a shadow extraction and compare. Agreement at or above 98 percent means the floor is holding. Below 95 percent raises an alert. Below 90 percent auto-triggers re-distillation on the last three months of data.
Shadow validation adds only about 3 to 5 percent to total cost, and it means the system monitors itself with no dedicated QA team required. The same monitor also spots promotions-if mid-tier and frontier models consistently agree on a field currently routed to Tier 2, that field is a candidate to fold into the next distillation cycle, so the system keeps shifting work to the cheapest tier on its own.
The Payoff Beyond the Bill
The reason this is worth doing well is not only the savings on an existing workload. Every percentage point off per-unit cost widens the set of use cases where automation is worth building at all. A workload that only penciled out for one of your five document types at naive pricing might clear the bar for three or four once the architecture is right.
Cost optimization is not a late-stage cleanup. Treat it as a first-class design requirement alongside accuracy, latency, and reliability, and more of what you want to automate becomes viable.
Open-Source Patterns
Two of the ideas above are available as small, dependency-light libraries under AWS Samples (MIT-0):
- aws-samples/sample-textract-field-memory - Spatial field location memory for document processing pipelines. Learns field positions, validates extractions, identifies document types by layout, detects drift, and monitors template health. Zero dependencies, pure Python.
- aws-samples/sample-prompt-correction-memory - Self-improving LLM document extraction on AWS. Each human QA correction improves future extractions via few-shot self-healing and deterministic rule graduation-no retraining, no redeployment.
Disclaimer: This is sample code for non-production usage. You should work with your security and legal teams to meet your organizational security, regulatory, and compliance requirements before deployment. When processing documents containing PII (e.g., SSNs), PHI, or payment data, ensure your implementation meets applicable compliance requirements (HIPAA, PCI-DSS, GDPR, etc.). The libraries store only field names and bounding-box coordinates-never field values or document content-but you remain responsible for securing the underlying documents.
Comments
No comments yet. Start the discussion.