Document model · Fast-moving · Intermediate
OCR and document models
Models that read text and layout from images and PDFs, including handwriting and complex tables.
What OCR and document models is
Modern document models return structure — reading order, tables, key-value pairs — rather than a flat string, which is what downstream automation needs.
How it works
Pipelines combine detection and recognition, or use a multimodal model that outputs structured markup directly. Confidence scores route low-certainty fields to human review.
Why it matters
It is the entry point for most enterprise document automation, and accuracy at this stage caps everything downstream.
Common uses
- →Invoice and form processing
- →Archive digitisation
- →Receipt and ID capture
- →Table extraction
Strengths
- ✓Handles messy real-world scans
- ✓Structure-aware output
Watch for
- ✓Handwriting and poor scans still fail
- ✓Language coverage varies
Continue exploring
More in this collection
Browse all AI Models