Handwriting OCR: choose by script, page and evidence
Start with the handwriting and output you actually need: cropped modern text lines, whole pages, historical manuscripts or structured forms.
Choose by workload
| Workload | Candidate path | Important limitation |
|---|---|---|
| Cropped English handwriting lines | TrOCR handwritten checkpoints as a reproducible local baseline. | Line recognition needs reliable segmentation; an IAM-fine-tuned checkpoint is not evidence for every script or whole-page input. |
| Whole pages and mixed printed/handwritten content | Document OCR such as Surya, plus a current multimodal model tested on identical pages. | Reading order, missing lines and hallucinated corrections must be scored separately. |
| Historical or specialized scripts | A transcription system adapted to that collection and script. | A modern English handwriting result cannot predict performance on historical spelling or another alphabet. |
| Forms, amounts and identifiers | Recognition plus field extraction and explicit review rules. | A low average CER can still conceal wrong amounts, IDs or field assignments. |
Candidate shortlist; row order does not represent an accuracy or speed ranking.
What the cited research actually covers
Benchmarking Large Language Models for Handwritten Text Recognition was first submitted in March 2025 and revised in June 2025. It compares proprietary and open models with Transkribus across modern and historical datasets in several European languages. Its findings are tied to those model versions and conditions.
That paper does not substantiate the former precise IAM results for GPT-5, Opus 4.7 or Gemini 3. Historical research can guide a candidate list, but current release names need their own model IDs, prompts, datasets and outputs.
The 2025 HTR study, version 3 ↗TrOCR: a documented line-recognition baseline
microsoft/trocr-base-handwritten is an IAM-fine-tuned checkpoint intended for single text-line images. Crop a line before inference rather than sending an entire multi-line page. The example follows the model card with the Transformers 4 API; newer major versions require their own compatibility check.
Documentation-based example; no runtime measurement is claimed.
# pip install "transformers<5" torch pillow
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
checkpoint = "microsoft/trocr-base-handwritten"
processor = TrOCRProcessor.from_pretrained(checkpoint)
model = VisionEncoderDecoderModel.from_pretrained(checkpoint)
image = Image.open("handwritten-line.png").convert("RGB")
pixels = processor(images=image, return_tensors="pt").pixel_values
ids = model.generate(pixels)
print(processor.batch_decode(ids, skip_special_tokens=True)[0])TrOCR handwritten model card and limitations ↗Whole-page OCR is a separate evaluation
Surya exposes page OCR and layout output through its current inference manager. It is a candidate to test, not a handwriting winner established here. Keep line segmentation errors, omitted regions and recognition errors separate when comparing it with a line-based model.
For multimodal APIs, pin the deployed model ID, prompt and image settings. Ask for literal transcription with an explicit unreadable marker. Score invented words and missing lines, and retain the original output before any language-model cleanup.
Surya current interface, backends and license ↗A credible CER or WER result needs a protocol
State the dataset version, official split, writer split, preprocessing, line crops, case and punctuation rules, Unicode normalization, decoding settings and language-model assistance. Define CER as character edit distance divided by reference character count, and WER using reference word count. Include empty-output handling and the aggregation method.
Evaluate writers and documents outside the tuning set. Report error distributions by writer, language and document quality, alongside human review needs. Publish predictions and scoring code so another person can reproduce the number. Never compare a line-level CER with whole-page recognition without accounting for segmentation and layout failures.
Primary sources
Documentation and licensing were checked on 7 October 2026. Pin package versions, model checkpoints and configuration in your own environment; upstream defaults can change.