Home / OCR / Handwriting OCR: choose by script, page and evidence
Practical guide · Sources checked 7 October 2026

Handwriting OCR: choose by script, page and evidence

Start with the handwriting and output you actually need: cropped modern text lines, whole pages, historical manuscripts or structured forms.

Evidence status. The previous IAM table assigned precise CER values to GPT-5, Opus 4.7 and Gemini 3 without corresponding runs in the listed references. Those scores and the resulting winner claims have been withdrawn. This page does not establish an October 2026 handwriting leaderboard.

Choose by workload

WorkloadCandidate pathImportant limitation
Cropped English handwriting linesTrOCR handwritten checkpoints as a reproducible local baseline.Line recognition needs reliable segmentation; an IAM-fine-tuned checkpoint is not evidence for every script or whole-page input.
Whole pages and mixed printed/handwritten contentDocument OCR such as Surya, plus a current multimodal model tested on identical pages.Reading order, missing lines and hallucinated corrections must be scored separately.
Historical or specialized scriptsA transcription system adapted to that collection and script.A modern English handwriting result cannot predict performance on historical spelling or another alphabet.
Forms, amounts and identifiersRecognition plus field extraction and explicit review rules.A low average CER can still conceal wrong amounts, IDs or field assignments.

Candidate shortlist; row order does not represent an accuracy or speed ranking.

What the cited research actually covers

Benchmarking Large Language Models for Handwritten Text Recognition was first submitted in March 2025 and revised in June 2025. It compares proprietary and open models with Transkribus across modern and historical datasets in several European languages. Its findings are tied to those model versions and conditions.

That paper does not substantiate the former precise IAM results for GPT-5, Opus 4.7 or Gemini 3. Historical research can guide a candidate list, but current release names need their own model IDs, prompts, datasets and outputs.

The 2025 HTR study, version 3 ↗

TrOCR: a documented line-recognition baseline

microsoft/trocr-base-handwritten is an IAM-fine-tuned checkpoint intended for single text-line images. Crop a line before inference rather than sending an entire multi-line page. The example follows the model card with the Transformers 4 API; newer major versions require their own compatibility check.

Documentation-based example; no runtime measurement is claimed.

# pip install "transformers<5" torch pillow
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

checkpoint = "microsoft/trocr-base-handwritten"
processor = TrOCRProcessor.from_pretrained(checkpoint)
model = VisionEncoderDecoderModel.from_pretrained(checkpoint)
image = Image.open("handwritten-line.png").convert("RGB")
pixels = processor(images=image, return_tensors="pt").pixel_values
ids = model.generate(pixels)
print(processor.batch_decode(ids, skip_special_tokens=True)[0])
TrOCR handwritten model card and limitations ↗

Whole-page OCR is a separate evaluation

Surya exposes page OCR and layout output through its current inference manager. It is a candidate to test, not a handwriting winner established here. Keep line segmentation errors, omitted regions and recognition errors separate when comparing it with a line-based model.

For multimodal APIs, pin the deployed model ID, prompt and image settings. Ask for literal transcription with an explicit unreadable marker. Score invented words and missing lines, and retain the original output before any language-model cleanup.

Surya current interface, backends and license ↗

A credible CER or WER result needs a protocol

State the dataset version, official split, writer split, preprocessing, line crops, case and punctuation rules, Unicode normalization, decoding settings and language-model assistance. Define CER as character edit distance divided by reference character count, and WER using reference word count. Include empty-output handling and the aggregation method.

Evaluate writers and documents outside the tuning set. Report error distributions by writer, language and document quality, alongside human review needs. Publish predictions and scoring code so another person can reproduce the number. Never compare a line-level CER with whole-page recognition without accounting for segmentation and layout failures.

Primary sources

Documentation and licensing were checked on 7 October 2026. Pin package versions, model checkpoints and configuration in your own environment; upstream defaults can change.

OCR benchmark registry · How we verify results