Blog
OCR & Scanning

Blog

Converting Handwritten or Scanned Exam Papers to Digital Text: An OCR Guide

What actually happens when a scanned or photographed exam paper gets converted to digital text, and why math and diagrams need special handling.

FFormatly Team4 min read

Not all OCR is equal

Basic OCR — the kind built into most scanner software — is tuned for plain prose: paragraphs of regular text in a single language, no math, no dense tabular layouts. Point it at an exam paper full of equations, chemical formulas, multi-column layouts and answer-option grids, and it tends to fall apart in predictable ways: fractions become a jumble of slashes and numbers, subscripts and superscripts disappear, and reading order across columns gets scrambled.

What good math-aware OCR looks like

The fix is OCR that specifically recognizes mathematical notation as math, not as loose text — a fraction gets reconstructed as an actual fraction, an integral or a matrix keeps its real structure, and chemistry notation like subscripted formulas survives instead of collapsing into plain digits. Formatly's conversion pipeline is built on this kind of OCR, which is also what makes it usable on genuinely rough source material — a slightly skewed phone photo of a textbook page, not just a clean flatbed scan.

Diagrams don't get left behind either

A circuit diagram, a graph, or a labelled figure that's part of a question can't be turned into text at all — it has to be cropped out of the page and carried through as an actual image, attached to the right question, option or solution. Losing that image (or worse, having it silently attach to the wrong question) makes a converted paper useless for anything involving physics, chemistry diagrams, or geometry. Getting that association right, even on a page where the source layout is dense or jumbled, is one of the harder parts of the whole pipeline — and one worth checking a converted file for before you trust it.