
PDF & documents6 min read
Getting tables out of a PDF and into Excel
A PDF has no cells, only positioned glyphs. How extractors infer columns from ruling lines and x coordinates, where that fails, and how to verify the result.
Merging, OCR, pulling tables into Excel and redaction.

A PDF has no cells, only positioned glyphs. How extractors infer columns from ruling lines and x coordinates, where that fails, and how to verify the result.

OCR accuracy is set by the image, not the engine: resolution, binarisation, skew, compression artefacts, language data and layout. Plus how to judge the output.