If you have ever worked with a scanned document, then you might understand the process by which scanned documents can become searchable and selectable. The answer lies in OCR technology.
PDF files contain data. When a PDF file is created from an image — such as a scanned page — it is stored as an image of the original text, rather than as actual characters. A PDF containing only images of a text page isn’t searchable or selectable unless OCR (Optical Character Recognition) is applied to it. OCR embeds the underlying text that makes up the original page so that it becomes searchable and selectable. You will be able to select and copy the text without any problem.
There are many tools available for performing OCR on your PDF. There are two basic approaches: either convert your PDF to a format like Word or plain text, then perform OCR on that file; or use a tool that performs OCR directly on the PDF itself. Many online services allow you to upload a file, then recognize the text automatically. For example, a free service provides this, along with some additional options. You’ll find other similar tools at ilovepdf.com.
The best way to perform OCR on your PDF is using Acrobat Pro, Adobe’s desktop product that includes OCR capabilities. If you’ve already converted your PDF to an accessible text file, however, then all that remains is to open that file and begin editing. But first, you must ensure that the file contains the correct information. There are several common problems that persist even after OCR has been applied:
Messed-up text spacing
Garbled characters
Inability to copy the text
Completely distorted formatting
To see if a PDF has been OCR’d, there’s no instant visual indicator that you can rely upon. The only practical method I’ve found is to attempt to extract text from the PDF. Either try opening it in a program like Adobe Reader, or use a tool like PDF2Text. If you don’t get text output, then you know that OCR hasn’t been performed yet.
I should also mention what most people don’t understand about converting a PDF to text: it does not mean that you now understand the content of the PDF. These are two completely different things. Converting a PDF to text is simply extracting the information from the file into another form of text — but nothing more. You may think that you understand what a document says, but it is very possible that you’ve missed out on key details that make a difference in how you interpret the document.