Can't Copy Text from a PDF? OCR, Permissions and Layout Explained
Learn why PDF text cannot be selected or pastes as gibberish, and choose the right fix for scans, permissions, fonts and complex layouts.
What to do first
Zoom in and try selecting individual characters. If selection draws one large box over the page, the PDF is probably a scan and needs OCR. If text is selectable but copying is blocked, check document permissions and request an authorized copy.
A PDF is a container, not a guarantee that every visible word is stored as usable text. A page may be a photograph, letters may be drawn with custom character maps, or copying may be restricted by the document owner.
The fix should match the cause. OCR helps image-only scans, but it does not override permissions. Converting a complex layout to Word may make editing easier, but it can change columns, tables and line breaks.
What this problem looks like
- Dragging selects the whole page as one image.
- No text cursor appears over visible words.
- Copied text pastes as boxes, symbols or wrong characters.
- Words copy in the wrong order across columns.
- The viewer explicitly reports that copying is not allowed.
The most common causes
Image-only scan
The page contains pixels but no text objects. It looks readable to a person, yet the computer has no characters to select.
Missing or unusual character mapping
Embedded subset fonts can display correctly while mapping glyphs to the wrong Unicode characters during copy and paste.
Reading-order problem
PDF stores positioned objects. Multi-column pages, sidebars and tables may not include a logical order for extraction.
Owner permissions
The PDF can restrict copying or extraction. Respect the restriction and obtain an authorized source or password.
Viewer limitation
A lightweight browser viewer may offer less reliable selection than a full reader, particularly with rotated or layered text.
How to fix the problem safely
- 1
Identify image versus text
Zoom to 300 percent. If letters become blocky pixels and selection covers a large rectangle, treat the page as a scan. If individual words highlight, text objects exist.
- 2
Check permissions
Open document properties and inspect security. If copying is not allowed, ask the owner for an accessible copy or the relevant password. Do not try to bypass restrictions without authorization.
- 3
Try a full PDF reader
Save the file locally and test selection in a current desktop reader. This rules out browser-selection bugs before any conversion.
- 4
Run OCR on scans
Choose the correct language and process only the required pages first. OCR adds a searchable text layer while keeping the page image. Review names, numbers and punctuation afterward.
- 5
Extract plain text for reuse
Use PDF-to-text when wording matters more than layout. Expect headers, footers and columns to require cleanup.
- 6
Convert to Word for structured editing
Use PDF-to-Word when paragraphs and images should remain broadly editable. Compare tables, footnotes and page breaks with the original.
- 7
Correct reading order manually
For columns or tables, extract one region at a time or use a table-specific extractor. Copying the entire page often mixes unrelated blocks.
OCR quality: what changes the result
OCR works best on straight, high-contrast pages around 300 dpi. Skewed phone photos, shadows, handwriting, decorative fonts and mixed languages reduce accuracy. More resolution is not always better; very large noisy images can slow processing without improving recognition.
Always proofread critical data. OCR can confuse O and 0, I and 1, decimal separators, hyphens and accented characters. It should accelerate transcription, not replace verification.
- Select the languages actually present in the document.
- Rotate pages upright before OCR.
- Crop dark borders and desk background where possible.
- Check names, account numbers, dates and totals manually.
Accessible PDFs and reliable reading order
A tagged PDF can contain headings, lists, alternative text and a reading order for assistive technology. OCR alone does not create a fully accessible document; it mainly supplies characters.
If accessibility or publication quality matters, return to the source document and export a properly tagged PDF. A repaired extraction copy is not a substitute for an accessible source.
Checklist before you finish
- Determine whether the page is an image or real text.
- Check copying permissions before processing.
- Use the correct OCR language.
- Process a short sample before a long document.
- Choose text extraction or Word conversion based on the goal.
- Proofread critical values against the page image.
When the file itself needs work
Work on a copy. Uploads use an encrypted connection and expire automatically after the short job window.
Frequently asked questions
Why can I select the page but not individual words?
The page is probably an image scan. OCR is required to create a selectable text layer.
Why does copied text become strange symbols?
The PDF may use an incomplete Unicode character map or a custom embedded font. OCR or conversion can create a new character mapping.
Does OCR change the appearance?
A searchable-PDF workflow normally preserves the page image and adds invisible text. Conversion to Word can change layout.
Can OCR bypass copy restrictions?
OCR can recognize pixels, but you must still respect copyright, privacy and document permissions. Obtain authorization before extracting restricted content.