PDF tools

How to Extract Text from a PDF

How to extract text from a PDF without the jumbled lines. Why copy and paste breaks, how to keep paragraphs intact, and what to do when the PDF is a scan.

6 min readUpdated Aug 20, 2026

Selecting text in a PDF reader and pressing Ctrl+C works right up until it does not. The text arrives with a line break after every line rather than at the end of each paragraph, columns interleave into nonsense, or nothing is selectable at all. None of that is your reader misbehaving - it follows from how a PDF stores text. This guide explains what is going on, and how to get clean text out with the PDF to Text converter.

Why copying text out of a PDF goes wrong

A word processor document stores a paragraph as a paragraph. A PDF does not. A PDF is a set of drawing instructions, and text is drawn the same way a rectangle is: put this glyph at these coordinates in this font at this size. There is no record of where a line begins, where a paragraph ends, or which block of text is a column and which is a footnote. All of that is something your eye reconstructs from the layout.

So extracting text means rebuilding structure that was never stored. Three consequences follow, and they explain almost every bad paste you have ever had:

  • Lines are separate objects. A paragraph that wrapped over five lines is five unrelated runs of text at five heights. Paste it and you get five hard breaks, which re-wrap horribly the moment you change font or page width.
  • Spaces are often not there. Many PDF writers position the next word instead of emitting a space character. Extract naively and you get 'InvoiceTotal' where the page clearly reads 'Invoice Total'.
  • Draw order is not reading order. Instructions arrive in whatever sequence the producing software emitted them, which for a two-column layout, a table or a form is often not the order a human reads them in.

A good extractor therefore works from geometry rather than from file order: it groups text by vertical position into lines, sorts each line left to right, and inserts a space wherever the horizontal jump between two runs is wider than the gap inside a word. That is what happens when you drop a file into the tool.

Extract the text in four steps

  1. Open PDF to Text and drop your PDF onto the box. It is read inside your browser tab, so the file is not uploaded.
  2. Pick a layout. Flowing paragraphs stitches wrapped lines back together; Keep line breaks leaves every line exactly where it was. The next section covers which to choose.
  3. Optionally type a page range such as 1-3, 7 to take the text from only part of the document, and tick Mark where each page starts if you want a marker between pages.
  4. Click Extract text. The result appears in an editable box - tidy it there if you like, then Copy it to the clipboard or Download .txt.

Flowing paragraphs or keep line breaks?

This one setting decides whether the output is pleasant or infuriating, and the right answer depends entirely on what the page contains. Take a short document with a wrapped paragraph followed by an address block:

Our records show the invoice was issued on 4 March and remains / unpaid as of today. Please settle it at your earliest convenience. / Accounts Department / 14 Residency Road / Bengaluru 560025

The first two lines are one sentence that happened to wrap; the last three are an address where the line breaks are the point. Flowing paragraphs rejoins the wrapped pair into one sentence, and the address - set off by a wider vertical gap - starts a new block. Keep line breaks gives you five lines instead, with the sentence still split after the word 'and'.

The rule of thumb is straightforward. Choose flowing paragraphs for anything you intend to re-read or re-flow: articles, reports, letters, book chapters. Choose keep line breaks wherever a line is a unit of meaning: addresses, tables, code, poetry, and any file a script will read line by line.

How does it tell a wrapped line from a new paragraph? By measuring. It takes the vertical gap between consecutive lines, finds the typical value for that document, and starts a new paragraph wherever a gap is noticeably larger. Documents vary in leading, so a fixed threshold in points would break on the next file; measuring the document against itself does not.

Taking the text from only some pages

Long documents rarely need extracting whole. A page range accepts individual pages and spans together - 1-3, 7, 12-14 - and is validated against the real page count as you type, so a typo is caught before anything runs. It is the quickest way to pull one clause out of a contract, and it keeps the output small enough to read.

When the PDF is a scan

If extraction returns nothing, the usual reason is that there is nothing to return. A scanned document, or a photograph of a page saved as a PDF, contains an image and no characters whatsoever. Your eye reads the words; the file holds pixels. The quick test is to try selecting a word in any PDF reader: if no selection highlight appears, there is no text layer, and no extractor of any kind can produce text from it.

The fix is optical character recognition, which looks at the picture and works out which letters it shows. Run the file through OCR PDF first to add a real, selectable text layer, then extract from the result. Accuracy depends on the scan: clean 300 DPI print does very well, while an angled phone snapshot or heavy handwriting needs proofreading.

What extraction cannot carry across

Plain text is plain by definition, so bold, italics, font sizes, colours and images are all dropped. That is usually the point - text is what you wanted - but it does mean you should pick the right tool for the job. If you want an editable document that keeps a document shape, PDF to Word produces a DOCX from the same extraction. If what you actually need is the numbers out of a table, PDF to Excel detects rows and columns from the same coordinates and gives you a spreadsheet, which is far better than trying to rescue a table from a wall of text.

Two smaller things are worth knowing. Text drawn inside an image is not text, so a chart with its labels baked into the picture extracts nothing even in an otherwise selectable document. And the downloaded .txt is UTF-8 with a byte order mark, which is what makes accents, currency symbols and non-Latin scripts open correctly in Notepad on Windows rather than as mojibake.

Nothing leaves your browser

Extraction runs entirely on your own machine. The PDF is opened, the text layer read and the .txt assembled by code running in the page you have open, and no part of the document is sent anywhere - which matters, because the files people most need text out of are contracts, statements and medical letters. You can confirm it the blunt way: load PDF to Text, disconnect from the internet, and extract a file. It still works.

Frequently asked questions

Why does my extracted text have a line break after every line?
Because the PDF does, in effect - each visual line is stored as its own run of text, with nothing recording that five of them are one paragraph. That is what the Flowing paragraphs layout is for: it measures the document's usual line spacing and rejoins lines that are merely wrapped, breaking only where the vertical gap is bigger than normal. If you are getting one line per line, switch the layout from Keep line breaks to Flowing paragraphs.
Can I extract text from a password-protected PDF?
Yes, if you know the password. Dropping a protected file in prompts you for it, unlocks the document in your browser and then extracts as usual - the password and the file both stay on your device. What no tool can do is extract from a file whose password you do not have; that is the whole purpose of the encryption.
Does the text come out in the right order for a two-column page?
Usually yes for straightforward layouts, because lines are ordered by their position on the page rather than by the order the file happens to draw them in. Genuine multi-column pages are the hard case: two columns sit side by side, so a line from the left column and a line from the right can share the same height and be read as one line. Newsletters, academic papers and magazines are worth checking over after extraction, and often extract more cleanly one column at a time.