> pdf to-text
Convert PDF to text
Get the text out of a PDF to copy it, search it or keep working with it.
Choose a PDF
or just drag it here.
.PDF ONLY · STAYS ON YOUR DEVICE
The password your PDF program asks for when opening the file. It stays in this browser: It never appears in the address bar, is not saved and is not sent. As soon as the file is open, the field is cleared again.
Single pages with a comma, ranges with a hyphen. “4-” means: from page 4 to the end. In the result, the pages are always in the order of the file.
You changed a reading setting. The result below still belongs to the old values. Tap “Extract again”.
How the text should look
Related: Count and clean up text · Compare PDFs
Your file and your password do not leave this device. The text is created in your browser and is not stored anywhere.
How it works
- 01Choose a PDF or drag it into the box. If the file asks for a password, type it in.
- 02Have all pages read or only certain ones, for example 1-3,5. Then tap “Extract text”.
- 03Next to the result, you set how the text should look. The result changes right away, the file is not read again for that.
- 04Copy it or save it as a TXT file.
Where the text comes from. A PDF does not store flowing text, but single pieces of text with coordinates on the page. This tool puts them back together: Pieces at the same height become one line (how much height difference is allowed depends on the font size), sorted from left to right within the line. A space appears where the gap is larger than a quarter of the font size, not between every piece. A larger space above, a short line before, a list item or a jump in font size starts a new paragraph. Ligatures like “fi” become “fi”, invisible characters are dropped.
Where a blank line goes. The tool only puts a blank line where there was a larger space in the PDF too. Otherwise the paragraph starts on the next line. Where it has to guess, it can be wrong: If all lines are the same length (an address, a poem, a list), they look like a single paragraph, and “Remove line breaks within paragraphs” glues them together. Then turn the setting off.
Columns. The tool detects two columns by a vertical gap in the middle of the page that almost no line runs through. Then it reads the left column first, then the right one. A heading across both columns and a table in the middle of the page stay in their place. If a paragraph runs from the left into the right column, the tool joins it when the line reaches the edge and ends without punctuation. This is an approximation. Not detected are three or more columns, text that flows around an image, pages with very few lines, and pages where the columns only take up part of the height with full width text above or below them. There the text is read across, and the lines of both columns get mixed. A narrow right column, such as amounts in the margin, does not count as a column. On which pages it assumed two columns is shown below the button after extracting. With “Detect multiple columns” you turn the search off.
Tables. Tables become text, not tables. A table row comes out as one line, the cells separated by a space. If a line with a large gap sits below a second one, the tool treats both as table rows and does not attach them to the text before. A single such line can also be justified text stretched over two short words, so it stays in the paragraph. If a cell contains text over several lines, the lines of a table row get mixed. For real tables, the result is not useful.
Hyphenation. If a line ends with a hyphen and the next line starts with a lowercase letter, the tool joins the word. If it starts with an uppercase letter (“Wi-” and “Fi”), the hyphen stays and the space goes. With “pre-” and “and post-war” (also with “or”), both stay. A real hyphen that is followed by a lowercase letter and sits exactly at the end of the line looks like hyphenation and is joined as well.
Across the page break. Without page separators, the tool joins a paragraph that runs over the end of the page if the continuation is clear: a hyphen at the end of the page, or the line reaches the edge and stops without punctuation. If the new page starts with a header, the next line must also continue in lowercase. Otherwise there is a blank line between the pages. With page separators, such a paragraph stays in two parts.
Headers and footers. The tool looks at the top and bottom three lines of each page and looks for lines that are the same on at least three pages and on at least 60 percent of the pages read. Digits count as the same: “Page 3 of 12” and “Page 4 of 12” are the same line. If a header on column pages comes in two parts (the date on the left, the title in the middle), both parts count as one line. They are only removed if you turn the setting on, and it first names what was found. A line that happens to be at the top or bottom of many pages is found too: Read the examples before you remove anything.
Pages without text. If a page has no text at all, the tool says so and names the page. If all pages read are empty, the PDF was probably scanned, and there is nothing to output.
Unreadable characters. Some PDFs contain fonts without a mapping to real letters. Copying then gives gibberish, in every program. The tool counts characters that are clearly not letters (replacement characters, characters from the private use area, non-printable characters) and warns from 5 percent of all characters, but only from at least three such characters: Otherwise a single broken character in a very short text would already raise an alarm. Characters that are wrong but look like normal letters are not detected.
Text in another direction. If part of the text runs at an angle or across (a watermark like “DRAFT”, a margin note), it is missing from the result. The notice names the pages and the number of characters. If all of the text on a page runs across, the tool reads it in reading direction.
Password and copy lock. If the PDF asks for a password to open, you need it: There is no guessing or cracking here. If only copying is locked, the tool reads the text anyway and tells you so. Only use this for files that belong to you or whose text you are allowed to use.
The file. The TXT file is UTF-8 and ends with a line break. Lines are separated by a simple line feed; very old programs then show everything on one line. With the byte order mark (BOM), Windows Notepad also reliably detects UTF-8. For a selection of pages the file is called “name-extract.txt”, otherwise like the PDF. Words and characters are counted as in Count text: characters including spaces and line breaks, words separated by white space.