PDF text extractor
Get selectable text out of a PDF, one page or a whole document at a time.
Your PDF stays on your device. Embedded text only, not OCR.Choose or drop one PDF
25 MiB · 300 source pages · 5,000,000 extracted Unicode characters
Embedded text only. No OCR, uploads or page rendering.
How to extract text from a PDF
- Choose one unencrypted PDF. Check the source page count.
- Use all pages or enter a range such as 1-3, 6, then choose Extract text.
- Review each page. Optional cleanup changes the output, not the original extraction.
- Copy a page or the combined text, or edit the filename and download the complete UTF-8 TXT file.
A simple example
A two-page PDF containing “Hello” on the first page and “World” on the second exports the following with the default page markers:
--- Page 1 --- Hello --- Page 2 --- World
With markers off, it becomes Hello, two LF newlines, then World. Selected pages retain their original physical numbers, so page 6 is labelled --- Page 6 ---.
Check the result, especially complex layouts
This is text extraction, not a faithful document conversion, table reconstruction, translation or accessibility repair. It does not join broken hyphenated words, invent missing characters, activate PDF links or scripts, or open attachments. Columns, rotated text and right-to-left layouts can be imperfect. Retain your original PDF and verify important information.
The preview displays at most 20,000 UTF-16 units per page; exports retain complete supported text. Clipboard operations are capped at 1,000,000 UTF-16 units. If access is blocked, select the preview manually or download the TXT file. Clear offers an in-memory undo of the file and settings; reload or closing the page discards them.
Frequently asked questions
Is my PDF uploaded?
No. Parsing, text extraction and cleanup happen in browser workers on your device. Files, filenames and extracted text are not sent to analytics, stored or uploaded.
Can this read scanned PDFs?
Not without an existing text layer. This tool does not perform OCR. A scan with a hidden text layer may work; an image-only page has no extractable text. A textless result does not prove a page is blank or scanned.
Will the text match the page layout?
Not always. A conservative content-stream and geometry heuristic separates words, lines and rotated blocks. Columns, tables, ligatures, mixed directions and unusual fonts may have imperfect reading order or spacing. Missing character mappings cannot be reconstructed. Check important text against the original.
Which PDFs are supported?
One unencrypted PDF up to 25 MiB and 300 source pages. Selected pages may contain at most 5,000,000 extracted Unicode code points. Signed PDFs and form documents can be read without modifying them, but annotation and form-field values are not extracted. Encrypted and password-protected files are rejected; this tool never removes protection.
What happens when a page fails?
Successful and textless pages remain available. Failed pages are identified separately, and partial exports require explicit acknowledgement. The combined file starts with a PARTIAL EXTRACTION status and the failed physical page numbers. A cancelled or over-limit run does not export a truncated result.
How are pages and cleanup handled?
Page numbers are physical positions counted from 1, not printed labels. Duplicates are removed and source order is preserved. Page markers are on by default. Without markers, pages are joined with two LF newlines. Trailing-space trimming and repeated-blank-line cleanup are optional and off by default; switching them off restores the original extraction.
How was this tool?
Help us improve PDF text extractor.
Opens your email app; you send it there. No email app? Copy your comment and email feedback@tooltulip.com.
Don't include sensitive data.
0/300