Extract Links from PDF

Last updated:

About Extract Links from PDF

Extract Links from PDF is a free browser tool that lists every web address in a PDF together with the page it appears on. Drop the file and the list is ready a moment later, with no button to press. The tool reads three sources. The first is link annotations, the clickable areas a PDF uses for hyperlinks, which often hide the real address behind text such as click here or a numbered reference. The second is the visible text, where URLs are frequently written out without being clickable, as in reference lists, footnotes and documents exported from Word without hyperlinks. The third is the page image itself: scanned pages have no text layer, and the tool offers to read them with OCR so that links printed on paper are found too. Internal links that jump to another page of the same PDF, such as a table of contents, can be listed with their target page. Addresses that the layout broke across two lines are joined back together, and punctuation that ends a sentence is trimmed from the end of a URL while brackets that belong to it, as in Wikipedia addresses, are kept. Duplicates are merged into one row that lists every page the link appears on, and links can be grouped by domain to see at a glance which sites a document cites most. Each link is marked when it exists only as text, which helps authors find references they forgot to hyperlink, and email addresses can be added to the list as mailto entries. The result can be filtered, copied or downloaded as CSV for Excel and Google Sheets, or as TXT with one URL per line. The PDF is processed entirely in your browser and is never uploaded. Extract Links from PDF is commonly used as a pdf link extractor and to get all urls from a pdf. To work with the rest of the document, PDF to Text extracts the full text of the document, and Duplicate Line Remover can merge link lists collected from several files.

Long PDFs collect links quickly. A thesis or research paper can cite dozens of sources in its bibliography, an annual report links to filings and investor pages, an ebook points to tools and further reading, and a marketing whitepaper carries tracked campaign URLs. Opening each link by hand, or copying them one by one out of a PDF reader, is slow and easy to get wrong, especially when the visible text says click here and the address is hidden in the link itself. Researchers and students use a link list to check that every cited source is still online, to build a reading list or to import references into a spreadsheet. SEO and marketing teams use it to audit the outbound links in a lead magnet or press kit, to confirm that UTM parameters are correct and to see which domains a competitor references. Editors use the text only marker to find addresses that should have been hyperlinked before publishing. Grouping by domain turns a long list into a short summary, for example how many links point to doi.org, github.com or a company site, and the filter box narrows the list to one site or path. Authors of long documents can also switch on internal page jumps to check that every table of contents entry and cross reference points to the right page. For scanned contracts, printed brochures and old reports, OCR finds the addresses printed on the page, which no ordinary link extractor can see. For a document that needs further work after the links are collected, PDF Splitter can separate the chapters, PDF Page Count Checker confirms the page count, and Fix Copied PDF Text cleans up passages copied from the same file.

How to use Extract Links from PDF

  1. Drop a PDF or click to select it
  2. Remove duplicates, group by domain, add internal page jumps or scan images with OCR
  3. Copy the URLs or download them as CSV or TXT

Frequently Asked Questions

How do I extract all links from a PDF?
Drop the PDF onto this page or click to select it. The tool scans every page and shows each link with the page numbers where it appears. Use Copy URLs to put the list on your clipboard, or download it as CSV or TXT. Nothing needs to be installed and the file stays on your device.
Does it find links that are not clickable?
Yes. Besides the clickable link annotations, the tool reads the text of each page and picks up addresses that start with http, https, ftp or www. Links found only in the text are marked text only, so you can see which references in the document were never turned into hyperlinks.
What happens to a URL that is split across two lines?
PDF layouts often break long addresses after a slash or a hyphen. When a line ends inside a URL and the next line clearly continues it, the two parts are joined back into one address. A URL followed by an ordinary word on the next line is left alone, so the next sentence is not glued onto the link.
What does the CSV file contain?
With Remove duplicates on, each row holds the URL, its domain, the pages it appears on, the number of pages and whether it was found as a clickable link, as text, with OCR or a combination. With Remove duplicates off, there is one row per page on which a link appears. The file is UTF-8 and opens directly in Excel, Google Sheets and Numbers.
Can it find links in a scanned PDF?
Yes. Scanned pages are images without a text layer, so the tool detects them and shows a Scan with OCR button. OCR reads the printed text of those pages in your browser and any URLs, www addresses and emails it finds are added to the list with an OCR label. OCR can misread a character, so check those links before using them. For a PDF that has text but also images with printed addresses, use Also look for links printed inside images.
Can it list internal links to other pages of the PDF?
Yes. Switch on Include internal page jumps and links that point to another page of the same document, such as entries in a table of contents or cross references, are listed as Jump to page N together with the page they are on. Both direct page targets and named destinations are resolved. They are off by default so that the list contains only web addresses.
Is my PDF uploaded to a server?
No. The PDF is opened with PDF.js inside your browser and all link detection runs locally. ToolBox never receives or stores the file, so the tool is safe for confidential reports, contracts and unpublished papers.

Related Tools