Last updated:
Extract Links from PDF is a free browser tool that lists every web address in a PDF together with the page it appears on. Drop the file and the list is ready a moment later, with no button to press. The tool reads three sources. The first is link annotations, the clickable areas a PDF uses for hyperlinks, which often hide the real address behind text such as click here or a numbered reference. The second is the visible text, where URLs are frequently written out without being clickable, as in reference lists, footnotes and documents exported from Word without hyperlinks. The third is the page image itself: scanned pages have no text layer, and the tool offers to read them with OCR so that links printed on paper are found too. Internal links that jump to another page of the same PDF, such as a table of contents, can be listed with their target page. Addresses that the layout broke across two lines are joined back together, and punctuation that ends a sentence is trimmed from the end of a URL while brackets that belong to it, as in Wikipedia addresses, are kept. Duplicates are merged into one row that lists every page the link appears on, and links can be grouped by domain to see at a glance which sites a document cites most. Each link is marked when it exists only as text, which helps authors find references they forgot to hyperlink, and email addresses can be added to the list as mailto entries. The result can be filtered, copied or downloaded as CSV for Excel and Google Sheets, or as TXT with one URL per line. The PDF is processed entirely in your browser and is never uploaded. Extract Links from PDF is commonly used as a pdf link extractor and to get all urls from a pdf. To work with the rest of the document, PDF to Text extracts the full text of the document, and Duplicate Line Remover can merge link lists collected from several files.
Long PDFs collect links quickly. A thesis or research paper can cite dozens of sources in its bibliography, an annual report links to filings and investor pages, an ebook points to tools and further reading, and a marketing whitepaper carries tracked campaign URLs. Opening each link by hand, or copying them one by one out of a PDF reader, is slow and easy to get wrong, especially when the visible text says click here and the address is hidden in the link itself. Researchers and students use a link list to check that every cited source is still online, to build a reading list or to import references into a spreadsheet. SEO and marketing teams use it to audit the outbound links in a lead magnet or press kit, to confirm that UTM parameters are correct and to see which domains a competitor references. Editors use the text only marker to find addresses that should have been hyperlinked before publishing. Grouping by domain turns a long list into a short summary, for example how many links point to doi.org, github.com or a company site, and the filter box narrows the list to one site or path. Authors of long documents can also switch on internal page jumps to check that every table of contents entry and cross reference points to the right page. For scanned contracts, printed brochures and old reports, OCR finds the addresses printed on the page, which no ordinary link extractor can see. For a document that needs further work after the links are collected, PDF Splitter can separate the chapters, PDF Page Count Checker confirms the page count, and Fix Copied PDF Text cleans up passages copied from the same file.