Guide

How to check a website for broken links and forms

Broken links and forms that fail without a sound are the defects visitors find first and report last. This is how to find them yourself, on the live site, with a record of what you saw.

Last reviewed: 29 September 2026

Run these checks against the site on its real domain. A link checker pointed at a developer machine or a staging host finds different problems from the ones your client's visitors meet, because redirects, trailing slashes and form delivery are all decided by the production setup.

1. Build the list of URLs

Start from three sources, not one

A crawl from the home page only finds pages that are linked. Pages reached from the sitemap, from old campaigns or from other sites are missed. Put three lists together: the URLs in the sitemap, the URLs a crawl from the home page finds, and the old URLs that still get traffic or backlinks.

HowExtract the sitemap URLs with one command, then add the crawl export and the analytics export of the most visited landing pages. Remove duplicates.

curl -s https://example.org/sitemap.xml | grep -oE "<loc>[^<]+" | sed "s/<loc>//" > urls.txt

Include the links nobody clicks during review

Footer links, legal pages, the cookie banner, the 404 page, confirmation pages and email templates carry links too, and they are the ones least often opened by hand. Make sure your crawl reaches them.

2. Check every status code

Internal links answer 200, directly

An internal link should reach its page in one request. A 404 or 500 is a broken link. A 301 or 308 is a detour: it still works, but it costs a request, and a chain of them eventually breaks. The usual cause is a trailing slash mismatch, a link to /about on a site that serves /about/.

HowRequest each URL on the list and write down the status and, for redirects, the destination. Anything other than 200 on an internal link goes on the fix list.

while read -r u; do
  printf "%s %s\n" "$u" "$(curl -s -o /dev/null -w "%{http_code} %{redirect_url}" "$u")"
done < urls.txt

External links still point at something useful

Links to partners, suppliers, social profiles and documents on other sites break without anyone on your side changing anything. Check that each one answers 200 and that the page is still the one the link text promises: a domain that has lapsed often answers 200 with a parking page.

HowSome sites refuse automated requests and answer 403 or 429 to a script while working fine in a browser. Open those few by hand before calling them broken.

3. The sitemap lists only live pages

Every URL in the sitemap should answer 200, be allowed by robots.txt and not be marked noindex. A sitemap is often generated from a list someone wrote once, and pages that were removed or never built stay in it.

From a Sightline run on our own site

On 29 August 2026, and again on 2 September, a Sightline run on seekingdatalabs.com filed this finding against https://seekingdatalabs.com/audit/:

The site lists a URL (in its sitemap or internal links) that returns an error. Remove or fix the dead URL so crawlers and visitors are not sent to it.

There was no page at /audit/; our sitemap listed it anyway. We removed it from the sitemap and added a test that holds every sitemap entry against the real pages, robots.txt and the redirect rules. On 29 September 2026 /audit/ still answers 404, as it should, and the live sitemap no longer lists it.

4. Each link goes where its text says

A link that answers 200 can still be wrong. If the text says "Pricing" and the link opens the contact page, a checker reports it as healthy and a visitor reports it as broken. This is a common result of late copy changes: the text is updated, the link is not.

HowExport every link with its visible text and its destination, sort by text, and read the list. Two links with the same text and different destinations, or a destination that does not match its label, stand out at once.

From a Sightline run on our own site

On 5 September 2026 a Sightline run on seekingdatalabs.com filed this against our home page:

The page claims 'Every report ships over the same envelope, delivered by the Reports API' with an arrow link to /app/billing. The Reports API is described as a delivery mechanism, but the link destination is the billing page, not documentation or a description of the API.

The link worked; it just went to the wrong place, and the footer link with the same label went somewhere else. On 29 September 2026 both "Reports API" links on the home page point to /docs/.

5. Anchors, mail and phone links

Every in-page anchor has a target

A link to #contact or /#pricing does nothing when the section was renamed or removed, and a crawler reports it as fine because the page itself loads.

HowFor each href that contains #, open the target page and confirm an element with that id exists. In the browser console, document.getElementById("contact") returns null when it does not.

Email and phone links are written correctly

Check that mailto: links contain one valid address with no spaces, and that tel: links use the international format with no letters or brackets, for example tel:+41441234567. Tap each one on a phone.

Links that open a new tab are marked

Links with target="_blank" should say so in their text or with an icon that has a text alternative, and should carry rel="noopener" so the new page cannot control the old one.

6. Redirects and the 404 page

Old URLs land on their new equivalent in one hop

When a site replaces an older one, each old URL with traffic or backlinks should answer 301 to the matching new page. Sending everything to the home page loses the visitor and the search ranking of the old page.

HowFollow each old URL to the end and count the hops. More than one redirect, or a final page that is not the equivalent content, goes on the fix list.

curl -sIL https://example.org/old-page | grep -iE "^(HTTP|location)"

A missing page answers 404

The error page should look like the rest of the site and offer a way on, and its status must be 404. A friendly error page that answers 200 gets indexed as a page of its own.

curl -s -o /dev/null -w "%{http_code}\n" https://example.org/this-page-does-not-exist

7. Images, files and scripts

A page can load with a missing image, a stylesheet that returns 404 or a script blocked by the browser, and still look almost right. Downloads such as PDFs and price lists break in the same quiet way.

  • Open the network panel in the developer tools, reload each template (home, a content page, a listing, a form page) and filter by status. Anything red is a broken asset.
  • Read the console. Blocked scripts, mixed content (a file requested over plain HTTP on an HTTPS page) and failed font loads are reported there and nowhere else.
  • Request every linked file (PDF, DOCX, ZIP) like any other URL and check that it is the current version, not last year's.

8. Forms, end to end

A form is working when the visitor sees a clear result and the right person receives what they typed. A checker cannot prove either, so each form needs one real submission.

List every form

Contact, quote request, newsletter, search, login, booking, and the forms inside pop-ups. Note for each one where its submissions should go.

The form posts to the production endpoint

HowView the page source and read the action of each <form>, or submit once with the network panel open and read the request URL. It should be on HTTPS and on the production host, not a staging URL or a test account at a form service.

A valid submission arrives, complete

Submit once with realistic data, including an accented name and a long message. Confirm that the message arrives in the client's inbox, not a developer's test address, that nothing was cut off or garbled, and that replying goes to the visitor's address.

Invalid input is handled in the open

Submit the form empty, then with a malformed email address. The error should name the problem, sit next to the field, and keep everything else the visitor typed. Then submit valid data twice in quick succession: the second click should not send a duplicate.

Spam protection does not block people

A challenge widget or a hidden honeypot field should stop scripts, not visitors. Submit once from a private window on a phone connection and once with the keyboard only. If either fails, real visitors fail the same way.

Failures are visible to someone

If the mail service rejects a message, the visitor should see an error rather than a thank-you page, and someone on your side should be told. Ask how the form reports a failed delivery, and test it if you can.

Fields are labelled and fill in correctly

Click each label: the cursor should move into its field. Check that email, phone and address fields carry the right type and autocomplete attributes, so phones show the right keyboard and browsers offer saved details.

9. Re-check after launch

Links break after launch, not only before it. Partner sites move, documents are replaced, and one edit to a menu can orphan a section. Run the status check on the whole URL list after every deploy that touches navigation or content, and on a schedule in between. Keep each run's output with its date: a list of URLs and status codes from last month is what tells you whether a broken link is new or was never right.