Technical SEO Checklist for Business Websites
A practical technical SEO checklist for business websites: crawling, indexing, canonicals, sitemaps, robots.txt, Core Web Vitals, structured data, redirects and JavaScript rendering.
Technical SEO is the part of search optimisation that most business owners never see, and the part that can quietly undo everything else. Excellent content does nothing if Google cannot crawl it, indexes the wrong version, or finds the page too slow and unstable to recommend. The good news is that most technical problems on business websites are common, well documented and fixable.
This checklist is organised in the order we work through a site: can search engines find the pages, can they index the right ones, do the pages perform well, and is the information clearly structured. For authoritative detail, Google Search Central documentation is the reference we rely on, and web.dev explains Core Web Vitals in depth.
Before you start: tools you need
- Google Search Console, verified as a domain property. The Page indexing report, URL Inspection tool, Sitemaps report and Core Web Vitals report are essential.
- A crawler that visits your site like a search engine and lists status codes, titles, canonicals and links.
- PageSpeed Insights for both field data from real users and lab diagnostics.
- Google's Rich Results Test for structured data.
- Access to the site's code or CMS, and ideally a staging environment for testing changes.
1. Crawling
Crawling is how search engines discover pages by following links and reading sitemaps.
- Check robots.txt at yoursite.com/robots.txt. Make sure it does not block important sections, CSS or JavaScript files needed to render pages. A leftover "Disallow: /" from a staging site is a surprisingly common disaster.
- Remember robots.txt controls crawling, not indexing. A blocked page can still appear in results if other sites link to it. Use noindex to keep a page out of the index, and do not block that page in robots.txt, or Google cannot see the noindex.
- Make sure every important page is reachable through normal HTML links, not only through search boxes, forms or JavaScript click events.
- Look for crawl traps: calendars, filters and sorting parameters that generate endless URL combinations.
- Fix internal links pointing to redirects or broken pages so crawlers and users reach the final URL directly.
2. Indexing
Indexing is whether a page is stored and eligible to appear in results.
- Review the Page indexing report in Search Console. Look at why pages are not indexed, and whether any important pages are among them.
- Check for accidental noindex tags or X-Robots-Tag headers on pages that should rank. These often slip in from CMS settings or plugins.
- Identify thin or duplicate pages, such as tag archives, empty category pages and printer versions, and decide whether to improve, consolidate, noindex or remove them.
- Use URL Inspection on key pages to see the version Google has indexed and when it last crawled it.
- Accept that Google does not index every page it finds. Focus on making important pages clearly valuable rather than forcing indexing.
3. Canonical tags and duplicate content
When several URLs show the same or very similar content, a canonical tag tells search engines which version you prefer.
- Each indexable page should have a self-referencing canonical tag with the full, absolute URL.
- Pick one version of the domain: https, with or without www, and with or without trailing slashes. Redirect other versions to it and use it consistently in canonicals, sitemaps and internal links.
- URLs with tracking or filter parameters should canonicalise to the clean version where the content is essentially the same.
- Do not canonicalise paginated pages to page one; each page in a series contains different items.
- Check in URL Inspection whether Google selected the same canonical you declared. If not, look for conflicting signals.
4. XML sitemaps
- Include only canonical, indexable URLs that return a 200 status. No redirects, no noindexed pages, no errors.
- Generate the sitemap automatically so it updates when pages are added or removed.
- Use accurate lastmod dates, updated only when content meaningfully changes.
- Split large sites into multiple sitemaps by content type, with a sitemap index file.
- Reference the sitemap in robots.txt and submit it in Search Console.
5. Redirects and status codes
- Use 301 or 308 redirects for permanent moves and 302 or 307 for genuinely temporary ones.
- Remove redirect chains: A to B to C should become A to C.
- Redirect removed pages to the closest relevant equivalent, not all to the homepage. If there is no equivalent, a 404 or 410 is acceptable.
- Watch for soft 404s: pages that say "not found" but return a 200 status.
- Keep a redirect map whenever URLs change during a redesign. Our guide on redesigning a website without losing SEO covers this process.
6. Core Web Vitals and page experience
Core Web Vitals measure real-user experience. Google's published "good" thresholds, measured at the 75th percentile of page loads, are:
| Metric | What it measures | Good threshold | Common fixes |
|---|---|---|---|
| LCP (Largest Contentful Paint) | How quickly the main content appears | 2.5 seconds or less | Faster server response, optimised hero images, preload key resources, remove render-blocking CSS and JS |
| INP (Interaction to Next Paint) | How quickly the page responds to taps, clicks and key presses | 200 milliseconds or less | Reduce heavy JavaScript, break up long tasks, limit third-party scripts |
| CLS (Cumulative Layout Shift) | How much the layout jumps while loading | 0.1 or less | Set image and video dimensions, reserve space for ads and embeds, careful font loading |
INP replaced First Input Delay as a Core Web Vital, so older audit reports mentioning FID are out of date. Prioritise field data from Search Console and PageSpeed Insights over lab scores, because field data reflects your actual visitors, many of whom may be on mid-range phones and mobile networks.
Also check the basics of page experience: HTTPS across the whole site, no intrusive pop-ups covering content on mobile, and a responsive layout that works on small screens.
7. Structured data
- Use JSON-LD format, as Google recommends.
- Add types that match the page: Organization or LocalBusiness on the relevant page, Article for blog posts, Product for products, BreadcrumbList for navigation, and others as relevant.
- Only mark up content that is visible on the page. Marking up hidden or false information is against Google's guidelines.
- Validate with the Rich Results Test and monitor the enhancement reports in Search Console.
- Remember structured data makes pages eligible for rich results; it does not guarantee them.
8. JavaScript rendering
Google can render JavaScript, but rendering can be delayed, and other search engines and AI crawlers may handle it less well.
- Use URL Inspection's rendered HTML view to confirm important content, links and metadata appear after rendering.
- Prefer server-side rendering or static generation for important content, especially on single-page applications.
- Use real anchor links with href attributes for navigation.
- Make sure titles, meta descriptions and canonicals are not set only by JavaScript that may fail.
- Avoid loading key content only after user interaction, such as clicking a tab that fetches content.
9. On-page technical hygiene
- Unique, descriptive title tags and meta descriptions on each important page.
- One clear H1 per page and a logical heading structure.
- Descriptive alt text on meaningful images, and compressed images in modern formats.
- Correct hreflang tags if you have language or country versions.
- A clear internal linking structure that puts important pages within a few clicks of the homepage.
Illustrative example
Illustrative example. A services company relaunches its website and sees enquiries from search drop. A technical review finds the staging robots.txt was copied to the live site, the new URLs have no redirects from the old ones, and the hero image on every page is a large uncompressed file causing slow LCP. Fixing robots.txt, adding a redirect map from old to new URLs and optimising images addresses the root causes. Search Console is then used to monitor recrawling and indexing over the following weeks.
Key takeaways
- Start with crawling and indexing; nothing else matters if important pages are not indexed.
- Use robots.txt to control crawling and noindex to control indexing, and do not mix them up.
- Keep canonicals, sitemaps, redirects and internal links consistent with one preferred URL format.
- Meet Core Web Vitals thresholds for LCP, INP and CLS based on real-user field data.
- Add valid structured data that reflects visible content, and make sure key content does not depend on client-side JavaScript alone.
If you would rather have an expert work through this for you, our technical SEO service covers audits and implementation, and our web development team can fix performance and rendering issues at the code level.
Frequently asked questions
How often should I run a technical SEO audit?
What is the difference between robots.txt and noindex?
Do Core Web Vitals directly affect rankings?
Is a JavaScript framework bad for SEO?
Worried something technical is holding your site back? Ask us for a technical SEO review.
Book a free consultation — we'll look at your situation and suggest the next practical step.