Technical SEO checklist before launch: from staging to Search Console
Blog
Technology

Technical SEO checklist before launch: from staging to Search Console

Indexing, robots.txt, sitemap, canonical, rendering, performance and structured data. What to check on the staging site, what on launch day and what in the first few weeks.

DualFroz - VulCode CEODualFroz - VulCode CEO·29 August 2026·23 min read

A new website that doesn't show up in Google for its first month often has no problem with its content, just a setting carried over from staging: noindex in a header, Disallow: / in robots.txt, or a canonical pointing to the staging address. That's why a technical SEO checklist before launch starts with the staging site, not with keywords. Each of these mistakes can be caught with a single command before go-live, yet after launch it can cost weeks.

The whole list comes down to four questions: can the crawler fetch every important page, does it get the page's content in the HTML, does it know which address is the right one, and will it stay away from versions that shouldn't be public? The items are in order: staging first, then launch day, and finally the first few weeks in Search Console. Each one comes with a way to check it that doesn't need paid tools.

In short
  • Lock the staging site with a server-level password, not with robots.txt. Robots.txt blocks crawling, not indexing.
  • Every page has one address. All other variants (http, without www, with a trailing slash) redirect to it with a single 301.
  • The sitemap contains only canonical URLs that return a 200. Google ignores priority and changefreq in it.
  • Content and links have to be in the HTML the crawler receives, not only after JavaScript has run.
  • On launch day: Search Console verified via DNS, sitemap submitted, key URLs inspected. Indexing takes anywhere from a few days to a few weeks.

A staging site that stays out of Google

For several weeks the website lives at a test address, for example staging.example.com. All it takes is someone pasting that address into a public bug report or a forum post, and Google may find it and index it. Then two versions of the same site appear in the results, and the staging one is often full of placeholder content.

The usual reaction is to add Disallow: / to the staging robots.txt, and that doesn't work the way people think. Google states in its robots.txt documentation that this file is not a mechanism for keeping a page out of search results. A URL blocked in robots.txt can still appear in results without a description if links point to it. On top of that, a crawler that can't fetch a page won't see the noindex tag on it, so adding a robots.txt block disables exactly the block that works.

Effective protection is a password at server level. The crawler gets a 401 response, sees no content and has nothing to index. An X-Robots-Tag: noindex header is worth adding as a second layer, in case someone briefly switches the password off to show the site to a client.

staging.conf
server {
    listen 443 ssl;
    server_name staging.example.com;
    ssl_certificate     /etc/letsencrypt/live/staging.example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/staging.example.com/privkey.pem;

    auth_basic "Staging";
    auth_basic_user_file /etc/nginx/.htpasswd-staging;
    add_header X-Robots-Tag "noindex, nofollow" always;

    root /var/www/staging;
}

The word always makes nginx send the header with the 401 response too. Where the block lives also matters: in the staging server configuration, not in the site's code. The code moves to production, the test server's configuration stays where it is, so the block doesn't travel with the deployment.

Settings stored in the database are a different story. In WordPress, the option under Reading settings that asks search engines not to index the site has, since version 5.3, added a <meta name='robots' content='noindex, nofollow' /> tag to every page, as described in a WordPress core team note. If you tick it on staging and then move the database to production, you move the noindex as well. That's why the first test on launch day is run on the production domain:

check-blocks.sh · bash
curl -sI https://www.example.com/ | grep -i '^x-robots-tag'
curl -s https://www.example.com/ | grep -io '<meta[^>]*robots[^>]*>'
curl -s https://www.example.com/robots.txt

On production, the first command should return nothing. The second should return nothing, or a tag with index, follow. The third should show a file without a Disallow: / line. The pattern in the second command doesn't assume a quote style, because WordPress uses single quotes and other systems use double quotes.

One address for every page: HTTPS, www and the trailing slash

The same home page can respond at four addresses: with http and https, with www and without. Add the variant with and without a trailing slash, and on some systems one with index.html as well. To Google these are separate URLs. The search engine will pick one of them as canonical, but external links get spread across the variants, and in reports the traffic of a single page can be split into several rows.

Which version you choose as the main one is up to you. What matters is that there's only one, and that every other variant redirects to it with a single 301, not a chain from http://example.com through https://example.com to https://www.example.com.

domain.conf
server {
    listen 80;
    server_name example.com www.example.com;
    return 301 https://www.example.com$request_uri;
}

server {
    listen 443 ssl;
    server_name example.com;
    ssl_certificate     /etc/letsencrypt/live/example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/example.com/privkey.pem;
    return 301 https://www.example.com$request_uri;
}

The certificate has to cover both names, with and without www, because the browser and the crawler set up the encrypted connection first and only then receive the redirect. You can check the result with a loop that queries all four variants:

variants.sh · bash
for url in http://example.com http://www.example.com https://example.com https://www.example.com; do
  curl -s -o /dev/null -w "$url -> %{http_code} %{redirect_url}\n" "$url/"
done

A correct result looks like this: three variants return a 301 to the same target, and the fourth returns a 200 with no redirect.

output.txt
http://example.com -> 301 https://www.example.com/
http://www.example.com -> 301 https://www.example.com/
https://example.com -> 301 https://www.example.com/
https://www.example.com -> 200

When choosing a canonical URL, Google prefers HTTPS unless something contradicts it. In the same way, settle on one convention for the trailing slash and stick to it in internal links, the sitemap and canonicals.

Status codes: 200, 301 and a real 404

Type in an address that definitely doesn't exist. The response must have a 404 or 410 status, not a 200 with a "not found" message, and not a redirect to the home page. Google calls the situation where the server says everything is fine but the content says the page doesn't exist a soft 404.

404.sh · bash
curl -s -o /dev/null -w '%{http_code}\n' https://www.example.com/this-page-does-not-exist-123

The expected result is 404. This is a common mistake in apps rendered in the browser, because the server returns the same HTML file with a 200 for every address, and the "page not found" message is only drawn by JavaScript. Google's JavaScript documentation gives two ways out: redirect in JavaScript to a URL for which the server returns a 404, or add noindex to such a page. The simplest option, though, is for the server to know the list of existing routes and respond with the right code straight away.

The second thing is maintenance windows. If the site shows a maintenance message for a few minutes during a deployment, it should return a 503, not a 200. On 5xx errors, Google temporarily slows down crawling and ignores the content of the response. A page with a 200 and the text "under maintenance" is, to a crawler, the new content of your home page.

robots.txt: a file that shouldn't block anything by accident

On a new company website, robots.txt is usually short. Its job isn't to hide pages but to save the crawler's effort on URLs with no value in search results: the cart, internal search results, filters that generate thousands of parameter combinations.

robots.txt
User-agent: *
Disallow: /cart/
Disallow: /search

Sitemap: https://www.example.com/sitemap.xml

A few rules from Google's robots.txt specification that matter at launch:

  • Scope. The file applies only to the host, protocol and port it lives on. https://www.example.com/robots.txt doesn't apply to https://shop.example.com.
  • Size. Google reads the first 500 KiB of the file and ignores the rest.
  • Caching. Google usually caches the file's contents for up to 24 hours, so a fix won't take effect a minute after you upload it.
  • Status codes. When robots.txt returns a 404, Google assumes there are no restrictions. When it returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the last known version for 30 days. A file that returns a 500 because of a configuration mistake can therefore halt crawling of the entire site.
  • Rule conflicts. The rule with the longest path wins, and in a tie, the less restrictive one.
  • Page resources. Don't block the CSS and JavaScript files without which the crawler can't understand the page's layout and content.

The most dangerous launch-day mistake is a robots.txt copied from staging together with Disallow: /. That's why the file should be generated separately for each environment or, even simpler, staging shouldn't need one at all, because it's locked behind a password.

Blocking crawlers that collect data for training models is a separate decision. According to Google's crawler documentation, the Google-Extended token doesn't affect a site's inclusion in Search and isn't a ranking signal, so blocking it won't hurt visibility. Blocking the regular Googlebot in the same sweep will.

XML sitemap: only the URLs that should be in results

A sitemap is a list of the URLs you want to see in search results, not of every file on the server. Each entry should meet three conditions: it's a full URL with protocol and domain, it returns a 200, and it's canonical, meaning it doesn't redirect and has no noindex.

sitemap.xml · xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/services/websites</loc>
    <lastmod>2026-08-27</lastmod>
  </url>
</urlset>

According to Google's guide to building sitemaps, a single file can hold at most 50,000 URLs or 50 MB uncompressed, and larger sites split the sitemap into several files linked by an index. The file must be UTF-8 encoded. Google ignores the priority and changefreq values, so there's no point working on them. lastmod is used, but only when it is consistently and verifiably accurate. A generator that stamps today's date on every page at every build teaches Google that the value means nothing.

The sitemap should be generated automatically from the same data the site's routes are built from. A hand-maintained file drifts away from reality with the first new page. Before launch, check that every URL in the sitemap returns a 200:

check-sitemap.sh · bash
curl -s https://www.example.com/sitemap.xml \
  | grep -o '<loc>[^<]*</loc>' \
  | sed 's/<[^>]*>//g' \
  | while read -r url; do
      printf '%s %s\n' "$(curl -s -o /dev/null -w '%{http_code}' "$url")" "$url"
    done \
  | grep -v '^200 '

An empty result means every URL responds with a 200. Each line printed is a redirect, an error or a page that doesn't exist. Also check that the staging domain hasn't crept into the sitemap. It happens when the site's base URL comes from an environment variable set on the test server.

Canonical: a hint that one variable can break

The <link rel="canonical"> tag says which URL is the main version of a page. Every page that should be in results ought to point to itself with a full URL without parameters. Then the version with ?utm_source=newsletter points to the clean URL and doesn't compete with it in the index.

Google treats the canonical as a strong signal, not a directive. In its guide to consolidating URLs it lists three methods: redirects and canonical as strong signals, and inclusion in the sitemap as a weak one. This works well when all the signals say the same thing. When the canonical points to one URL, the sitemap to another and internal links to a third, Google chooses for itself, and it doesn't have to choose what you wanted.

Typical launch mistakes look like this:

  • The staging domain in the canonical. The base URL comes from an environment variable nobody changed. Every production page then says its main version lives on a password-protected test server.
  • One canonical across the whole site. The template puts the home page URL on every page, which tells Google that everything else is a duplicate of the home page.
  • A canonical changed by JavaScript. Google advises against setting a different URL in JavaScript from the one in the original HTML.
  • A relative URL. href="/contact" instead of a full address. Google recommends absolute URLs.
  • Noindex instead of a canonical. Google explicitly advises against using noindex to indicate which duplicate should stay, just as it advises against using robots.txt for that.

The check is one command per template type: home page, service page, blog post, contact.

canonical.sh · bash
curl -s https://www.example.com/services/websites | grep -io '<link[^>]*canonical[^>]*>'

Titles, descriptions and headings on every page

This part sits on the border between technical SEO and content, but the mistakes here are technical: a template that gives every page the same title, or a title field in the CMS that nobody filled in. In its guide to title links, Google recommends that every page has its own <title> element, without boilerplate text repeated across pages and without lists of keywords. If the title doesn't match the content, Google may show something else in results, such as the H1 heading.

The meta description isn't always displayed, but it's a candidate for the snippet under the title in results. Write one for every important page so that it answers the question someone arrives at that page with. Pages without a description get a snippet cut automatically from the content.

While you're at it, check that every page has a single H1 heading that says what it's about, and that the title is in the same language as the page content.

Rendering: is the content in the HTML before JavaScript runs

Google runs JavaScript in an up-to-date version of Chromium, but it does so in a separate step. The page is crawled first, then waits in the rendering queue, and only after rendering does its content go into the index. According to Google's documentation, the wait in the queue can be a few seconds, but it can also take longer. The same documentation calls server-side rendering or prerendering a good idea, because the site is faster for both people and crawlers. Dynamic rendering, a separate version for crawlers, is according to Google a workaround, not a long-term solution.

The simplest test needs no SEO tool. Fetch the page without a browser and search it for a sentence you can see on screen:

render.sh · bash
curl -s https://www.example.com/services/websites | grep -c 'We design websites'

A result of 0 means the sentence isn't in the HTML sent by the server and only appears after JavaScript runs. Google may still see it, but the page then depends on the rendering queue and on the whole script running without errors on the crawler's side. The page you're reading is generated as static HTML at build time.

For apps written in React, Vue or Angular, check four more things:

  • Links are <a href> elements. An element that moves to another page only through a click handler isn't a link to a crawler and doesn't lead it to discover more URLs.
  • URLs without a hash. Fragment-based routing (/#/services) makes it hard for the crawler to read URLs. Google recommends the History API.
  • Noindex in the original HTML. If the server sends noindex and JavaScript removes it, Google may never run rendering at all and stick with the noindex.
  • Content behind interaction. Google indexes the mobile version of the page and doesn't load content that requires a click, swiping a carousel or typing text. An accordion is fine as long as its content is in the HTML from the start.

The URL Inspection tool in Search Console shows how a page looks from Google's point of view. The "Test live URL" option fetches the page live and shows the rendered HTML and a screenshot, so you can see straight away whether something failed to load on the crawler's side.

Performance and the mobile version before launch

Since July 2024, Google crawls and indexes all sites with its smartphone crawler. The documentation on mobile-first indexing states plainly that the mobile version of the content is used for indexing and ranking. The mobile version therefore has to have the same content, titles, descriptions and structured data as the desktop version. Text removed from the mobile markup "for readability" doesn't exist for Google.

Performance is measured with Core Web Vitals, and Google publishes the thresholds for a good score:

MetricWhat it measuresGood score
LCPTime to render the largest element in the viewport2.5 s or less
INPTime from an interaction to the next frame on screenbelow 200 ms
CLSLayout shifts not caused by the userbelow 0.1

A new site doesn't yet have data from real users, so before launch you're left with a lab measurement in PageSpeed Insights. Measure the built production version and each template type separately, because the heaviest page is often one nobody looks at during acceptance. The usual culprits are: an above-the-fold image marked loading="lazy", images without width and height attributes, and a font that changes the width of the text once it loads.

It's worth knowing the proportions. In its description of page experience, Google says it always tries to show the most relevant content, even if the page experience is weaker, and that there is no single "page experience" signal. A fast page won't outrank a more relevant one, but when there are many similarly good results, speed helps. The websites we deliver score 95+ in PageSpeed. That's a quality standard for the site, not a promise of rankings.

Structured data: what makes sense on a company website in 2026

Structured data describes a page in a format a machine can read without guessing: a company, its logo and tax number, an article with its publication date. Google recommends the JSON-LD format and sets one condition that must not be worked around: the data has to describe content that is visible on the page where it's placed.

On a typical company website, four types make sense:

  • Organization on the home page or the "about us" page. Google doesn't require any fields in it, and among the recommended ones lists logo, address, vatID, taxID and sameAs, among others.
  • LocalBusiness instead of Organization, if the business serves customers on site. It includes the address, phone number and opening hours.
  • BreadcrumbList on subpages, to describe their place in the site structure.
  • Article on blog posts, with the publication date and author.
organization.html · html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Company",
  "legalName": "Example Company sp. z o.o.",
  "url": "https://www.example.com/",
  "logo": "https://www.example.com/logo.png",
  "vatID": "PL0000000000",
  "address": {
    "@type": "PostalAddress",
    "streetAddress": "1 Example Street",
    "postalCode": "00-001",
    "addressLocality": "Warsaw",
    "addressCountry": "PL"
  },
  "sameAs": ["https://www.linkedin.com/company/example-company"]
}
</script>

Valid structured data is a prerequisite for rich results, not a guarantee of them. One change this year matters for planning: since 7 May 2026, Google no longer shows FAQ rich results, including for government and health sites, which were the last ones still eligible. Adding FAQPage markup in the hope of expandable questions under your result no longer has any justification.

You can check validity in the Rich Results Test, which tells you whether a given type is eligible to appear in Google, and in the Schema Markup Validator, which checks conformance with the schema.org vocabulary itself.

What you don't need to do before launch

Every item below comes from Google's documentation, not from opinion:

TaskWhy you can skip it
The meta keywords tagGoogle states plainly that it doesn't use it and that it has no effect on indexing or ranking
Setting priority and changefreq in the sitemapGoogle ignores both values
Submitting the sitemap via the "ping" endpointGoogle retired this mechanism; the announcement from June 2023 said it would be switched off after six months. Search Console and the Sitemap line in robots.txt remain
An llms.txt file for searchAccording to Google's documentation from June 2026, it isn't needed in Search and doesn't affect visibility or rankings
FAQPage markup for expandable questionsFAQ rich results haven't been shown since 7 May 2026
Requesting indexing for every URL one by oneThe tool has a daily quota, and repeating the request for the same URL doesn't speed up indexing

Sources: unsupported meta tags, retirement of the ping endpoint, changes to Google's documentation and asking Google to recrawl.

Search Console on launch day

It's best to set up Search Console before launch, as a Domain property verified with a DNS record. It covers all subdomains and protocols at once, so it will also show you whether anything from staging made it into the index. DNS verification doesn't depend on the site's code, so it won't disappear with the next deployment.

On launch day itself, the order looks like this:

1
Blocks removed on the production domain

the X-Robots-Tag header, the meta robots tag and robots.txt checked with the commands from the first section.

2
Address variants

all four domain variants, each with a single 301 to the main version.

3
No traces of staging

searching the build output for the test domain, for example grep -rl "staging.example.com" dist/, returns nothing.

4
A real 404

a random URL returns a 404, not a 200 or a redirect.

5
Sitemap

every URL returns a 200, and the sitemap is submitted in the Sitemaps report.

6
Key URLs

the home page and the most important pages checked with the URL Inspection tool, with indexing requested for a few of them.

7
Performance measurement

PageSpeed Insights on production for each template type, with results saved as a baseline.

Then you have to wait. Google says that crawling can take anywhere from a few days to a few weeks, and requesting indexing doesn't guarantee that the page will appear in results immediately, or at all.

The first weeks after launch: how to read the Page indexing report

The Page indexing report in Search Console splits URLs into indexed and not indexed, and gives a reason for the latter. In the first few weeks most of the reasons are normal. The problem starts when the reason doesn't match what you intended.

Status in the reportWhat it meansWhat to do
Discovered - currently not indexedGoogle knows the URL but hasn't crawled it yetOn a new site, nothing: it's the queue. If it lasts for months, check internal links to those pages
Crawled - currently not indexedCrawled, but out of the index for nowCheck whether the content is thin or very similar to another page
Duplicate, Google chose different canonical than userYour canonical was ignoredAlign the signals: canonical, sitemap, internal links, redirects
Alternate page with proper canonical tagA variant, for example with a parameter, points to the main versionNothing, this is how it should be
Excluded by 'noindex' tagThe page excludes itselfCheck whether that's intended. On a service page it's almost always a mistake
Blocked by robots.txtThe crawler can't fetch the URLCheck that the rule doesn't cover more than it was meant to
Soft 404A page with a 200 looks empty or like an error messageReturn a 404 for non-existent URLs, or add content
Page with redirectThe URL redirects elsewhereNothing, if these are domain variants. Remove such URLs from the sitemap

The status names come from Search Console Help. Alongside the indexing report, keep an eye on the Performance report: the Queries tab shows what Google associates the site with before it starts showing it high up. If you have access to server logs, check them for 404 and 5xx responses to Googlebot.

During the same period the site gets its first real traffic and its first changes from the client's team: plugins, marketing scripts, new pages. What else to watch at this stage is covered in our post on monitoring and maintenance after launch.

Frequently asked questions about technical SEO at launch

How long does it take for a new website to appear in Google?

There's no fixed timeframe. Google says crawling can take anywhere from a few days to a few weeks, and being crawled doesn't yet mean being indexed. A submitted sitemap and links from other sites help the crawler find URLs, but faster indexing can't be bought or forced.

Do I have to submit a new website to Google?

You don't have to, because Google finds sites through links. It's worth doing anyway, because Search Console with a sitemap shows which URLs Google knows, which it hasn't indexed and why. Without it, you'll only find out about a noindex mistake from the lack of traffic.

Is robots.txt enough to hide a staging site?

No. Robots.txt blocks crawling, not indexing, so a staging URL can end up in results without a description. A staging site is locked with a server-level password.

Will a score of 100 in PageSpeed improve my rankings?

Not directly. Google primarily shows relevant content, and page experience helps when there are many similarly good results. The PageSpeed score is a lab test, while Core Web Vitals describe the experience of real users. Chasing the last few points makes sense when one of the metrics is outside its threshold, not for the number itself.

Is technical SEO done once, before launch?

Most of the work is done before launch. But every later deployment can change a template, the base URL or the server configuration, and every CMS plugin can add its own tags. It's worth running the commands from this article after every major change and reviewing the Page indexing report regularly.

How to put this into the project scope

"SEO optimisation" in a quote is a phrase that can mean everything or nothing. In our post on how to read a software house quote, we showed how to turn phrases like that into a description that can be verified at acceptance. This checklist gives you ready-made items for that description. If you're still preparing your request, add these items to your project brief so that every vendor quotes the same thing.

The websites we deliver score 95+ in PageSpeed, and with them you receive the rights to the code and documentation. The whole path from brief to launch is described on the How we work page, and you'll find the scope of our services in the offer. If your website is already live and doesn't show up in Google the way it should, write to us with its address.