Website maintenance after launch: monitoring, fixes and upkeep for sites and apps
Blog
Technology

Website maintenance after launch: monitoring, fixes and upkeep for sites and apps

What only shows up under real traffic, what is worth monitoring, how to report bugs, when to deploy fixes and what the documentation needs so another team can take the project over.

DualFroz - VulCode CEODualFroz - VulCode CEO·10 August 2026·22 min read

A website that passed every test on staging will still show things nobody has seen before in its first week in production. Not because the tests were bad, but because no test replaces real users, real mailboxes and real search traffic. The question in website maintenance after launch is not whether something will come up, but who notices it first: you, your customer or a monitoring system.

This article covers the post-launch period from the practical side: what to watch, how to set up alerts so that someone actually reads them, how to report bugs so a fix does not take three days, when to deploy changes and how to protect your data. At the end we compare the options for when free monitoring ends, including the ones where our maintenance plan is not a good choice.

In short
  • After launch our projects get free monitoring, usually for 90 days, and the project price includes a buffer of 15% of the project time for fixes.
  • External monitoring checks availability and response time. The certificate, domain, backups and scheduled jobs need separate checks.
  • An alert without an assigned action and recipient is noise. Every rule should say what to do when it fires.
  • A backup nobody has tried to restore is an assumption, not a backup.
  • Full rights to the code only make sense if the documentation lets another team run and deploy the project without asking the authors.

What only shows up in production

The first group is real data. On staging, forms get filled in with data like "John Smith, 1 Test Street". Real users have surnames with apostrophes, addresses longer than the field allows and phone numbers with a country code and spaces, and they paste text from an editor full of invisible formatting characters into the message field. Validation that let through everything it should during testing starts rejecting valid data, or letting through data that breaks the export to the accounting system.

The second group is email. Contact form notifications sent from a server without correctly configured SPF, DKIM and DMARC records for the domain land in spam or get rejected by the recipient's servers. In testing the messages arrived, because the test address was on the same domain or in a mailbox that does not filter. After launch the client spends a week thinking nobody is getting in touch, while the messages sit in a salesperson's spam folder.

The third group is settings left over from staging. A noindex tag or Disallow: / in robots.txt, added so that search engines would not index staging, carried over to production, blocks indexing of the new site. Test keys for the payment gateway accept orders nobody pays for. An API address pointing at the test backend works until someone clears the test database.

The fourth group appears when an old website is replaced with a new one. The old page URLs are indexed by search engines, linked from other sites and saved in bookmarks. If the new site has a different URL structure and nobody has prepared 301 redirects from the old addresses to the new ones, search traffic lands on a 404 error page. The Pages report in Google Search Console shows such URLs as 404 errors once the crawler hits them, and it is one of the first places worth checking after launch.

➜
Tip

Before switching the domain, check three things in production: that the site has no noindex tag and no block in robots.txt, that the form delivers a message to an external mailbox, for example in Gmail, and that the payment gateway runs in live mode. Each of these checks takes a minute.

What monitoring covers in the first 90 days

With us, free post-launch monitoring usually runs for 90 days. Its foundation is availability and response time. Both signals are best checked from the outside, because such a test has an advantage over monitoring on the server itself: a server that has lost its network connection will not send a message saying it has lost its connection. A test from another location sees the same thing the user sees.

Availability and response time are not everything that can break without any change to the code, though. The table below collects the signals worth watching in a typical website or app, with how to check them and the threshold at which an alert should go off. Which of them make sense in a given project depends on what the project does, so the scope is best written down rather than assumed.

SignalHow to checkWhen to alert
AvailabilityAn external HTTP request every minute, from more than one locationSeveral failed attempts in a row from different locations
Response timeThe same test, recording the time to first byteAn increase that persists for a dozen or so minutes, not a single spike
TLS certificateExpiry date read from the certificateTwo weeks before expiry, if automatic renewal has not worked
DomainExpiry date in the domain registryA month before, if auto-renewal is not set up
Application errorsShare of 5xx responses in the logs, JavaScript errors collected from browsersA new type of error or a clear rise in the number of known errors
Disk spaceServer metricsWhen, at the current rate, the disk will fill up within a few days
BackupsJob status and backup file sizeNo successful backup in the expected window, or a suspiciously small file
Scheduled jobsAn "I'm alive" signal sent at the end of every runNo signal within the expected time

Two rows of this table are particularly treacherous. A scheduled job that has stopped running generates no error at all: simply nothing happens. Reminders do not go out, invoices are not generated, old temporary files are not cleaned up. That is why such jobs are monitored the other way round from everything else, with an alert on a missing signal rather than on an error.

The second treacherous row is backups. A backup job can finish successfully and create an empty file, because, for example, the database password has changed and the script does not check the exit code of the dump tool. A file size that suddenly drops from several hundred megabytes to a few kilobytes is a simple and effective sign that something is wrong.

An alert someone will read

Monitoring can stop working without any technical failure. Alerts arrive so often, and so often turn out to be false, that people stop reading them, and then a real alert gets lost among the rest. Well-configured monitoring speaks up rarely, and every time for a reason that needs a response.

The first tool against false alarms is a duration condition. A single failed request can come from a momentary network problem between the probe and the server. An alert should fire only when the problem persists for several minutes and, for availability, ideally when tests from different locations confirm it. Below is an example of Prometheus rules using the blackbox exporter, which performs HTTP requests and reads certificates:

site-alerts.yml · yaml
groups:
  - name: site
    rules:
      - alert: SiteDown
        expr: probe_success{job="blackbox-http"} == 0
        for: 3m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.instance }} has not responded for 3 minutes"
          action: "Check the container status and the latest deploy, roll back the version if needed"

      - alert: CertificateExpiring
        expr: probe_ssl_earliest_cert_expiry{job="blackbox-http"} - time() < 14 * 86400
        for: 1h
        labels:
          severity: warning
        annotations:
          summary: "The certificate for {{ $labels.instance }} expires in less than 14 days"
          action: "Check the automatic certificate renewal logs"

The for field is exactly that duration condition: the rule has to be met continuously for the given time before the alert moves from pending to firing. The severity label lets you route alerts in different ways: critical ones to a channel with phone notifications, warnings to a channel someone reviews once a day. The action annotation is not required by Prometheus, but it is the most important line in the rule, because it tells the person woken up by the notification where to start.

The second tool is the rule that every alert has a recipient and an action. If nobody knows what to do when a rule fires, or the only response is "I'll check later", the rule should be removed or turned into an ordinary chart. Metrics worth looking at that do not require an immediate response are better kept on a dashboard, for example in Grafana, than in the alert channel.

The third is a review after every false alarm. An alert that fired for no reason is information that the threshold is set wrong. Fixing it straight away costs a few minutes, while leaving it costs trust in the whole system.

A buffer of 15% of project time for fixes

The price of every project we do includes a buffer for fixes equal to 15% of the project time. On a project planned at 100 hours of work that is 15 hours, and on a 400-hour project it is 60 hours. A buffer calculated from project time rather than as a fixed number of days grows with the project: a panel with a dozen or so views has more places where real use will reveal the need for a fix than a one-page landing page does.

The line between a fix and a new feature is where misunderstandings after launch arise most easily, so it is worth drawing it clearly before the first request comes in. A fix is a change that brings the project to the state described in what was agreed, or removes a problem that surfaced in real use. A new feature is something that was not in the agreement.

Fits within the fix bufferThis is new scope
A bug in a feature that was in the project scopeA new feature that was not in the agreement
Validation rejecting valid data or letting invalid data throughA new integration with an external system
Reordering fields, labels and messages after the first days of useReworking the layout of sections or a new graphic design
Display breaking on a specific device or browserSupport for a new type of user with their own permissions
Small changes to text and imagesNew subpages, a new language version

The second column does not mean such changes are frowned upon. After a few weeks of using a new website or panel, ideas naturally come up for things nobody foresaw at the planning stage. Such changes are quoted separately, like any other scope, so that the fix buffer stays available for what it is meant for.

The changes in the third row of the first column show well why the buffer exists. A form that looked logical in the mockup turns out in daily use to have its fields in an awkward order, and an error message that makes sense to a developer means nothing to a receptionist. You cannot catch such things in testing, because they only come out after the hundredth time the form is filled in by someone who does it for a living.

How to report a bug so the fix takes minutes, not days

A report saying "the form doesn't work" sets off a series of questions: which form, on what device, what exactly happened, when. Every question and answer is another exchange of messages, and every exchange means hours of delay. A report that contains the answers up front lets us start by looking for the cause, not by working out what to look for.

bug-report.txt
Page address:        https://yourdomain.com/contact
What I did:          filled in the form, attached a PDF (4 MB), clicked Send
What happened:       the button spins for about 30 seconds, then the message "An error occurred"
What I expected:     confirmation that the message was sent
Device:              iPhone, Safari
When:                today, around 14:20
Does it repeat:      yes, three times in a row; works without the attachment
Attachments:         screenshot of the message

The most valuable field is the time. The server logs every request with an exact timestamp, so "around 14:20" makes it possible to find the specific error in the logs within minutes instead of going through the whole day. The "does it repeat" field separates bugs that can be reproduced and fixed straight away from one-off problems, for example with the user's connection.

In the example above, one detail almost points to the cause: without the attachment the form works, with a 4 MB file it does not. It may be a request size limit on the server, an upload timeout or file type validation. Without that information the search would start by checking email delivery, which is not the problem at all.

A screenshot is good, and a short screen recording is better, because it shows the order of steps, which users often do not remember. Both phone operating systems have built-in screen recording, so there is nothing to install.

Deploying a fix: in the middle of the day and in small steps

We deploy regularly, and preferably in the middle of the day. The reason is simple: if something goes wrong after a deploy, in the middle of the day everyone is at their computer, including the people on the client side who can check the change live. A Friday evening deploy means that a problem discovered on Saturday morning waits for whoever happens to pick up the phone.

A fix goes first to the staging environment, which the client has access to. There it can be checked on the same data on which the bug occurred, before the change touches production. With fixes reported by the client, the client knows best whether the problem has gone, so their confirmation on staging is a better test than any automated one.

Small deploys are safer than big ones for a simple reason: when something breaks after deploying one change, you know which change did it. When you deploy twenty changes collected over a month, the search for the cause starts with working out which of the twenty it was. A small deploy is also easier to roll back, because restoring the previous version does not undo nineteen other fixes along the way.

Database changes need separate attention. Rolling back code is quick, but rolling back a migration that dropped a column can be impossible without a backup. That is why changes to the database structure are made in two steps: first you add the new column and code that handles both versions, and the old column is removed only in a later deploy, once the new version runs stably. This arrangement lets you roll back the code without touching the data.

Dependency updates and security patches

A website or app in which nobody changes the code still ages, because its dependencies age. Libraries get security patches, runtime versions lose support and the server's operating system needs updating. A project left without updates for two years then has to jump several versions at once, which is much more expensive than regular small updates, because the changes from several versions have to be reviewed and tested at the same time.

Updates carry different weight, according to the version numbering. A patch version, the third number, should only fix bugs. A minor version, the second number, adds features without breaking compatibility. A major version, the first number, can change behaviour and requires reading the list of changes. Not every library sticks to these rules, so even a patch update should go through tests before deployment.

dependency-review.sh · bash
npm outdated
npm audit --omit=dev

composer outdated --direct
composer audit

npm audit --omit=dev skips dependencies used only for building and testing, which never reach production, and that separates real threats from noise. composer outdated --direct shows only the libraries added directly to the project, not the whole dependency tree. Tools such as Dependabot or Renovate open pull requests with updates automatically, and automated tests check them before anyone clicks merge.

The runtime has its own calendar. Node.js publishes a support schedule for every LTS version, and PHP for every major and minor version. It is worth knowing the end-of-support date of the version your project runs on and planning the move well ahead, because after that date security holes stop being patched.

Backups you can actually restore

A backup only makes sense if you can restore a working system from it. That sounds obvious, but a backup nobody has ever tried to restore may be incomplete, corrupted or saved in a format for which the right version of the tool is missing. This comes out at the worst possible moment: during an outage.

A full backup of a website or app is three things, not one. The database, the files uploaded by users, such as product photos and attachments, and the configuration needed to bring the system up from scratch. The code itself lives in the repository, so it does not need a separate backup, but secrets and environment variables should not be in the repository, so they need their own secure place.

The 3-2-1 rule says: three copies of the data, on two different media, one of them in another location. A server snapshot taken by the hosting provider is convenient, but it sits in the same infrastructure as the server. If the problem is your account with the provider or a data centre failure, the snapshot is unavailable along with the server.

A restore test does not have to be complicated. For a PostgreSQL database, it is enough to restore the dump into a temporary database and check that it contains data:

backup-restore-test.sh · bash
FILE=/backup/store-$(date +%F).dump

pg_dump --format=custom --file="$FILE" store

createdb store_restore_test
pg_restore --dbname=store_restore_test "$FILE"
psql --dbname=store_restore_test --command="SELECT count(*) FROM orders;"
dropdb store_restore_test

With backups, it is worth agreeing on two numbers. The first is the maximum amount of data the business can afford to lose: if a backup runs once a day, a failure just before the next backup means losing almost a full day of orders. The second is how quickly the system has to be back up and running. For an online store that sells mostly in the evening, both numbers look different than for an internal panel used during office hours.

After 90 days: a maintenance plan, your own team or ad hoc work

When free monitoring ends, there are three sensible routes, and none of them is best for everyone. The choice depends on how often the project changes, whether the company has its own developers and how costly an hour of downtime is for it.

OptionFor whomWhat you getWhat to watch out for
A maintenance plan with the contractorA company without its own developers, a project that generates revenue or serves customersMonitoring, updates and fixes done by people who know the codeThe scope of the plan has to be written down: what the price covers and what is a separate job
Handover to your own teamA company with a developer or an IT teamFull control and no dependence on an outside companyRequires documentation and knowledge transfer, and the team needs time for maintenance alongside its other work
Ad hoc jobsA website that rarely changes and handles no transactionsYou pay only for the work doneNobody watches the site between jobs, so a problem comes out when someone from the outside reports it

We price our maintenance plans individually, because maintaining a landing page and maintaining an online store integrated with a warehouse are completely different scopes of work. If you have a developer in the company who knows the project's technology, you most likely do not need a maintenance plan from us: it is better to invest in a proper handover and documentation, and keep us for bigger changes.

Ad hoc jobs make sense for a simple company website without a store and without logins, provided the basic things happen on their own. The certificate renews automatically, the domain has auto-renewal switched on and a free uptime monitoring tool sends notifications to the address of someone in the company. That is the minimum, which takes an hour to set up and protects against the silliest outages.

For projects that earn money or serve customers, having nobody watch the system is a risk that is easy to put a number on: just estimate what a day costs in which the store takes no orders or the panel does not let customers log in.

Handing the project over to another team: what the documentation must contain

From us, the client gets full rights to the code, access to the repository and documentation that makes it possible to take the project over. Rights to the code mean nothing if a new team needs a week to get the project running on their machine. The test for documentation is simple: can a developer who has never seen the project run it locally and deploy a change without asking the authors a single question.

Handover documentation should contain at least:

  • Running locally. The required tool versions, install commands and how to seed the database with test data.
  • Environment variables. A list of names with a description of what each one is for, without the values. Production values are handed over through a separate, secure channel.
  • Deployment procedure. Step by step, including how to roll back a version when something goes wrong.
  • Architecture in brief. What parts the system consists of, how they communicate with each other and where they run.
  • External services. Payment gateway, email delivery, maps, partner APIs, with information on whose account each is registered to.
  • Domain and DNS. Where the domain is registered, where the DNS records are managed and which records are needed for email.
  • Backups. Where they are, how often they run and how to restore them.
  • Known issues. Things that work but need attention, and technical decisions worth knowing before changing the code.

The point about external services matters more than it looks. A payment gateway account, a domain or an email delivery account registered to the contractor means that taking over the project requires their cooperation, whatever the contract says about the rights to the code. The server, the domain and the external services should be registered to the client from the start, with the contractor having access to them, not the other way round.

The last element is a conversation. Even the best documentation will not convey everything, for example why one of the integrations is written in an unusual way. An hour-long conversation in which the authors walk the new team through the project saves the new team days of guessing.

Website maintenance costs after launch, whoever built the site

Some maintenance costs have nothing to do with who built the project, and they exist even when nobody changes a single line of it. It is worth having them written down before launch, so that none of them takes you by surprise during the year.

  • Domain. Renewal every year, or for several years in advance. An expired domain takes down the website and email at the same time.
  • Server or hosting. A monthly or annual fee, depending on the resources needed.
  • Business email. If the mailboxes are not part of the hosting, a separate fee for each user.
  • Paid APIs. Maps, text messages, transactional email, address recognition. Many of them have a free tier that stops being enough as traffic grows.
  • Payment provider fees. Charged on every transaction rather than as a fixed fee.
  • Licences. Paid plugins, fonts and stock photos often come with an annual licence or one limited to a specific use.

For each of these items, write down three things: whose account it is registered to, which card or bank account pays for it and when the next renewal is due. It is easy to miss an outage whose cause is not a bug in the code but an expired company card, which meant the domain or server renewal did not go through, while the notifications went to the mailbox of someone who no longer works at the company. One shared renewal calendar, visible to more than one person, closes that gap.

A TLS certificate does not have to be a cost. Let's Encrypt issues free, short-lived certificates that need renewing every few weeks, and servers such as Caddy, or the certbot tool, renew them automatically. That is exactly why the certificate expiry date is in the monitoring table: automatic renewal works until it stops working, for example after a change to the DNS or firewall configuration.

If you are planning a website or app and want to know from the start what the post-launch period will look like, we have described the whole course of a project on our How we work page. The scope and starting price of company websites, from PLN 1,199, are on our website offer page. If you already have a project that has gone live and you are looking for someone to maintain it or take it over, write to us. We reply within an hour.