Skip to main content

Cloud & DevOps

Production Readiness Checklist for Web Apps: Test Every Item

A production readiness checklist for web apps: backups, monitoring, secrets, rate limits, rollback and on-call, each with a quick test to prove it works.

Engineer reviewing a production readiness checklist with backup status, alerts and rollback steps for a web app
|Oct 10, 2026|Production ReadinessDevOpsWeb AppsReliability

What production ready actually means

Short answer — the key takeaway (TL;DR): In short, the main answer: A production readiness checklist for web apps: backups, monitoring, secrets, rate limits, rollback and on-call, each with a quick test to prove it works. Bottom line, that is the summary before the detail. Who this is for: readers researching this topic before choosing an approach.

Published: Oct 10, 2026 · Last updated: Oct 10, 2026

A web app is production ready when you have tested, not assumed, that you can recover data, spot failures before users report them, protect secrets, handle abuse, roll back a bad release and reach a named person at 2 a.m. This checklist groups those checks by area, and every item comes with a quick test you can run this week.

Most launch checklists only ask whether something is set up. That is the wrong question. A backup job that has never been restored, an alert that goes to an inbox nobody reads and a rollback button nobody has pressed all look fine on a dashboard. Then they fail on the worst possible day.

Use one rule throughout: if an item has no test you have run in the last 30 days, mark it as not done. Saying it is set up proves nothing. A test that passed does.

This is written for founders preparing a first launch, CTOs signing off on a release, and teams moving a prototype built with AI tools onto real traffic. If that last one is you, read our guide on taking a vibe-coded app to production alongside this list.

Backups and restore drills

A backup you have never restored is a hope, not a backup. The checks below confirm that you can get data back, how much you would lose and how long it takes.

Checklist
• Automated database backups run at least daily, with point-in-time recovery turned on if your database supports it.
• Backups are stored in a separate account or region from production, so one stolen credential cannot delete both.
• User uploads in object storage have versioning or their own backup, not just the database.
• Backups are kept long enough to catch silent data corruption that someone only notices a week later.
• Someone gets an alert when a backup job fails or produces a suspiciously small file.

Quick test: the restore drill. Pick yesterday's backup. Restore it into a fresh, isolated database. Point a staging copy of the app at it and log in with a real test account. Time the whole process. That number is your real recovery time, and it is usually longer than anyone guessed.

Quick test: the point-in-time check. Create a test record, note the exact time and delete it. Then restore to one minute before the deletion. If the record comes back, point-in-time recovery works. If nobody knows how to do this, write the steps down as you go. That document becomes your runbook.

Decision rule: agree on two numbers before launch. How much data can you afford to lose (recovery point objective), and how long can you be down (recovery time objective)? If the drill misses either number, fix backups before anything else on this list. Managed databases make this easier, but you still have to test. Our notes on Supabase production scaling cover the details for that stack.

Monitoring and alerting

Monitoring answers one question: will you know something is broken before a customer emails you? Dashboards alone do not count. Only alerts that reach a person count.

Checklist
• Uptime checks hit a real endpoint every minute from outside your cloud provider.
• A health endpoint checks dependencies such as the database, cache and queue, not just that the process is running.
• Error tracking catches exceptions on both server and browser, tagged with the release version.
• Logs are structured, searchable and kept long enough to investigate a problem reported days later.
• Key routes report request rate, error rate, response times at p95 and p99, and resource usage such as memory and database connections.
• Alerts fire on problems users feel, like rising errors or slow checkout, not on every CPU spike.

Quick test: break it on purpose. In staging, or in production during a quiet window, cut the database connection or make the health endpoint return a 500 error. Start a timer. How long until a phone buzzes? Who got the alert? Did the message say what is broken and where to look? If nobody got it, or it arrived as an email digest the next morning, alerting is not done.

Quick test: the thrown error. Add a hidden route that throws an exception, deploy it and call it once. Confirm the error shows up in your tracker within a minute, tagged with the release, with a readable stack trace instead of minified code. Missing source maps are one of the most common gaps we find in pre-launch reviews.

Trade-off: too many alerts are as dangerous as too few. If the on-call person gets several pages a week that need no action, they start ignoring all of them. Delete or downgrade noisy alerts and keep only the ones that need a human.

Secrets handling

Secrets are API keys, database passwords, signing keys and webhook tokens. The goal is simple: no secret in the code, in the browser or in a chat log, and every secret can be replaced calmly when needed.

Checklist
• No secrets in the git repository, including its full history.
• No server secrets in the code sent to the browser. Treat any variable exposed to the frontend as public.
• Secrets live in a secrets manager or in the hosting platform's encrypted environment settings.
• Production and staging use different keys, so a leaked staging key cannot touch real data.
• Each service gets only the permissions it needs. A read-only reporting job does not get an admin key.
• Only named people can access production secrets, and access is removed the day someone leaves.

Quick test: the repo scan. Run an open source scanner such as gitleaks or trufflehog over the entire repository history. Every key it finds must be rotated, meaning replaced with a new one. Deleting it from the file is not enough, because the old key still lives in the history and in every copy of the repository.

Quick test: the bundle search. Open the production site, open developer tools and search the loaded JavaScript for strings like sk_, secret, service_role or private_key. Anything you find there can be read by every visitor. This happens often in apps generated by AI builders. Our guide to vibe coding security lists the usual causes.

Quick test: the rotation drill. Pick one non-critical key and replace it end to end. Note every place that had to change and how long it took. If replacing one key needs a code change and a full redeploy, fix that now, before a real leak forces you to do it under pressure.

Rate limiting and timeouts

Rate limiting protects you from abuse, runaway scripts and your own buggy clients. Without it, one bad actor or one retry loop can take the app down for everyone, or quietly use up your quota on a paid third-party API.

Checklist
• Login, signup, password reset and OTP endpoints have strict limits per IP address and per account.
• Public APIs have limits per key or per user, and return a clear 429 response with a Retry-After header.
• Expensive routes such as search, exports, file processing and AI calls have tighter limits than cheap ones.
• Limits are enforced at the edge or in a shared store such as Redis, not in one server's memory, so they still hold when you run several servers.
• Calls to outside vendors have timeouts, so a slow provider cannot tie up all your database connections.

Quick test: the hammer. Use a load-testing tool such as k6 or hey to send 100 login attempts in 10 seconds from one machine against staging. Once you pass your limit, you should get 429 responses, while a normal user on another network is unaffected. Then repeat with two app servers running to confirm they share the limit.

Quick test: the slow dependency. Point one outside integration at an endpoint that never responds. The request should fail when it hits your timeout, for example 10 seconds, and the rest of the app should stay responsive. If pages across the whole site hang, a timeout is missing.

One proof point from our own work: 10M+ requests per minute handled in production (exam platform). At that kind of traffic, limits and timeouts are built in from day one, not added after the first outage. The same habit pays off at much smaller scale.

Error budgets

An error budget turns reliability into a simple rule. You choose a target, for example 99.5% of sign-in requests succeed in under one second over 30 days. The gap between that target and 100% is the budget you are allowed to spend on failures.

Why founders should care: it settles the ongoing argument between shipping features and fixing stability. While budget remains, ship. When it is spent, the team pauses risky releases and works on reliability until it recovers. Nobody has to win a debate. The number decides.

Checklist
• Set two or three reliability targets (service level objectives) on the journeys that matter: sign in, the core action and payment.
• Measure them from real traffic or automated checks, not from server uptime.
• Agree in writing what happens when the budget runs out, and who is allowed to override it.
• Review budget usage weekly for the first month after launch.

Quick test: the dashboard question. Ask anyone on the team how much of this month's error budget has been used. If they can answer in under a minute from a dashboard, the budget is real. If they have to guess, the target only exists in a document.

Start with a loose target. A new product aiming for 99.99% will spend the whole budget on its first bad deploy, and from then on everyone will ignore the rule. Pick a target you can actually meet, then tighten it as the product matures.

Rollback paths

Sooner or later, a release will go wrong. What matters is whether you can get back to the last good version in minutes, without a meeting.

Checklist
• Every deploy is a versioned build that never changes after it is made, so the previous version can be redeployed exactly.
• Rollback is one documented command or button and does not depend on one person's laptop.
• Database changes are made so that the old code still works with the new database: add new columns first, deploy code, and remove old columns in a later release. Skipping this step is the most common reason rollbacks fail.
• Risky features ship behind feature flags, so they can be switched off without a deploy.
• Background jobs keep working while old and new code run side by side during a deploy.

Quick test: the timed rollback. Deploy a harmless change to production, then roll it back while someone times it. Confirm the app works afterwards, including pages that use the latest database change. If it takes longer than about 10 minutes or needs one specific engineer, simplify the process and run the drill again.

Quick test: the flag kill. Switch a feature flag off in production and confirm the feature disappears for users within the expected time, with no redeploy.

Apps shipped straight from AI builders often have no rollback path at all, because deploys happen from inside the tool. Our cloud and DevOps work usually starts by setting up a proper deploy pipeline, versioned releases and a rollback that has actually been tested.

On-call ownership

Every alert needs a name attached. Not a team, not a shared inbox: one reachable person who knows they are on call this week.

Checklist
• A written on-call schedule covers every hour you promise to be available, including weekends and public holidays.
• Alerts go to a paging tool or phone. If the first person does not respond within a set time, for example 15 minutes, the alert passes to a second person.
• Each critical alert links to a short runbook: what it means, the first three things to check and how to roll back.
• The on-call person can reach production from wherever they actually are.
• A status page or a customer message template is ready before the first incident.
• After each incident, a short blameless review records what happened and one specific fix.

Quick test: the 2 a.m. page. Without warning, trigger a test alert outside working hours. Measure how long it takes someone to respond, check whether it passed to the backup person when it should have, and ask the responder to open the runbook and complete the first step. Everything that slows them down becomes a ticket.

Quick test: the vendor list. Ask the on-call person to name the support contact and status page for your hosting, database, payments and email providers. If they would have to search for these during an incident, add them to the runbook today.

If you work with an outside development partner, agree in writing who owns on-call after launch, and do it before go-live. Unclear ownership is how small incidents turn into lost customers.

How to run this checklist in one week

You do not need a month. A small team can run every test in this article in about five working days.

Day 1: restore drill and point-in-time check.
Day 2: break-it-on-purpose alert test and thrown-error test.
Day 3: repo scan, bundle search and one key rotation.
Day 4: rate limit hammer and slow dependency test.
Day 5: timed rollback, flag kill and the 2 a.m. page. Agree on error budget targets in the same session.

Record each result as pass, fail or not tested, with the date. Every failure gets an owner and a deadline. Run the full set again every quarter and after any major architecture change, because settings change over time without anyone noticing.

Pair these drills with functional testing of your key user journeys. Our software testing and QA team covers that side. Across 50+ products shipped, Geminate Solutions sees the same pattern: the calm launches are the ones where these drills ran before real users arrived, not after.

YK
Written by

CEO and co-founder of Geminate Solutions, a software and product development partner. He has led teams shipping custom web apps, mobile apps, SaaS platforms, and AI products that serve over 250,000 daily active users.

Free production readiness review

Find the launch gaps before your users do

Share your app URL and stack. Geminate Solutions will check it against this checklist and send a prioritized list of gaps, each with a quick test your team can run.

  • Backup, restore and rollback review
  • Monitoring and alert routing check
  • Scan of your public site code for exposed secrets
  • Rate limiting and on-call gaps flagged

Request your free readiness review

NDA before we talk. We reply within 24 hours.

Reply in 48 hours. Free, no pitch, no commitment. By submitting, you agree we may use your details to reply, under our legitimate interest and stored via EmailJS. We never sell your data. Privacy Policy.

5.0 Google reviews5.0 ClutchTop Rated on Upwork
Trusted by Volvo, L&T and the Government of Gujarat.
FAQ

Frequently asked questions

What is a production readiness checklist?
A production readiness checklist is a list of checks a web app must pass before real users depend on it. It covers backups that restore, alerts that reach a person, secrets kept out of code, rate limits, error budgets, a tested rollback and a named on-call person. Good checklists attach a test to every item, so the team proves each one works instead of assuming it does.
How often should we test backup restores?
Run a full restore drill before launch, then at least once a quarter. Run it again after any change to your database, hosting or backup tools. Restore into an isolated environment, connect a staging copy of the app and time the whole process. The time you measure is your real recovery time. If it misses your target, fix the backup setup before anything else.
What should a new web app alert on?
Alert on problems users feel: the site being down, rising error rates on key routes, slow sign-in or checkout, failed backup jobs and queues filling up. Do not page someone for every CPU spike or single error, because people start ignoring noisy alerts. Each alert should reach a named person through a paging tool or phone and link to a short runbook with the first steps.
Do early-stage startups need error budgets?
Yes, in a simple form. Pick two or three targets on the journeys that matter, such as sign-in and the core action. Agree in advance what happens when the budget is spent. This gives founders a clear rule for when to pause features and fix stability. Start with a target you can actually meet, then tighten it as the product matures.
How do we know our rollback actually works?
Deploy a harmless change to production, then roll it back while timing the process. Check pages that use the latest database change, because database changes the old code cannot handle are the most common reason rollbacks fail. If the rollback takes longer than about ten minutes, depends on one engineer or needs manual database fixes, simplify it and repeat the drill until it is routine.
Can Geminate Solutions review our app before launch?
Yes. Geminate Solutions reviews web apps against this checklist: backups, monitoring, secrets, rate limiting, rollback and on-call setup. We sign an NDA before we talk, clients own 100% of the code and IP, and we reply within 24 hours. You get a prioritized list of gaps along with the tests used to find them, so your team can check every fix.
FREE WEBSITE REVIEW

Get a free 24-hour review of your website

Send us your website link on WhatsApp. Within 24 hours we tell you exactly what is costing you customers and what we would fix first. No obligation and no sales script.

Send my website for review

4.9 rated · 50+ products shipped · 250K+ daily users served

GET STARTED

Already built something, and it is starting to break?

Most teams that reach us have a working product and a growing list of things that scare them. We read the code first and tell you what actually needs fixing, including the parts that do not. Rebuilding from scratch is rarely the honest answer.

Related Articles