What production ready actually means
Published: Oct 10, 2026 · Last updated: Oct 10, 2026
A web app is production ready when you have tested, not assumed, that you can recover data, spot failures before users report them, protect secrets, handle abuse, roll back a bad release and reach a named person at 2 a.m. This checklist groups those checks by area, and every item comes with a quick test you can run this week.
Most launch checklists only ask whether something is set up. That is the wrong question. A backup job that has never been restored, an alert that goes to an inbox nobody reads and a rollback button nobody has pressed all look fine on a dashboard. Then they fail on the worst possible day.
Use one rule throughout: if an item has no test you have run in the last 30 days, mark it as not done. Saying it is set up proves nothing. A test that passed does.
This is written for founders preparing a first launch, CTOs signing off on a release, and teams moving a prototype built with AI tools onto real traffic. If that last one is you, read our guide on taking a vibe-coded app to production alongside this list.
Backups and restore drills
A backup you have never restored is a hope, not a backup. The checks below confirm that you can get data back, how much you would lose and how long it takes.
Checklist
• Automated database backups run at least daily, with point-in-time recovery turned on if your database supports it.
• Backups are stored in a separate account or region from production, so one stolen credential cannot delete both.
• User uploads in object storage have versioning or their own backup, not just the database.
• Backups are kept long enough to catch silent data corruption that someone only notices a week later.
• Someone gets an alert when a backup job fails or produces a suspiciously small file.
Quick test: the restore drill. Pick yesterday's backup. Restore it into a fresh, isolated database. Point a staging copy of the app at it and log in with a real test account. Time the whole process. That number is your real recovery time, and it is usually longer than anyone guessed.
Quick test: the point-in-time check. Create a test record, note the exact time and delete it. Then restore to one minute before the deletion. If the record comes back, point-in-time recovery works. If nobody knows how to do this, write the steps down as you go. That document becomes your runbook.
Decision rule: agree on two numbers before launch. How much data can you afford to lose (recovery point objective), and how long can you be down (recovery time objective)? If the drill misses either number, fix backups before anything else on this list. Managed databases make this easier, but you still have to test. Our notes on Supabase production scaling cover the details for that stack.
Monitoring and alerting
Monitoring answers one question: will you know something is broken before a customer emails you? Dashboards alone do not count. Only alerts that reach a person count.
Checklist
• Uptime checks hit a real endpoint every minute from outside your cloud provider.
• A health endpoint checks dependencies such as the database, cache and queue, not just that the process is running.
• Error tracking catches exceptions on both server and browser, tagged with the release version.
• Logs are structured, searchable and kept long enough to investigate a problem reported days later.
• Key routes report request rate, error rate, response times at p95 and p99, and resource usage such as memory and database connections.
• Alerts fire on problems users feel, like rising errors or slow checkout, not on every CPU spike.
Quick test: break it on purpose. In staging, or in production during a quiet window, cut the database connection or make the health endpoint return a 500 error. Start a timer. How long until a phone buzzes? Who got the alert? Did the message say what is broken and where to look? If nobody got it, or it arrived as an email digest the next morning, alerting is not done.
Quick test: the thrown error. Add a hidden route that throws an exception, deploy it and call it once. Confirm the error shows up in your tracker within a minute, tagged with the release, with a readable stack trace instead of minified code. Missing source maps are one of the most common gaps we find in pre-launch reviews.
Trade-off: too many alerts are as dangerous as too few. If the on-call person gets several pages a week that need no action, they start ignoring all of them. Delete or downgrade noisy alerts and keep only the ones that need a human.
Secrets handling
Secrets are API keys, database passwords, signing keys and webhook tokens. The goal is simple: no secret in the code, in the browser or in a chat log, and every secret can be replaced calmly when needed.
Checklist
• No secrets in the git repository, including its full history.
• No server secrets in the code sent to the browser. Treat any variable exposed to the frontend as public.
• Secrets live in a secrets manager or in the hosting platform's encrypted environment settings.
• Production and staging use different keys, so a leaked staging key cannot touch real data.
• Each service gets only the permissions it needs. A read-only reporting job does not get an admin key.
• Only named people can access production secrets, and access is removed the day someone leaves.
Quick test: the repo scan. Run an open source scanner such as gitleaks or trufflehog over the entire repository history. Every key it finds must be rotated, meaning replaced with a new one. Deleting it from the file is not enough, because the old key still lives in the history and in every copy of the repository.
Quick test: the bundle search. Open the production site, open developer tools and search the loaded JavaScript for strings like sk_, secret, service_role or private_key. Anything you find there can be read by every visitor. This happens often in apps generated by AI builders. Our guide to vibe coding security lists the usual causes.
Quick test: the rotation drill. Pick one non-critical key and replace it end to end. Note every place that had to change and how long it took. If replacing one key needs a code change and a full redeploy, fix that now, before a real leak forces you to do it under pressure.
Rate limiting and timeouts
Rate limiting protects you from abuse, runaway scripts and your own buggy clients. Without it, one bad actor or one retry loop can take the app down for everyone, or quietly use up your quota on a paid third-party API.
Checklist
• Login, signup, password reset and OTP endpoints have strict limits per IP address and per account.
• Public APIs have limits per key or per user, and return a clear 429 response with a Retry-After header.
• Expensive routes such as search, exports, file processing and AI calls have tighter limits than cheap ones.
• Limits are enforced at the edge or in a shared store such as Redis, not in one server's memory, so they still hold when you run several servers.
• Calls to outside vendors have timeouts, so a slow provider cannot tie up all your database connections.
Quick test: the hammer. Use a load-testing tool such as k6 or hey to send 100 login attempts in 10 seconds from one machine against staging. Once you pass your limit, you should get 429 responses, while a normal user on another network is unaffected. Then repeat with two app servers running to confirm they share the limit.
Quick test: the slow dependency. Point one outside integration at an endpoint that never responds. The request should fail when it hits your timeout, for example 10 seconds, and the rest of the app should stay responsive. If pages across the whole site hang, a timeout is missing.
One proof point from our own work: 10M+ requests per minute handled in production (exam platform). At that kind of traffic, limits and timeouts are built in from day one, not added after the first outage. The same habit pays off at much smaller scale.
Error budgets
An error budget turns reliability into a simple rule. You choose a target, for example 99.5% of sign-in requests succeed in under one second over 30 days. The gap between that target and 100% is the budget you are allowed to spend on failures.
Why founders should care: it settles the ongoing argument between shipping features and fixing stability. While budget remains, ship. When it is spent, the team pauses risky releases and works on reliability until it recovers. Nobody has to win a debate. The number decides.
Checklist
• Set two or three reliability targets (service level objectives) on the journeys that matter: sign in, the core action and payment.
• Measure them from real traffic or automated checks, not from server uptime.
• Agree in writing what happens when the budget runs out, and who is allowed to override it.
• Review budget usage weekly for the first month after launch.
Quick test: the dashboard question. Ask anyone on the team how much of this month's error budget has been used. If they can answer in under a minute from a dashboard, the budget is real. If they have to guess, the target only exists in a document.
Start with a loose target. A new product aiming for 99.99% will spend the whole budget on its first bad deploy, and from then on everyone will ignore the rule. Pick a target you can actually meet, then tighten it as the product matures.
Rollback paths
Sooner or later, a release will go wrong. What matters is whether you can get back to the last good version in minutes, without a meeting.
Checklist
• Every deploy is a versioned build that never changes after it is made, so the previous version can be redeployed exactly.
• Rollback is one documented command or button and does not depend on one person's laptop.
• Database changes are made so that the old code still works with the new database: add new columns first, deploy code, and remove old columns in a later release. Skipping this step is the most common reason rollbacks fail.
• Risky features ship behind feature flags, so they can be switched off without a deploy.
• Background jobs keep working while old and new code run side by side during a deploy.
Quick test: the timed rollback. Deploy a harmless change to production, then roll it back while someone times it. Confirm the app works afterwards, including pages that use the latest database change. If it takes longer than about 10 minutes or needs one specific engineer, simplify the process and run the drill again.
Quick test: the flag kill. Switch a feature flag off in production and confirm the feature disappears for users within the expected time, with no redeploy.
Apps shipped straight from AI builders often have no rollback path at all, because deploys happen from inside the tool. Our cloud and DevOps work usually starts by setting up a proper deploy pipeline, versioned releases and a rollback that has actually been tested.
On-call ownership
Every alert needs a name attached. Not a team, not a shared inbox: one reachable person who knows they are on call this week.
Checklist
• A written on-call schedule covers every hour you promise to be available, including weekends and public holidays.
• Alerts go to a paging tool or phone. If the first person does not respond within a set time, for example 15 minutes, the alert passes to a second person.
• Each critical alert links to a short runbook: what it means, the first three things to check and how to roll back.
• The on-call person can reach production from wherever they actually are.
• A status page or a customer message template is ready before the first incident.
• After each incident, a short blameless review records what happened and one specific fix.
Quick test: the 2 a.m. page. Without warning, trigger a test alert outside working hours. Measure how long it takes someone to respond, check whether it passed to the backup person when it should have, and ask the responder to open the runbook and complete the first step. Everything that slows them down becomes a ticket.
Quick test: the vendor list. Ask the on-call person to name the support contact and status page for your hosting, database, payments and email providers. If they would have to search for these during an incident, add them to the runbook today.
If you work with an outside development partner, agree in writing who owns on-call after launch, and do it before go-live. Unclear ownership is how small incidents turn into lost customers.
How to run this checklist in one week
You do not need a month. A small team can run every test in this article in about five working days.
Day 1: restore drill and point-in-time check.
Day 2: break-it-on-purpose alert test and thrown-error test.
Day 3: repo scan, bundle search and one key rotation.
Day 4: rate limit hammer and slow dependency test.
Day 5: timed rollback, flag kill and the 2 a.m. page. Agree on error budget targets in the same session.
Record each result as pass, fail or not tested, with the date. Every failure gets an owner and a deadline. Run the full set again every quarter and after any major architecture change, because settings change over time without anyone noticing.
Pair these drills with functional testing of your key user journeys. Our software testing and QA team covers that side. Across 50+ products shipped, Geminate Solutions sees the same pattern: the calm launches are the ones where these drills ran before real users arrived, not after.







