A product team had a routine they had followed for years: finish the week's work on Friday afternoon, put the latest version live, and go home. It worked often enough that nobody questioned it — until the evening when a broken release went out and the phone started ringing at nine o'clock.
The code was not the problem
What followed was the classic worst-case recovery. Nobody knew exactly what had changed, because changes were made directly on the live system. Nobody could say how to get back to the previous version, because there was no previous version saved anywhere — only the current files, overwritten. Two engineers spent the night rebuilding from memory while the service stayed partly broken.
The code was not the problem that night. The delivery was.
Three changes, none of them glamorous
We changed three things, and none of them were glamorous. First, every version is now built into a package and stored, so that returning to the previous one means running a single command rather than reconstructing files by hand. Second, changes go to a staging copy of the site first, where the same checks run automatically — does it start, do the main pages respond, do the critical functions complete. Third, releases stopped being on Fridays. Not out of superstition: a change that breaks is far cheaper to fix on a Tuesday morning when everyone is awake.
None of this makes failures impossible. It makes them short. A rollback that used to take a whole night now takes a couple of minutes, and it is a known procedure rather than an improvised rescue.
Failures become short, not impossible
Teams often treat release process as bureaucracy — the part of the work that has nothing to do with the product. The night when something breaks is the moment it turns out to be the whole product, because it decides whether you sleep or whether you explain yourself to customers at breakfast.
If nobody on your team can answer the question "how do we go back to yesterday's version", that is not a technical detail. That is the plan.