On the last night of September, one of the largest cloud platforms lost gateway connectivity across eighteen regions for almost six hours. Outages like that are loud: dashboards turn red, status pages update, a postmortem follows. They are not the failures that worry us most. The ones that worry us are quiet — a system that is wrong underneath while every check says it is fine. We nearly built one last week, on purpose, for good reasons.
Moving the files while people are using them
Every conversation on RoboDesk can carry files: the photo of a damaged delivery, a voice note, a PDF invoice, the video of a blinking error light. Those live in object storage, one bucket per customer. Last week we finished moving all of them from a self-hosted store we ran ourselves to a managed cloud service — more than thirty customers, switched one at a time over seven days.
The order was deliberate. Back everything up first. Apply the agreed ninety-day retention before copying, so we did not pay to move files we had promised to delete. Copy each customer’s bucket, then switch that customer outside their own working hours — which meant collecting those hours tenant by tenant, because one runs around the clock and another closes at three in the afternoon. After each switch, a person ran the same sixteen checks by hand: send and receive an image, a voice note, a file, a video, several attachments at once; preview and download each; upload a logo; and open an attachment in a conversation from months ago. Customers whose AI reads images and voice notes got a few more.
Those checks earned their keep on the first night. The first customer to be switched failed two of them: several attachments sent together returned an “invalid parameter” error, and one less common image format rendered broken. We stopped, left the next two customers on the old storage, fixed both, and carried on. One later customer’s old media would not open at all — and would not open from the old store either. Those files had gone missing long before the move, which is a different problem, and a useful thing to know.
The slowest part was not where anyone expected. The first batch of three buckets ran as sixteen parallel workers with a twelve-hour limit, and hit the limit. Once a worker started copying, it moved its share in a minute or two. Most of the time went on asking the old store for the list of what it held. Moving data is cheap; finding out what you have can be the whole job.
The fallback, and what it can hide
A move like this has an awkward middle. A customer has been switched, but a file they uploaded an hour before the switch may not have been copied yet. So reads got a fallback: if the new store says it does not hold a file, ask the old one. Writes do not fall back; new files only ever go to the new store. That is what let us switch customers without a freeze, and it is the right design for the job.
The engineers who wrote it knew the risk. A comment in the code explains why network errors must not trigger the fallback: if the new store were simply unreachable, every read would quietly go to the old one, and the customer would look healthy while being served entirely from pre-migration storage. That is exactly the quiet failure — and in review we found a second door into it. A missing bucket was being treated like a missing file. Mistype one bucket name in a customer’s configuration, or delete the bucket, and every read falls back and succeeds. Old attachments open. Agents notice nothing. Only new uploads fail, which sends whoever is debugging in the wrong direction, because “storage is down” does not match a screen full of working files.
Right about the risk, wrong about the fix
The review proposed a one-line change: stop counting the “no such bucket” error as a miss. The engineer who owned the code agreed with the risk and then showed that the change would have done nothing. A missing bucket never arrives under that name. Downloads treat any “not found” response as a missing file without reading what the store says, and existence checks are requests with no body at all, so there is nothing to read. The gap was not one wrong entry in a list. It was every provider, through every path.
So the fix moved. Instead of trying to tell a missing bucket from a missing file one request at a time — which would have meant matching on error messages each library words differently — check that the bucket exists when the service starts and whenever a customer is switched, and refuse to run if it does not. That fix is in review now. It replaces a guess made on every read with a certainty established once, at the moment a mistake is cheapest to see.
There is a second reason to care, and it is the question now on the table: when can the old store be switched off? As long as the fallback can serve reads silently, nobody can say for certain that it is not still serving some. The fallback that made the migration safe is also what makes its last step hard to call. The same reasoning applies to where customer data is allowed to live: “we moved it” is only true if nothing is still reading from the old place.
What generalises
Fallbacks are where systems learn to lie politely. They exist to keep a customer working while something underneath is wrong, and they do that job exactly as well as they hide it. Three habits follow. Decide in writing which failures a fallback may absorb — a missing file, yes; a missing bucket or an unreachable store, never — and check configuration once, loudly, at start-up rather than inferring it from errors on every request. Count every read a fallback serves, because a fallback nobody can see is a second production system nobody is running on purpose. And give it an end date: the migration is finished the day that count reaches zero and stays there, not the day the last bucket is copied.
