The Slack message came in at 6:42 on a Tuesday morning. One of my senior engineers, a person who almost never messages before nine, had written a single line: "I don't think we can do this by Q3." We were four months into migrating our core ledger off a fifteen-year-old Oracle stack onto Postgres, and she was right. We couldn't. What followed was the hardest six weeks of my time as a CTO, and also the stretch where I learned the most about leading people who are tired, scared, and quietly convinced the whole thing is going to blow up on their watch.
Migrations are where good engineering leadership either shows up or gets exposed. Anyone can lead a greenfield project where every commit feels like progress. Leading a team through the slow, thankless grind of moving a live system that processes real money, where success looks like "nothing changed" and failure looks like a regulator's phone call, is a different job entirely.
Why Migrations Break Teams, Not Just Systems
The technical risk of a migration is usually well understood. You have a schema on one side, a schema on the other, a plan to move the data, and a rollback. What people underestimate is the human cost of extended uncertainty. A feature ships and there is a small hit of satisfaction. A migration drags on for two or three quarters with no visible payoff, and morale erodes in a way that is hard to see until someone messages you at 6:42am.
On a data team of nine, I watched two of my best engineers start quietly interviewing elsewhere. Not because the work was boring, but because they couldn't tell if it was ever going to end. That is the real failure mode. You don't lose a migration to a bad query plan. You lose it because the people carrying it stop believing it will finish, and their care drains out of the code one shortcut at a time.
So the first thing I tell any leader heading into one of these: your primary job is not the schema. It is keeping a group of skilled, anxious humans oriented and intact through a long stretch of low visibility. The engineering will get done if the people are okay. It will not if they aren't.
Name the Pain Out Loud
The worst thing you can do at the start of a painful migration is pretend it won't be painful. Engineers can smell manufactured optimism from across the office, and it corrodes trust instantly. When we kicked off the Oracle-to-Postgres move, I stood up in front of the team and said plainly that this would be miserable in stretches, that we would find horrors buried in the old system, and that some weeks would feel like we were going backwards.
That honesty bought me enormous credibility later. When things did go sideways, nobody felt lied to. We had already agreed it would be hard, so a bad week was just the thing we predicted, not a sign the plan was falling apart. There is a strange calm that comes from having named the monster before it shows up.
The fastest way to lose a team is to promise them a smooth ride and then hand them a rollercoaster. Tell them it's a rollercoaster and hand them the safety bar.
The Strangler, Not the Big Bang
I have a strong opinion here, formed from one very bad experience early in my career. Big-bang cutovers are almost always a mistake for anything that touches money. I once sat through a weekend cutover that was scoped for eight hours and ran to thirty-one. We had a war room, cold pizza, and a rollback plan that turned out to depend on a backup nobody had tested. We got lucky. I have never trusted luck since.
For the ledger, we went with a strangler pattern. New writes went to both systems in parallel, we reconciled the two continuously, and we moved read traffic over one account segment at a time. It was slower and it was more code, but it meant that at any given moment we could stop, and the blast radius of a mistake was a few thousand accounts rather than all of them.
- Every migration step had to be independently reversible within one hour, or it didn't ship.
- We reconciled balances between old and new systems every fifteen minutes and alerted on any drift beyond a single cent.
- No segment moved to Postgres-only until it had run in dual-write mode for at least two weeks with a clean reconciliation record.
- We kept the old system fully warm and capable of taking all traffic back until the very last segment was done.
Those rules slowed us down, and I would make the same choices again without hesitation. When you are moving a ledger, boring and reversible beats fast and clever every single time.
Protect the Team From the Business
Halfway through, the pressure from the commercial side became intense. Sales wanted a feature the old ledger simply couldn't support cleanly, and there was a real temptation to build it twice, once in the dying system and once in the new one. That way lies madness. I said no, and I said it loudly enough that it stuck.
Part of leading a migration is being the person who absorbs the organizational pressure so your engineers don't have to. If every VP who wants something can walk up to a developer's desk and lean on them, the migration will bleed out through a thousand small compromises. I made myself the single point of contact for scope decisions during that period. Every "can we just also" went through me, and the answer was almost always "after cutover."
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
My engineers didn't need to be diplomats. They needed to be left alone to do careful work. Shielding them from that noise was one of the more valuable things I did, and it barely shows up in any technical postmortem.
The Week I Got It Wrong
I want to be honest about a mistake, because migration stories that are all wisdom and no scar tissue are useless. About ten weeks in, I pushed a deadline I shouldn't have. We were behind, I felt the board watching, and I told the team we'd move three account segments in one weekend instead of one. It was a bad call driven by my own anxiety, not by the state of the work.
We hit a reconciliation mismatch at 2am on the Saturday, spent six hours chasing it, and eventually rolled the whole thing back. Nothing broke for customers, thanks to the dual-write safety net, but the team was wrecked. On Monday I stood up and owned it. I said the deadline was mine, the pressure was mine, and the bad judgment was mine. Then we went back to one segment at a time.
That apology mattered more than any architecture decision. If you want people to take smart risks and tell you the truth when things are going wrong, you have to model it yourself when your own call is the thing that failed. Leaders who never admit a bad decision end up surrounded by people who hide theirs.
Make Invisible Progress Visible
The cruelty of a migration is that most of the work produces nothing a human can see. You can spend a week untangling a stored procedure that encoded some undocumented business rule from 2011, and to the outside world absolutely nothing happened. That invisibility is poison for morale over a long project.
So we built ourselves a scoreboard. A simple dashboard on a screen in the office showed the percentage of accounts running on the new ledger, the number of reconciliation mismatches per day trending toward zero, and a count of legacy stored procedures retired. When we killed a particularly nasty piece of the old system, we rang an actual bell. It sounds silly. It worked. People need to feel the boulder moving, even an inch.
The other thing that helped was writing down the horrors we found. Every buried business rule, every "why does this table have a column called flag_2," went into a running document. It turned frustration into a kind of shared archaeology. Instead of "this system is insane," the mood became "wait until you see what I found today."
Pace It Like a Marathon
You cannot sprint a migration that takes two quarters. I have seen leaders try, running their team hot for months, and the result is always the same: burnout right at the end, when you most need people sharp for the final cutover. The last mile of a migration is where the subtle, dangerous bugs live, and you want fresh eyes on it, not exhausted ones.
We deliberately kept a sustainable pace. No heroics as a default. When we did have to pull a hard weekend for a specific cutover, it was planned, it was compensated with real time off afterward, and it was rare. I would rather the whole thing take three extra weeks than have my best data engineer make a tired mistake on the balance reconciliation logic because I'd been grinding her for two months.
Rest is not a reward you dole out at the end. On a long migration it is a piece of operational risk management, exactly like your rollback plan. A rested team catches the thing that a tired team ships to production.
The Cutover and the Quiet After
The final segment moved on a Thursday morning, deliberately not a Friday and deliberately not a weekend. I wanted the whole team awake, caffeinated, and available, with a normal business day ahead to catch anything strange. It was almost boring. We had rehearsed the steps so many times in the lower segments that the last one felt like the hundredth take of a scene, not opening night. That boredom was the entire point.
The trap right after a successful migration is to immediately swing the team onto the next big thing. Don't. We spent two weeks doing nothing but hardening, writing the documentation we'd skipped, and letting people breathe. We also held a proper retrospective, and I made sure the person who'd sent me that 6:42am message got explicit credit for raising the alarm early. She'd been right, and rewarding that kind of honesty pays dividends for years.
If you're keeping score at home: the migration slipped past our original Q3 target and finished in the second week of Q4. Nobody remembers that now. What they remember is that no customer balance was ever wrong for a second, and that we didn't torch the team to get there.

Conclusion
Here's what I didn't understand until I'd led a few of these: a painful migration is one of the best team-building tools you will ever get, and no offsite comes close. You go through something hard together, you're honest about the cost, someone owns a mistake, and you come out the other side having proven you can survive the worst kind of work there is. The engineers who suffered through that ledger move are the ones I'd trust with anything now. The database is just the byproduct. What you're really building, in all that reconciliation and grind, is a group of people who know they'll tell each other the truth when it counts. Keep the receipts on who spoke up at 6:42am. Those are the people you promote.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (0)
Leave a Comment
No comments yet. Be the first to comment!

