Last month, we rolled back the Sofia platform to a build from six weeks earlier.
That’s not the kind of sentence most AI companies publish. We’re publishing it anyway — because the decision to roll back is exactly the kind of signal you should be looking for when you evaluate an AI vendor.
Here’s what happened, what we learned, and the three questions we now ask before we ever make a move like that again.
The Question Every RIA Principal Should Be Asking
When a firm comes to us about building or deploying an AI platform, the technical questions come fast: What models are you using? How is our data protected? What does the integration look like?
But the question that actually matters most rarely gets asked out loud:
“How do I know this vendor won’t fall apart six months after I sign?”
It’s a fair question. The AI space is littered with platforms that launched fast, broke quietly, and left their clients holding the bag. Regressions don’t always appear on dashboards. They appear in client conversations — in the moment a system gives bad advice, misses a task, or behaves unpredictably in front of someone who trusted you.
That’s the real risk. Not the demo. The six months after the demo.
What Actually Happened
We had been iterating fast on Sofia — our internal AI orchestration platform — adding capabilities, expanding agent workflows, and pushing new features on a tight cadence. The system was more capable than it had ever been.
It was also less stable than it had been six weeks earlier.
The regressions weren’t catastrophic. They were the quiet kind — small inconsistencies in agent behavior, edge cases that had been handled gracefully and now weren’t, outputs that were technically correct but no longer trustworthy in a production context.
We had a choice: keep patching forward, or stop and restore a known-good baseline.
We chose the baseline.
Why That Was the Right Call
There’s a discipline in knowing when forward momentum has become the problem.
Most engineering cultures reward shipping. Velocity is celebrated. Rollbacks feel like defeat — an admission that something went wrong. So teams patch forward, and the complexity compounds, and eventually the system is so entangled that no one is sure what’s stable and what isn’t.
We’ve seen this pattern across client engagements. A firm adopts an AI tool, the vendor keeps pushing updates, and six months later the firm’s team has quietly stopped trusting it — not because it failed dramatically, but because it became unpredictable.
Predictability is the product. Not features.
When we rolled back Sofia, we weren’t retreating. We were choosing to operate from a position of certainty rather than a position of optimism. There’s a difference.
The Three Questions We Now Ask Before Any Major Platform Decision
Coming out of this, we built a simple framework. We apply it before any significant platform change — and we think it’s worth sharing.
1. Do we know exactly what “stable” looks like right now?
Not in theory. In practice. Can you point to a specific build, a specific date, and say with confidence: this is a known-good baseline? If you can’t answer that question cleanly, you don’t have a stability baseline — you have an assumption.
2. Is the complexity we’re adding exceeding the value we’re delivering?
Every new capability adds surface area. More agents, more integrations, more edge cases to handle. The question isn’t whether the new feature is valuable in isolation — it’s whether the system as a whole is more trustworthy after you add it. When the answer starts being “probably yes,” that’s the warning sign.
3. Who would catch a quiet regression before it reaches a client?
This is the process question. Not the technical one. Loud failures get caught. Quiet regressions — the kind that make a system 15% less reliable in a way that doesn’t trigger any alerts — only get caught if someone is specifically looking for them. If you don’t have an answer to this question, you’re relying on luck.
What This Means If You’re Evaluating an AI Vendor
Ask your vendor about their rollback history.
Not as a gotcha — as a signal. A vendor who has never rolled back anything either hasn’t operated at real production scale, or they’ve been patching forward and hoping for the best.
The mature answer sounds like: “Yes, we’ve done it. Here’s what triggered the decision. Here’s what we restored. Here’s what the process looks like.”
Operational maturity isn’t the absence of problems. It’s the presence of discipline when problems show up.
That’s what we’re building at FINdustries. Not the fastest platform. The most trustworthy one.
Don Finley is the founder of FINdustries, an AI consulting and technology company helping wealth management firms navigate AI transformation. The Sofia platform is FINdustries’ internal AI orchestration system, built on Anthropic Claude.