I have spent much of my career believing that some rules are simply rules.
You don’t knowingly put an untested financial process into production. You don’t experiment with live customer transactions. You build it, test it, reconcile it, test the exceptions, obtain user acceptance and then when you have reasonable confidence that it works you carefully move it into production. Pretty basic stuff.
So I had a predictable reaction recently when I found myself discussing an extraordinarily aggressive technology development path. AI and modern development tools allow new modules to be built at a pace I had never experienced before. Rather than wait until an entire legacy system could be replaced, the idea was to move pieces of functionality into production incrementally, perhaps processing only a small portion of existing live transaction volume through the new system to start. In effect, we would be doing some of the learning in production.
My reaction was immediate. You don’t test with live transactions. That is a cardinal sin.
I had plenty of reasons for believing that. In the world I come from, a transaction isn’t an interesting piece of data moving through an elegant technology stack. It is somebody’s money. It affects an account balance. It may create a fee, a commission, a remittance obligation or an accounting entry. It may later be reversed. It may interact with another transaction that arrives five minutes later. Eventually everything has to reconcile. A software error isn’t just a software error once it gets into this world. So I pushed back hard.
But then we started working through actual failure scenarios, rather than arguing about development philosophy. Suppose, for example, that a new payment-processing module accidentally processes the same payment twice. Clearly that is bad. Under my traditional way of thinking, it also proves the point: this is exactly why we don’t test unfinished systems using live transactions.
But what if another control independently examines the day’s transactions and identifies the duplicate almost immediately? What if the duplicate can be reversed before the payment processor actually clears it through the banking network? The customer never sees the erroneous transaction. The financial activity still reconciles. The development team now has a real example of the condition that caused the duplicate and can correct the logic so it doesn’t happen again.
We have discovered a defect in production. But have we experienced a control failure? I am no longer sure the answer is yes. In fact, we may have learned more, and learned it faster, by exposing a very limited population to real operating conditions than we would have learned by trying to reproduce every possible combination of data, timing, history and system interaction in an artificial test environment.
That realization bothered me a little, because it forced me to reconsider something I had been treating as a principle. Maybe “don’t test in production” isn’t a principle at all. Maybe it is a rule. Rules and principles aren’t the same thing.
The Building Code
I’ve written before about foundations in technology and the importance of understanding the underlying business, data and processes before attempting to transform them. This is something different. This is the building code.
Building codes contain very specific requirements. Those requirements change as materials, engineering techniques and our understanding of risk change. But behind those requirements are more durable principles. Buildings shouldn’t collapse. People need to be able to escape a fire. Electrical systems shouldn’t electrocute the occupants. Water shouldn’t undermine the structure.
The particular rules we use to accomplish those things can and should change as technology improves. But when we remove an old requirement, we had better understand why it was there in the first place.
I think the same thing applies to technology controls. For decades, separating development, testing and production has been one of our building-code requirements. But why?
The underlying principle wasn’t that software somehow becomes safe merely because it spent time in an environment labeled “TEST.” The principle was that we should not expose customers or the business to uncontrolled consequences while we are still discovering whether an unproven system works. Once I framed it that way, the question changed. Perhaps modern technology really does allow us to satisfy that principle differently.
What If We Can Fail Safely?
AI dramatically accelerates development. Modern data platforms can compare enormous transaction populations almost instantly. Automated monitoring can identify exceptions far faster than people reviewing reports the following morning. Transactions can be traced, stopped and sometimes reversed before they create an external consequence. Those capabilities matter.
Perhaps instead of trying to eliminate every possible failure before production, we can design systems that allow tightly controlled failures to occur while ensuring that they are immediately detected, contained, understood and corrected. That is a very different proposition from “move fast and break things.” It is closer to move fast, but know immediately what broke. Contain it before anyone is harmed. Understand why it broke. Fix it. Then move faster.
If we can genuinely do that, perhaps some of the old development rules should change. But there is an enormous “if” buried in that sentence.
The New Technology Doesn’t Get a Free Pass
It would be easy to take the argument too far in the other direction. AI lets us develop faster, therefore traditional controls are obsolete. No, that confuses technological capability with operational readiness. If an old control is going to disappear, the burden should be on the new design to demonstrate how it addresses the risk the old control was intended to mitigate.
Take the duplicate transaction example. It isn’t enough to say that we can reverse it. How do we know it happened? How quickly do we know? Is detection independent of the same system that created the error? Can we stop the transaction before it reaches the customer? Can we trace exactly what happened? Can we reconcile the complete population afterward? What happens if the error isn’t a duplicate but something we didn’t anticipate? And perhaps most importantly, what happens when the new system doesn’t know that it is wrong?
That last question may become increasingly important in an AI-enabled world. An obvious failure is relatively easy to control. A plausible-looking wrong answer is much harder.
The Real World Is Part of the System
There is another side to this that traditional development thinking sometimes overlooks. Test environments aren’t the real world. Real systems contain years of history, strange data, unusual sequences, external processors, human behavior, timing differences and exceptions created by exceptions created by earlier exceptions.
We can attempt to reproduce all of that in test. We can never really reproduce all of it. At some point, every new system has to encounter reality. Perhaps the question isn’t whether that learning happens in production. Some of it inevitably will. The better question may be how much uncertainty we allow into production at once, and what mechanisms we have built to prevent that uncertainty from becoming an uncontrolled business event.
That suggests a different model. Start with a deliberately small population. Know exactly which transactions are going through the new process. Maintain complete lineage from input through final disposition. Monitor the results independently. Define the exceptions that require intervention. Establish thresholds that automatically stop the experiment. Make transactions reversible wherever possible. Reconcile the entire population. And only increase the volume when the evidence demonstrates that the system deserves a larger operating envelope.
That isn’t abandoning control. It may actually be a more sophisticated form of control.
Rules Are Easier Than Principles
There is comfort in rules. “Never test in production” requires very little judgment. Neither does “AI lets us move faster.” One belongs to the traditionalist. The other belongs to the technology enthusiast. Both can become intellectually lazy.
The harder question is: Why did the old rule exist? What failure was it designed to prevent?
Does that risk still exist? And if the technology now allows us to abandon the rule, what has replaced the protection the rule provided? That seems to me a better way to think about many of the changes AI is bringing to business.
Some practices we have treated as immutable may turn out to be artifacts of technological limitations that no longer exist. We should be willing to challenge them. But the principles underneath them deserve much more respect. Completeness still matters. Accuracy still matters. Reconciliation still matters. Traceability still matters. Protecting the customer still matters. Knowing when something has gone wrong still matters. AI doesn’t repeal any of those things. It may simply give us much better tools for achieving them.
I started with the wrong question. My initial question was essentially why are we testing this in production? I still think that’s a reasonable question. I’m just no longer sure it’s the best one. A better question is more fundamental. Show me how this production experiment protects us from the risks that testing outside production was designed to prevent.
If the answer is convincing, then perhaps the building code should change. If the answer isn’t convincing, then development speed isn’t the solution. We have simply found a much faster way to build something unsafe. That distinction seems particularly important right now.
AI is giving us capabilities at a pace that makes many established practices look painfully slow. Some of those practices probably should disappear. But before tearing them out, it might be worth looking underneath them.
Sometimes an old rule is merely an old rule. Sometimes it is carrying the accumulated lessons of failures we have forgotten. The challenge isn’t preserving the old building code forever. It is remembering why we wrote it.
Leave a comment