You have 1 article left to read this month before you need to register a free LeadDev.com account.
Estimated reading time: 8 minutes
Key takeaways:
- Years without incidents were luck, not proof of safety: the team was careful, not protected, so the engineering gap in the database stayed hidden.
- It was a business case, not an engineering one, that won the team over.
- The fix exposed what the old setup had absorbed, from a silent two-month failure to duplicate code.
A few years ago, I joined a team at a large bank where the business logic for every external integration – pension funds, tax authorities, retail partners, government services – lived as Groovy scripts stored in database rows. The surrounding engineering infrastructure that many teams take for granted was absent entirely, from build pipelines to deployment artifacts to any form of version control.
You opened the database, edited a script, and it went live in production the moment you hit ‘save’. For many of these integrations, there was no development or test environment, only production.
Although many don’t like to admit it, these systems can be very common, and they survive for reasons that have little to do with technical negligence. The people who work closely and comfortably with these systems usually cannot see the risk, and those who can rarely have the standing to change anything. I spent months arguing for basic engineering hygiene in a team that had lived without it for years and felt no need for it.
Your inbox, upgraded.
Receive weekly engineering insights to level up your leadership approach.
How the system worked
Groovy scripts were stored in a database as plain text, keyed by script ID. At runtime, the Java Virtual Machine (JVM) loaded them, compiled them on the fly, and executed them within a shared host process.
A typical script did roughly the same shape of work: connect to an external source over HTTPS or FTPS, decrypt the payload, transform and enrich the data, then push it onward to our internal systems.
Changes were production deployments in everything but name. A business analyst would spot a problem in production, open the row, fix the line, and move on. On its own terms, the feedback loop was fast and effective.
Scripts were also kept in Git, but only nominally. The database was the source of truth, and it was routine for someone to fix a script in production and forget to mirror the change back into the repository. When the same script came up for work again weeks or months later, the difference between Git and the database raised questions nobody could answer cleanly and necessitated extra communication.
The scripts were loaded and compiled at runtime, but they ran within a shared runtime – an “interpreter” host process that provided common resources: a database connection pool, a shared HTTP client with rate limiters, crypto modules, and so on.
As the scripts lived in the database rather than alongside the “interpreter” code, any change to the interpreter was painful. We had to be confident that every script in production would continue to work, which slowed down anything we wanted to do on scalability or performance.
Several interpreter versions ran in parallel, which made it worse. There was steady technical debt pressure to migrate old scripts to the newer version, and any urgent fix had to be ported into every version still in use.
Why nobody complained
Instead of engineers, business analysts owned the scripts. Their workflow was direct and satisfying: spot an error in production, open the database, edit the row, and watch the fix work.
The system maintained a version history of every script in a dedicated table, and rollbacks occurred from time to time. That, paradoxically, was part of why Git met so much resistance. It seemed redundant, storing code and tracking its history while missing everything underneath: CI/CD, tests, and the ability to refactor safely. There was no visible graveyard of bad changes because bad changes were overwritten on the spot.
When I proposed moving scripts to Git, introducing merge requests, and requiring tests before anything reached production, the response was the same: “what is this process for, and why do we need it?” Years of uneventful operation had been read as proof that the architecture was sound, when it was really proof that the team had been careful and lucky.
More like this
The risks were structural
The problems were structural, even if they had not yet produced a headline incident. Without proper version control, root-cause analysis after incidents would be guesswork. The in-database history table showed past versions of a script but not who changed it, why the change was made, or what else changed alongside it.
Without tests, the impact of any edit stayed unknown until production revealed it. Without a review step, one person under deadline pressure was the only thing standing between a typo and a broken integration with the tax authority.
No serious incident had occurred, but that outcome rested on the caution of a small group of analysts rather than on anything the architecture itself was doing.
What the fix required
Repairing this meant moving the source of truth out of the database and into a Git repository and introducing CI/CD processes. Scripts became files, changes became merge requests, and a small testing framework was built from scratch so that scripts could be validated in CI.
This testing framework proved its value to us as soon as we used it to bump a library version to address a known Common Vulnerabilities and Exposure (CVE). Before that framework existed, a similar change would have meant testing every single script manually in production, one at a time.
The hardest work was social
Convincing a team to adopt review and testing practices they had never needed and whose absence had never hurt them took far longer than the technical work itself. Tests came first, then versioning, and each required harder conversations than I had expected going in.
The argument that finally moved the conversation was a business one rather than a technical one, built on three points that leadership could weigh directly.
- Validated changes would lower incident risk before anything reached customers or regulators.
- Real version control and tests would speed up diagnosis when something did go wrong, giving us an actual trail to investigate.
- The bank would need a clear audit trail for compliance, regardless of the team’s preferences.
Framed that way, the changes went from an engineering preference to an operational requirement.
One argument the analysts made was harder to dismiss than I had expected. The external systems we integrated with were genuinely unstable, and their contracts could change without notice.
The team’s established response was to run to production and edit the script by hand, which had worked for them. The risk in that quick-fix workflow had never materialized, but it had never gone away either. We settled on a compromise: direct database edits remained possible, but anything in Git would overwrite them on the next deployment.
Once CI/CD matured, the merge-request path turned out to be only marginally slower than editing in production, and considerably safer – tests caught regressions and the code at least had to compile before it ran. Gradually, manual edits trailed off, merge requests took over, and nobody had to be argued into it a second time.
What the fix surfaced
Once the new workflow was running, the infrastructure paid for itself in ways we had not anticipated. Mostly, it exposed problems the old setup had been quietly absorbing. Three stand out:
A script that had been failing silently for 2 months
A change to the interpreter introduced a bug; the release went out without enough testing, and the script began exiting cleanly – no output, no error. What made it memorable was that the downstream team consuming the data never noticed. The data was not critical, but the fact that a two-month outage could pass unremarked got everyone’s attention.
A memory leak we had known about for years and never managed to fix
The application had to be restarted weekly to keep it from running out of memory. These leaks were spread across old library versions, multiple interpreter builds, and the scripts themselves. The only real fix was a clean interpreter version with every script migrated onto it – often through backward-incompatible library changes that meant rewriting parts of the script code. The prospect was demoralizing on its own. It was worse once you accepted that any failed attempt would mean retesting every script again.

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
With the test framework in place, coverage improving, and Git as the source of truth, that kind of large change became feasible. The drop in memory usage on the post-release graph was one of the more satisfying things the team had seen in a while.
The sheer volume of duplicated code we found once we could safely look at it
Given how painful interpreter and library changes had been, it was unsurprising. Copy-paste had been the path of least resistance for years, and nothing had pushed back against it.
Once Git was authoritative and tests existed, we could finally extract shared modules and consolidate. The average script length dropped by roughly a third during the cleanup, and we found several bugs introduced by careless copying along the way.
The lesson extends beyond one bank
Legacy architecture of this kind survives because changing it requires someone to absorb the transition costs before a crisis makes the need undeniable. Most organizations do not reward that kind of pre-emptive work. An absence of visible problems gets read as an absence of risk, but the two are very different things.
Engineering leaders should treat unreviewed, unversioned, and untested code paths in production as incidents waiting to be scheduled. Ask your teams where the business logic actually lives and how a change gets from someone’s head to a running system.
If the answer involves editing a database row, the work ahead is considerable, and it will go better if it happens before the regulator, the auditor, or the outage forces the timing.