The Forgotten Art of Maintenance: Why Software Is Rotting

Started by NightOwl, Yesterday at 01:49 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: The Forgotten Art of Maintenance: Why Software Is Rotting   Views(Read 83 times)
Active members in this topic:
NightOwl(1)

NightOwl

Every product keynote celebrates the same thing, a new feature, a redesigned interface, a faster version of something that already worked. Almost nobody stands on a stage to announce that a team spent six months quietly fixing the parts of a system nobody was allowed to see break. That asymmetry, glamorous new features get funded and celebrated, unglamorous maintenance gets deferred and ignored, has produced a genuinely enormous, largely invisible crisis sitting underneath the world's banks, hospitals and government agencies, and 2026 is turning out to be the year several major institutions run directly into its consequences.

The scale of the problem: how much of the world still runs on 1960s code

The specific technology at the centre of this crisis is COBOL, a programming language developed in 1959, specifically designed for business data processing, and never intended to still be running the world's financial system sixty five years later. And yet it is. Widely cited industry estimates put the total footprint at roughly 220 billion lines of COBOL code still active in production systems worldwide, processing an estimated 3 trillion dollars in commerce every single day. Reuters reporting has put the figure even more starkly for one specific transaction type, roughly 95 percent of ATM card swipes worldwide still pass through a COBOL system at some point in the transaction chain. This is not a fringe legacy problem confined to a handful of forgotten institutions, IBM has reported that 92 of the world's top 100 banks still rely on mainframes as their core operating system, all of the world's top 10 insurers depend on the same underlying architecture, and more than two thirds of the top 25 global retailers run critical operations through it as well. As of 2025, roughly 70 percent of banks globally were still relying on some form of legacy banking system, with more than 43 percent specifically still running COBOL at their core.

Why does this ancient code persist rather than simply getting replaced? Not, as outside observers often assume, out of pure institutional laziness. These systems are frequently embedded in decades of accumulated, extraordinarily specific business logic, tax rules, regulatory compliance requirements, edge cases discovered and patched one at a time over half a century, that nobody can fully document, let alone confidently rewrite from scratch without risking a catastrophic gap in coverage. A widely cited Japanese government report from the Ministry of Economy, Trade and Industry has warned that if this kind of legacy technical debt continues being ignored, Japan alone could face up to 12 trillion yen in annual economic losses beginning in 2026, and separately found that a third of organizations surveyed already cite loss of system knowledge, meaning nobody left who genuinely understands how the existing system actually works, as their single biggest barrier to modernization.

The incentive problem: why maintenance never gets funded

The deeper structural issue is not really technical at all, it is an incentive problem baked into how organizations decide what engineering work gets prioritized and funded. New features generate visible, measurable business outcomes, a product launch, a quarterly growth number, a press release, that executives and boards can point to directly. Maintenance work, patching a vulnerability nobody has exploited yet, refactoring a module so the next feature is easier to build, upgrading a dependency before it becomes unsupported, produces no such visible outcome when done well, its entire success condition is that nothing bad happens, which is a genuinely difficult thing to put in a slide deck or justify against a competing feature request with an obvious revenue story attached. Industry surveys of mature codebases consistently find that 30 to 40 percent of total developer time already goes toward technical debt related work, patching around problems that better upfront decisions would have prevented entirely, time that leadership would generally prefer was spent building something new and visible instead.

The scale of accumulated debt reflects this dynamic playing out across the entire industry for over a decade. One widely cited estimate found global technical debt roughly doubled between 2012 and 2023, growing by approximately 6 trillion dollars over that period, and separate industry surveys have found nearly 70 percent of organizations now view their accumulated technical debt as having a high level of impact on their actual ability to innovate, a genuinely ironic outcome given that debt typically accumulates specifically because organizations prioritized innovation speed over the maintenance work that would have kept that speed sustainable. Research into why developers actually introduce technical debt in the first place consistently points to the same handful of causes, unrealistic deadlines, workload pressure and pressure from management to ship something visible on schedule, rather than any lack of skill or awareness among the engineers themselves, developers in these same studies frequently report knowing exactly what the right long term solution would be and being denied the time to build it.

What happens when the debt comes due

The consequences of deferred maintenance tend to stay invisible for years, right up until a sudden spike in demand or a routine vendor decision forces a reckoning all at once. The COVID-19 pandemic in 2020 produced exactly this kind of forced reckoning for several US state governments, when a sudden, massive surge in unemployment claims overwhelmed COBOL based mainframe systems that had never been designed or resourced to handle that kind of volume. New Jersey's governor Phil Murphy made a public plea for volunteers who could still program in COBOL, a language most computer science graduates had never been taught, in states including New York, Florida and Ohio, the same underlying legacy architecture visibly slowed the delivery of emergency economic relief to people who needed it immediately, turning what should have been an invisible backend limitation into a genuinely visible public failure with real human consequences.

The banking sector is facing its own version of a forced reckoning on a specific, hard deadline rather than a sudden crisis. Major legacy operating systems and foundational tools running on IBM mainframes are scheduled to lose official vendor support between 2026 and 2027, forcing institutions to either complete migration work now or continue running trillion dollar financial systems entirely without vendor support, a genuinely serious operational and compliance exposure for institutions of that scale. Layered on top of that hard technical deadline is a demographic one, the average COBOL programmer today is roughly 55 years old, and industry analysts project that by 2030 the generation of engineers who originally wrote these systems in the 1970s and 1980s will have fully exited the workforce. If banks have not meaningfully decoupled their core operations from COBOL by that point, even routine regulatory updates, a tax code change, a new compliance requirement, could genuinely take months or years to implement rather than weeks, simply because too few people alive still understand how to safely modify the underlying system. Banks are already responding to the early stages of this talent shortage with what industry observers have nicknamed the boomerang strategy, rehiring their own retired engineers back as contractors at rates commonly ranging from 100 to 300 dollars or more per hour purely to handle routine patches, while junior developers willing to specialize in mainframe systems can now command starting salaries north of 125,000 dollars specifically because so few people are entering the field compared to the overcrowded pipeline of web and app developers.

Can this actually be fixed, and what would it take

The honest answer is that wholesale replacement of these systems is rarely realistic on any timeline that matters, 220 billion lines of code encoding six decades of accumulated business logic simply cannot be safely rewritten from scratch without an unacceptable risk of silently breaking something nobody remembers is there. The more credible path forward that modernization specialists actually recommend is incremental, wrapping legacy systems in modern interfaces that let newer software talk to the old core without requiring the old core itself to be torn out immediately, migrating specific high risk components one at a time, and building genuine institutional discipline around treating maintenance as a first class, continuously funded activity rather than a discretionary cost that gets cut whenever budgets tighten. Some organizations have started building specific rituals to normalize this shift, regular sessions where teams openly share maintenance improvements the way they would a new feature demo, technical debt ratios tracked and reported to leadership the same way a growth metric would be, and deliberate reward structures for engineers who reduce risk rather than only for those who ship something new and visible. None of this is as exciting as a product launch, and that is precisely the point, the entire crisis described in this piece exists because excitement, not actual risk or actual cost, has quietly been the deciding factor in how engineering time gets allocated for decades, and the institutions now facing hard 2026 and 2027 deadlines are the ones finding out what that tradeoff actually cost them all along.

Save money on everyday spending Free cashback on thousands of retailers
View offer