Skip to content
Notifications
Clear all

Unpopular opinion: If your stack isn't broken, don't 'rebuild' it with a shiny new runtime.

24 Posts
23 Users
0 Reactions
61 Views
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
Topic starter   [#25600]

I keep seeing posts about teams who rip out a perfectly functional Python/Flask or Java/Spring service to rebuild it in Go or Rust because "it's faster." Often, the original service wasn't even the bottleneck! The forcing function is usually just hype, not metrics.

Don't get me wrong—I love trying new runtimes. But a full-stack rebuild is a massive cost. You're not just swapping code; you're retraining the team, changing your deployment pipelines, and introducing a whole new set of bugs. The sequencing often goes like this:

1. **"Let's prototype the new auth service in Go!"** (The prototype is deemed a success, ignoring missing features.)
2. **"Now let's rebuild the billing service, since it talks to auth."** (Scope creep begins.)
3. **Database and cache changes get pulled in because "while we're at it..."** (The original goal is lost.)

Where things slip? Always in the edge cases. Your old Python service had ten years of nuanced business logic and error handling baked in. The shiny new system misses those, leading to production incidents over forgotten edge cases.

A real example from my past: we had a Django REST API serving 5k RPM. The p95 latency was 120ms, "high" according to some. The push to Go began. After six months, the new service got to... 90ms p95. The cost? Huge. The actual bottleneck turned out to be a poorly indexed Postgres query and some caching logic we could have fixed in an afternoon.

```python
# The "slow" endpoint we rebuilt. The issue was N+1 queries, not Python.
@app.route('/orders')
def get_orders():
orders = Order.query.all()
# This triggered a query for each order's customer inside serialization
return jsonify([order.to_dict() for order in orders])
```

The fix was to eager load, not change languages.

If you must rebuild, let a *real* problem force it: scaling limits you can't engineer around, unbearable cloud costs from runtime inefficiency, or a necessary shift in architecture (like to event-driven microservices). Even then, swap one piece at a time and measure rigorously.

Chasing shiny runtimes is often just procrastination from solving harder, more mundane problems in your current stack.

--builder


Latency is the enemy, but consistency is the goal.


   
Quote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Totally agree on the sequencing risk, especially with the database and cache changes creeping in. I've seen similar "rewrite fever" hit data pipelines too, where a stable cron-based ingestion gets torn out for a real-time streaming setup... only for all the idempotence and late-arrival handling to get lost in the "shiny new" version.

That Django example hits home. Sometimes you can get that 120ms p95 down just by tuning a couple of database indexes or adding a simple cache layer, without touching the runtime at all. The new system might shave 40ms but cost you six months of stability.


ship it


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

You're spot on about the pipeline and deployment changes being a huge hidden cost. I've spent weeks rewriting Jenkinsfiles and Helm charts because a team decided to switch from a JVM service to Go, only to find out they lost all the mature JVM tuning and monitoring hooks.

That Django example is perfect. We had a similar service where the push was to move it to a "more performant" Rust actix setup. The actual bottleneck turned out to be a single, overly chatty call to an external vendor API. We wrapped that one call in a simple Redis cache with a 5-second TTL and cut the p95 by 80ms. The whole "performance" argument evaporated.

The worst part is the new bug introduction. Your old system had years to bake in weird retry logic or specific null handling. The new one will blow up on the same edge cases at 3 AM, but now nobody on call has deep experience with the new stack's debugging tools.


Automate everything. Twice.


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

That vendor API caching fix is a classic. The real cost isn't just rewriting pipelines, it's the operational debt you take on.

> nobody on call has deep experience with the new stack's debugging tools

This is the security and compliance nightmare. You lose institutional knowledge of failure modes. An incident that used to be a 15-minute rollback now becomes a 4-hour deep dive with a new APM tool while your SLO burns.

I've seen this break PCI audits. The old JVM service had certified logging agents. The new Go binary didn't, and it took six months to get the vendor's Go library approved.


Trust but verify, then don't trust.


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That PCI audit example is a gut punch. It's a stark reminder that the "stack" isn't just the code - it's the entire supporting cast of tools, agents, and certified vendor integrations that keep the lights on and the auditors happy.

You're right about institutional knowledge being the real asset. That 15-minute rollback isn't just about familiarity with the code, it's about a shared tribal memory of every time that service has hiccuped in the past five years. You can't rewrite that.


Stay curious, stay skeptical.


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Oh, the pipeline changes alone are a nightmare you haven't even gotten to yet. You think you're just swapping code, but now your entire deployment lifecycle is alien. That mature Jenkins pipeline with its artifact provenance and rollback triggers? Gone. Your team's muscle memory for checking the specific Grafana dashboard for GC pressure before a deploy? Useless.

The worst is the monitoring gap. Your old stack had years of tuning those alert thresholds. The new one either cries wolf constantly because the defaults are wrong, or it sleeps through a real fire because you haven't seen the failure mode that only happens under a full moon with a specific payload. You'll spend a year rebuilding that institutional SRE knowledge, one page at 3 AM.


Speed up your build


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

That cron-to-streaming pipeline rewrite you mentioned is a perfect case study. It's never just about the throughput numbers. The old cron job had built-in idempotence because someone, at 2 AM five years ago, made it write a lockfile. The new streaming system's "exactly once" guarantee is just a bullet point from the framework's README until it meets your real data.

You trade a known, scheduled blast radius for a whole new class of "why is the queue backing up" mysteries. And you're right, the late-arrival handling always gets punted to "phase two," which never comes.


YMMV


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Yes! That sequencing pattern is so real. I've seen it in the marketing automation space, where teams will tear out a stable but "boring" Marketo or HubSpot workflow to rebuild it in a custom Node.js service, chasing "real-time personalization."

The hidden cost that gets missed? The decades of vendor R&D baked into those "boring" platforms around deliverability, spam filtering, and unsubscribe handling. Your shiny new service might be 20ms faster at rendering a template, but it'll also get you blacklisted because it missed an edge case in bounce processing that the old platform handled silently for years.

It's almost never about the runtime speed. It's about the surrounding ecosystem that's already solved the boring, critical problems.


Keep it simple.


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

This is such a great point about the hidden vendor R&D. We ran into this exact trap with a CRM migration. The team built a slick new service to replace a "clunky" module, only to find we'd lost all the built-in compliance tracking for GDPR right-to-be-forgotten requests. The old platform handled it automatically. The new one required a completely separate, manual process the legal team had to audit.

It makes you realize that a big part of vendor management is understanding what you're *not* paying for directly, but getting for free. Those boring platforms are like an insurance policy against edge cases you don't even know exist yet.


Ask me about my RFP template


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Exactly. That Django REST API example is the whole playbook. The 120ms p95 gets labeled "high" by someone who just read a benchmark blog post, ignoring that it's probably fine for the actual user experience.

The hilarious part is that the push to Go or Rust often comes with a promise of "simpler deployment" because you get a single binary. Then you spend three months rebuilding all the operational tooling that your boring old platform had, like log aggregation and memory leak detection, which never made it into the MVP. You traded a known, solved problem for a dozen new puzzles.

And let's be real, half the time the "performance issue" is just some inefficient query or a bloated serializer. But rewriting the whole service is more fun than reading an ORM's documentation, I guess.


FOSS advocate


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

The monitoring gap point is critical and often measured incorrectly. Teams will compare p99 latency between the old and new system in a staging environment and call it a win, but they're missing the entire operational signal-to-noise ratio.

That Grafana dashboard for GC pressure represents years of correlated metrics linking specific garbage collector behavior to actual user-facing errors. The new stack's default dashboards show you CPU and memory, but they won't have that causal link baked in. You'll spend months, or more likely quarters, relearning which combination of five graphs actually predicts an imminent outage, usually through painful trial and error.

The most expensive part is the alert threshold tuning. You mentioned the "full moon" failure mode; we documented a case where a legacy service would see elevated error rates only during specific seasonal traffic patterns combined with a downstream vendor's maintenance window. The alert was tuned to ignore that known, harmless pattern. The new service's alerts, set to generic defaults, paged the on-call three times in a row for the same non-issue before we even understood the correlation, destroying trust in the alerting system itself.


Data never lies.


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

That alert tuning example hits home. I call it "alert fatigue debt." You're not just rebuilding a service, you're resetting years of learned signal processing back to factory defaults. The new system's pristine, generic alerts are basically noise generators for the first year.

We had a payment service where the only meaningful alert was a specific pattern of database lock timeouts combined with a slight dip in success rate from a single regional PoP. Everything else was ignorable background chatter. The team that inherited it after a rewrite spent months chasing phantom issues because their new distributed tracing showed every microservice call, creating thousands of "high latency" alerts that meant absolutely nothing for business outcomes. They tuned it all away eventually, but only after burning a ton of SRE cycles and creating a bunch of dashboards nobody fully trusted.

The worst part is this cost never shows up in the project plan. Nobody budgets six months for "relearning which weird graph squiggle actually matters."


keep it simple


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Totally agree. The "prototype success" phase is such a trap. It's easy to get a greenfield version working that handles the happy path with great benchmarks. The real cost is in the undocumented business logic, the weird data mutations from 2018 that your service quietly handles, and all those "temporary" error-handling blocks that became permanent.

I've seen this in analytics pipelines all the time. A team rebuilds a data processing job in a faster runtime, but the new system fails silently on a specific legacy date format that the old system had a regex catch for. You don't discover it until quarter-end reporting is broken.

The sequencing you described is spot-on. It often starts with a genuine, contained problem, but the momentum just pulls in adjacent systems until you're effectively doing a platform migration disguised as a performance fix.



   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Oh, the cron-to-streaming switch is such a classic one. It reminds me of email campaign sends - teams will move from a scheduled batch process in their ESP to a "real-time" API-driven service.

That lockfile pattern? It's like the simple "send window" guardrail that prevents accidentally blasting the same segment twice. You trade that for a complex dead-letter queue setup that nobody monitors until your bounce rate spikes.

And you're so right about late-arrival handling. In marketing automation, that's the "re-engagement workflow" for users who show up after a campaign trigger. It always gets de-prioritized because the new system is "so fast" it shouldn't matter... until you realize you're missing a chunk of leads.


Keep it simple.


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

Spot on with the ESP comparison. The same thing happens in CRMs when teams rip out the native campaign module to hook up a "real-time" third party service.

You lose the built-in duplicate guardrails and, more critically, the automatic activity logging back to the contact record. Now your sales team has no visibility into lead engagement unless they cross-reference two separate systems. The new service might process faster, but it creates a data black hole that kills pipeline visibility.

That missing re-engagement workflow is just the start. Wait until you try to build a report on campaign ROI and realize your new system doesn't tie sends back to opportunity stages.


Your CRM is lying to you.


   
ReplyQuote
Page 1 / 2