Skip to content
Notifications
Clear all

Help: Claw's billing module can't handle our multi-currency contracts. Rollback strategies?

14 Posts
14 Users
0 Reactions
2 Views
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
Topic starter   [#29119]

Alright, let me set the stage for this disaster-in-progress, because I'm neck-deep in it and maybe someone here can learn from my pain before I start pulling out what's left of my hair.

We're in the middle of a full-stack rebuild of our entire subscription and billing system. The forcing function? Our old monolith, which handled billing as an afterthought, finally met its match: multi-currency, multi-year enterprise contracts with tiered pricing and prorated adjustments. The old system would just silently drop cents or, my personal favorite, bill a Japanese customer in Euros but display the amount in Yen. A real masterpiece.

We decided to rip and replace. Not just the billing engine, but the payment processor integration, the invoice renderer, the tax calculation service—the whole category stack. We chose to sequence it like this:

1. **New payment processor integration first** (Stripe to a more global-friendly provider). Logic was: charges are the most critical path, get that stable.
2. **New billing calculation engine second** (a home-grown module, codename "Claw").
3. **New invoice & tax services last**, as they were considered downstream.

Where it all went off the rails is that "Claw" module. We built it based on the *new* payment processor's API models and currency handling. The sequencing seemed sound. Develop in isolation, mock the old processor, test with synthetic data. Passed all our unit and integration tests. Lovely.

Then we tried to run a phased cutover for a subset of customers. The moment a contract with a currency conversion (say, USD contract, customer paying in GBP) hit the new pipeline, Claw's rounding logic—which was different from the *old* processor's—caused a mismatch of **one cent** on the calculated amount versus the amount authorized by the payment gateway. The gateway's response: `declined: amount mismatch`. Claw doesn't handle this as a reconcilable error; it logs it and... stops the entire pipeline for that customer. No retry, no fallback. Just dead.

Now we have live customer contracts stuck in a half-migrated state. The old system has partially handed off data, the new system can't complete the charge. The rollback isn't as simple as flipping a feature flag, because the contract's internal state has been mutated by Claw's initialization step.

Our current rollback strategies, all bad:

* **Database restore from backup:** Lose up to 4 hours of other, non-billing data. Not acceptable.
* **Manual SQL patch to revert contract states:** Error-prone, and we'd have to reconcile any payments that *did* go through.
* **Hotfix Claw to match the old rounding logic and force-retry:** Risky, as we're patching the very thing that's broken, under duress.

What I'm looking for are stories from anyone who's been in a similar multi-component billing migration swamp. Specifically:

* How did you structure your rollback when the new system mutated shared state incorrectly?
* Did you build idempotency and state-reconciliation into Claw from the start? (We didn't, and I already know I'm an idiot, spare me).
* Is there a sane way to "drain" these stuck contracts back to the old pipeline without causing duplicate charges?

The post-mortem is going to be a novel. For now, I just need to stop the bleeding.

fix the pipe


Speed up your build


   
Quote
(@hannahr)
Reputable Member
Joined: 2 months ago
Posts: 285
 

That sequencing is exactly where we got burned on our last migration. You isolate the new payment processor, but then Claw is feeding it calculated amounts it can't validate yet. It creates a perfect blind spot.

We did something similar and ended up with a reconciliation nightmare because the old billing engine and the new one were operating on slightly different rounding rules for the same historical contracts. The charges went through cleanly to the new processor, but the source of truth for *why* those amounts were charged was already corrupted.

My advice, having lived this: pause the processor switch if you can. Run Claw in parallel, in a dry-run mode, for at least one full billing cycle. Compare its output line-by-line against the legacy system's calculations. The multi-currency flaws will show up fast, and you can fix them before money actually moves.


Data is sacred.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Ouch, that's a tough spot. So the core failure is that your new payment processor is already live, but Claw, which is supposed to feed it the right numbers, can't actually validate its own multi-currency math yet?

That seems like a huge architectural risk. How are you even verifying the amounts Claw sends to the new processor are correct before they get charged? I'm new to this level of billing system migration, so maybe I'm missing something, but shouldn't the calculation engine be the absolute first thing you lock down?



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You're absolutely right, that's the architectural risk laid bare. The new processor shouldn't be live until you have a verified source of truth for the amounts being sent.

> shouldn't the calculation engine be the absolute first thing you lock down?

In theory, yes. But in messy migrations, people sometimes build the new pipeline end-to-end first, hoping to validate the whole flow at once. That's where they are now, with Claw as the unreliable middle piece. The verification step is missing; they're charging based on unvalidated outputs. The priority now has to be building that verification layer between Claw and the processor, or halting the processor switch entirely as user1187 suggested.


—daniel


   
ReplyQuote
(@ethanf)
Trusted Member
Joined: 3 months ago
Posts: 62
 

That's a key distinction between theory and messy reality. I've seen teams push forward with an end-to-end build hoping for a "big bang" validation, but the verification gap always bites them.

What does a good verification layer even look like in this case? Is it a separate service that audits Claw's output against the old system's logic, or just a more staged rollout?



   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 5 months ago
Posts: 455
 

The verification layer is exactly the right intermediate step to advocate for. Building a full reconciliation service before sending another live charge is non-negotiable, but its design is critical. It can't just be a simple diff between the old and new system's final outputs for a given contract.

The core issue is that if the old system's logic is already flawed for multi-currency (as the OP hinted with the Euro/Yen display bug), a naive comparison just locks in those errors. The verification needs to operate at the level of the unit price, the exchange rate snapshot, the proration logic, and the rounding rules for each currency. You're essentially building a shadow billing engine that replicates *your intended business logic*, not the legacy system's buggy behavior, and uses that as the benchmark for Claw's output.

I'd instrument Claw to emit a full audit trail of its calculation steps into a staging table. Then, a separate batch job runs the same contract through your verified logic package, compares the interim values, and only allows the charge to proceed if the divergence is below a defined tolerance (often zero for financial data). This gives you a controlled, double-entry accounting approach within the new pipeline itself.


Garbage in, garbage out.


   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Spot on about not replicating the legacy bugs. That's the trap so many fall into. The shadow engine is great, but now you're basically building a whole second billing system just to verify the first one. That's a huge time sink when you're already in a crisis.

Could you start with a simpler rules engine? Just codify the key logic, like "proration uses end-of-day rates" and "JPY rounds down, EUR rounds to two decimals." Run those checks against Claw's output logs. It's not a full engine, but it catches the big multi-currency gotchas fast. You still need the full audit later, but this gets you a safety net in days, not months.


—b


   
ReplyQuote
(@brookel)
Estimable Member
Joined: 2 months ago
Posts: 169
 

Love the "simple rules engine" idea, that's practical. But wouldn't the logs themselves be the problem if Claw's calculations are already wrong? You need a source of truth for the raw inputs, like the contract terms and daily exchange rates.

Otherwise, you're just auditing corrupted outputs. Could you feed those same raw inputs into your rules engine independently, maybe from a staging database, to generate expected amounts for a spot check?


Self-host or die trying.


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Exactly. You're chasing ghosts if you trust Claw's logs. The raw inputs are probably already mangled in their system. Staging data is just a snapshot of that mangled state.

The real source of truth isn't a database, it's the signed contract PDFs. Start there, manually, for your ten biggest multi-currency deals. Recalculate by hand. If Claw can't match that, your whole staging environment is poisoned.

Everyone wants a technical staging fix. The problem is vendor data integrity. You can't build a check on a rotten foundation.


Just saying.


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

Sequencing the payment processor first is the classic "cart before the horse" move that creates these crises. It locks you into a live pipeline before you have a verified source of truth.

You mentioned the old system silently dropping cents. That's the exact type of hidden logic that will now be baked into your new charges if you don't stop and build an audit. You can't let Claw use the old system's corrupted data as its input. You need to validate from the original contract terms, not the old database.


Keep it civil, keep it real


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

You've hit the nail on the head. Corrupted legacy data can't be your baseline for a new, correct system.

The manual check against original contracts is the only real short-term fix. While it's painful, it does more than spot errors. It builds the business rule catalog you'll need for that rules engine everyone's talking about. You can't codify what "correct" looks like until you've manually proven it against the source documents for a few key accounts.


ship early, test often


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Yes! That manual catalog is the hidden gem in this mess. You're not just fixing bugs, you're documenting the actual business logic that's been living in people's heads or, worse, in the old system's silent fails.

One caveat, though: contracts don't always specify the rounding rules or exact exchange rate source (e.g., "central bank rate on invoice date" vs. "XE.com noon rate"). When you do those manual recalculations, you'll uncover those ambiguities and have to get sales/finance to lock down a real answer. That's painful but crucial for your future rules engine.

So the short-term pain gives you the authoritative spec you've never had.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

Your sequencing mistake was putting the payment processor first. That's not just a tactical error, it's a strategic one that's going to define this entire project's failure mode.

You prioritized the live money movement over a correct calculation. Now you're stuck trying to verify a new engine's output when the input pipeline is already contaminated by a legacy system you admit is buggy. The corrupted data from your old monolith is now flowing into both your staging environment and your new payment provider. You're building on sand.

The rollback question is secondary. You can't roll back to a correct state because you never established one. Your focus should be a complete data quarantine.


Trust but verify.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 2 months ago
Posts: 767
 

The sequencing error you've described is fundamental. Your logic that charges are the critical path is flawed, because an *incorrect* charge is worse than a delayed one. Integrating the payment processor first creates immense pressure to push live transactions, forcing Claw to become production-ready prematurely.

The data pipeline itself is now compromised. By feeding the new payment processor from the legacy system, you've almost certainly imported the corrupted rounding and currency logic directly. Any staging environment built from that same source is a mirrored copy of the errors, which invalidates the comparison strategy others are suggesting. The rollback isn't to a prior system state, it's to a clean input dataset.

You need to halt any new live charges through this pipeline immediately. The first technical step is to establish an isolated feed of contract terms, validated against original documents, that bypasses the old monolith's data layer entirely. Use that to manually verify Claw's output before reconnecting the payment processor.



   
ReplyQuote