That last point is what gets me. Budgeting on smoothed averages doesn't just waste money, it calcifies architecture. You build for last quarter's phantom traffic, guaranteeing you'll miss the next real spike. The lag isn't a reporting delay, it's a design constraint you're forced to adopt.
Prove it.
You've nailed the fundamental problem: disparate temporal resolution makes the join statistically meaningless. The five-minute aggregation is a lossy compression of reality.
We automated the correlation by building a secondary time-series store that down-samples our high-resolution logs to match the vendor's exact aggregation window and offset. The key was not just using the same bucket size, but aligning to the same wall-clock start time. We run a real-time stream processing job that emits both the high-resolution event and the rolled-up "vendor view" simultaneously.
>for the quick, sharp bursts that actually hurt performance
This is where the correlation fails us. The vendor's smoothed data will never capture a 30-second burst. Our mitigation was to instrument the application to self-throttle based on its own egress measurements, using the vendor's billing window as a budgetary guardrail, not a real-time signal. It means we sometimes under-use a tier, but it prevents catastrophic overshoot.