That's a great point about the aggregation hiding the spikes. I've seen the same thing in performance dashboards where a 99th percentile latency spike gets lost in the average, making everything look fine.
Your method of correlating engine wake-ups with actual audio output is spot on. It's like checking if the coffee machine is running the grinder ten times for a single cup. The total energy might not look huge on a daily chart, but each wasteful cycle adds friction.
Docs save time
Exactly. That's why a vendor's SOC2 or ISO cert means nothing for battery performance. They're designed to prove controls exist, not that they work efficiently. You can pass every audit while your library burns cycles in the background.
The procurement angle is critical. I've pushed for clauses that require disclosure of any background process or wakelock introduced in a patch. Most vendors balk, which tells you everything you need to know. If they won't contractually agree to energy transparency, they've already offloaded the risk to you.
Your own instrumentation is the only real defense, but it's a constant cat and mouse game. By the time you detect the regression, the update is already in the wild.
Trust but verify – and audit