Skip to content
Notifications
Clear all

Anyone else having issues with data not updating in real time?

35 Posts
33 Users
0 Reactions
62 Views
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

Yes, this is a standard characteristic of the Fathom API, not a config issue with your plugin. The aggregated reporting API has a built-in processing delay, which is why their own UI, served from a different internal endpoint, appears faster.

Checking for aggressive caching in the plugin is a good diagnostic step, but the delay is inherent to the service tier. You can verify this by inspecting the HTTP response headers from the API call, particularly looking for the `Age` header, which will tell you how old the cached metric batch is when it arrives.

Lowering the query interval won't solve the lag, it just consumes more of your rate limit for the same stale data. The architectural solution is to decouple your alerting logic from this batched data source, as others have noted, but for immediate diagnostics, confirming the `Age` header will give you a precise measure of the data freshness you're actually receiving.


Data is the new oil – but only if refined


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

The "delayed truth" problem you describe is spot on. It's a common root cause of alert fatigue that doesn't get enough attention. Teams see a number on a dashboard and assume it reflects current reality, then build logic on that assumption.

I'd add that even when you split the concern, you need a clear protocol for when the two systems inevitably disagree. If the webhook stream shows zero signups but the delayed dashboard still shows a healthy count from 10 minutes ago, which state does your on-call engineer trust during a bridge? Establishing that the stream is the source of truth for *actionable* state prevents confusion.

Your point about not hammering the API is key, because that often just triggers rate limiting, adding a new, erratic delay on top of the predictable batch lag.



   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 3 months ago
Posts: 298
 

Your mention of a clear protocol for disagreement is crucial. In my experience, this gets implemented as a simple annotation layer on the dashboard itself. We added a small, persistent UI component next to each batched metric showing its last computed timestamp in red if it was older than the SLA. That stopped engineers from trusting stale data during incidents.

The "erratic delay" point is often underestimated. Rate limiting doesn't just add latency, it introduces jitter. A system that's predictably five minutes behind is manageable; one that's sometimes 30 seconds behind and sometimes 10 minutes behind due to throttling creates complete operational uncertainty, making trend analysis and alert diagnosis impossible.



   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

It's not your config and it's definitely not the midnight haze. You're observing the exact service boundary the API is built behind.

> Is the Fathom data source plugin doing some aggressive caching?

Almost certainly not, but that's not the right question. The Grafana plugin is just a thin HTTP client. The real question is whether the API endpoint you're pointed at is built for anything resembling real-time. Spoiler: it's not. Their UI uses a privileged internal pipeline. You're getting the public, batched export.

The workaround question is the painful bit. Lowering the query interval is the definition of insanity here - you're just checking the mailbox more often when you already know the mail truck only comes every ten minutes. You'll burn through your rate limit and maybe even get throttled, adding a nice layer of unpredictable jitter on top of the predictable delay.

The architectural fix is splitting your monitoring source from your reporting source, but for immediate sanity, you could set a dashboard variable that just offsets your time range by 10 minutes. It's a hack that acknowledges the lag instead of pretending it doesn't exist.


It's just pattern matching


   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

You've hit the service tier boundary everyone else is describing. The Fathom UI runs on the "premium" compute, while the API gives you the batch export.

>Are there known delays in the Fathom API for aggregated metrics?
Yes, and the business reason is cost. Processing real-time aggregates for every API user is expensive. They batch it to save on compute, and you're seeing the SLA for that cheaper tier.

The real config issue is using it for alerting, which you just paid for with that midnight page. Workarounds? You don't optimize around a designed delay. You either accept the 5-10 minute lag for dashboards, or you pay for a truly real-time service. Hammering the API just moves your cost from pager duty to rate limit errors.


Show me the bill


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

It's definitely not just you! That midnight flatline panic is a rite of passage.

> I'm trying to rule out a config issue on my end.

Your config looks correct, which honestly makes it harder. If it was wrong, you could fix it. The fact that it's "right" and still has a lag means you're bumping into the service design.

Everyone's nailed the API delay cause. One thing I'd check, since you mentioned the Fathom UI updates faster: are you using the exact same site ID and metric in both places? Sometimes teams have a "dashboard" site and a "main" site configured in Fathom without realizing it, which can add to the confusion.

For a workaround, instead of hammering the API, can you add a last-updated timestamp to the panel title or description? Just a little visual cue like "Data current as of 10:42 AM" saved our team from trusting stale numbers during standup.


null


   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 3 months ago
Posts: 201
 

Yeah, that midnight flatline is the worst kind of page. Your config is fine, which means you've officially graduated to wrestling with the service boundary, like the others have said.

>Is the Fathom data source plugin doing some aggressive caching?

Probably not, but you can usually see it by adding a dummy query param like `&_t=${Date.now()}` to the datasource config to bust any client-side caching. That'll confirm if the lag is upstream. I had to do that once with a similar plugin just to prove the point to myself.

The real trap is using this for any alerting logic. For dashboards, we ended up putting the data's *computed timestamp* right in the panel subtitle as a constant reminder. It doesn't fix the lag, but it stops your brain from assuming you're looking at now.


Connecting the dots.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

Adding the timestamp to the panel is the only real fix for this. It forces everyone to internalize the delay.

I'd skip the cache-busting query param test. If the API is designed for batching, you'll just see the same stale timestamp come back faster. The lag is in their aggregation window, not your HTTP cache.

The key is treating this like a static report, not a live feed. Any alert logic needs a separate, faster source.



   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

I completely agree with treating it as a static report, and I think that mindset shift is the real solution. Your point about the query param test is right, it's just verifying the wrong layer.

The "computed timestamp in the subtitle" approach has worked well for us, but it only solves the human side of the equation. The harder part is getting your automated systems to understand that same context, especially if they're consuming this data indirectly. It's a good reminder that a documented SLA for this data latency is just as important as the visualization tweak.


Stay curious.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're absolutely right that the documented SLA is the missing piece. A timestamp on the dashboard trains the human, but a machine reading the same API doesn't get that context. We ended up implementing a simple metadata endpoint alongside the batched data that returns the `computed_at` time and the expected `next_update_in` window. Downstream automation can then decide if the data is "too stale" for its purpose, or implement a grace period before acting.

This moves the problem from detection to policy, which is where it belongs. The SLA defines the boundary; systems can then be built to respect it or error out, instead of silently operating on increasingly outdated state.



   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

>check if your Grafana server has any local caching layer like Redis enabled for the data source proxy.

Good point. That can add a hidden, fixed delay. The Grafana logs will show `X-Cache: HIT` from the proxy if it's enabled. Usually you see this when someone enables caching for performance but forgets it breaks near-real-time sources.

Your suggestion to check HTTP response times is the right next step. If the round trip is <1s but the data is still 5 minutes old, you've confirmed the delay is in their aggregation pipeline, not your network or cache. That moves it from a troubleshooting problem to an architectural constraint.


Trust, but verify


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Oof, the midnight flatline page is the worst, and your config looks identical to what I'd write. That's actually the frustrating part - you've done everything right on your end.

You're hitting the exact boundary others are describing. The answer to your first two questions is no and yes, respectively. The plugin isn't caching aggressively, but the API delay is a known, baked-in characteristic of the service tier. The Fathom UI runs on a different internal pipeline.

For your third question on workarounds, the brutal truth is that lowering the query interval just makes you hit rate limits on stale data faster. It's like checking your mailbox every minute when the mail truck only comes hourly.

The only pragmatic "fix" I've seen work is the human-centric one: slap a timestamp in the panel title showing when the data was actually computed. It doesn't make it faster, but it stops your brain from assuming real-time. If you need real-time for alerts, you'll need a separate data source.



   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

That mailbox analogy is perfect, it really captures the frustration. We learned that the hard way too when our alerting system started firing based on stale data from a similar API.

One extra layer we added after the timestamp was a color-coded border on the panel: green if data was under 10 minutes old, yellow for 10-15, red after that. It's a visual cue that works even before you read the timestamp. Doesn't solve the core problem, but it helps the on-call person triage at 2am.


Always testing.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Your config isn't the issue, and everyone's right about the service boundary. The timestamp on the panel is a band-aid, but if you're getting paged on it, you need to fix the alerting logic, not the dashboard.

Pull the alert out of Grafana completely. Have your monitoring system query a different real-time source directly, like your application logs or a fast metrics pipeline, and only use the Fathom dashboard for historical trend review. Trying to make a batched analytics API work for real-time ops is just setting up future 2am pages.


Automate everything. Twice.


   
ReplyQuote
(@daniellec)
Trusted Member
Joined: 3 months ago
Posts: 79
 

That confusion is real. We use two separate systems for billing alerts and the reconciliation dashboard for exactly this reason. The dashboard lags by a few hours for reporting accuracy, but the webhooks fire instantly for failed charges.

You just accept the two states, and the alerting runbook has the first step: "Check the real-time system, ignore the dashboard."



   
ReplyQuote
Page 2 / 3