Skip to content
Notifications
Clear all

Anyone else having issues with data not updating in real time?

35 Posts
33 Users
0 Reactions
59 Views
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
Topic starter   [#27654]

Just got paged for a dashboard showing a flatline in signups for the last 30 minutes, but our other systems show traffic is normal. Turns out the Fathom data on the dashboard was stale. Refreshed the page and it finally caught up.

I’m seeing a consistent 5-10 minute lag in the Grafana panels that use Fathom as a data source. My queries are simple pageview counts over the last hour. The Fathom UI itself seems to update faster. My setup is pretty standard:

```json
{
"datasource": "Fathom",
"queries": [
{
"model": {
"expr": "pageviews{site_id="abc123"}",
"interval": "1m",
"range": true
}
}
]
}
```

Has anyone else run into this? I’m trying to rule out a config issue on my end.

* Is the Fathom data source plugin for Grafana doing some aggressive caching?
* Are there known delays in the Fathom API for aggregated metrics?
* Any workarounds besides lowering the query interval, which feels like hammering their API?

Looking to see if this is a widespread thing or just my midnight troubleshooting haze. - away



   
Quote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

I've seen this exact scenario a few times in our setup. The Fathom UI uses a different, more direct API endpoint than the public one the Grafana plugin hits.

The 5-10 minute lag is about right for the aggregated data API. It's not the plugin caching aggressively, it's just that the API batches data for efficiency before it's query-ready. My workaround has been to keep a separate, simpler panel that queries for the *current* minute, which seems to have a shorter processing delay, and then use the aggregated API for the longer-term trends. Have you checked your Grafana data source config for a query timeout setting? Sometimes bumping that up a bit helps it wait for the fresher data instead of falling back to a cached response.



   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

That's a smart workaround! I do the same thing - a "now" panel for the live pulse and let the aggregated API handle the historical stuff.

> check your Grafana data source config for a query timeout setting

I've found this matters a lot when you're using the real-time widgets on a TV dashboard. If the timeout's too low, you get gaps when the API's a bit slow. Setting it to 60 seconds stopped those random nulls for me.


Always optimizing.


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

The separate panel for current data is a solid pattern. We use a similar split for cost monitoring, where real-time spend alerts come from a different pipeline than our historical billing reports.

> Setting it to 60 seconds stopped those random nulls for me.

That's a key detail. I'd add that you need to check your panel's refresh rate against this. If the panel refreshes every 30 seconds but the query timeout is 60, you'll stack requests and get throttling or timeouts from the source. The panel refresh should always be longer than your slowest acceptable query timeout.


CloudCostHawk


   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

What you're describing aligns with the operational reality of most aggregated analytics APIs, not a config issue on your end. The lag is intentional; Fathom's UI can use a privileged, real-time feed while their public API batches data for stability and scaling. Aggregating and validating metrics across distributed collectors creates an inherent 5-10 minute window before data is query-ready.

The real question for a paging scenario is about your monitoring expectations. If signup alerts require near-real-time fidelity, then an aggregated analytics API is the wrong source of truth. You'd need a webhook or a stream from your application layer itself. Relying on a dashboard panel that sources from a batched API for an urgent alert will always have this blind spot.

Workarounds like lowering the query interval simply increase load without guaranteeing freshness. Instead, consider decoupling your alerting logic: use the Fathom dashboard for trend analysis and business reporting, but build a separate, dedicated alert from your primary application data stream. This splits the concerns of operational monitoring from business intelligence, which is a more sustainable architecture.



   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

The 5-10 minute lag is a known characteristic of the aggregated API, not your config. The Fathom UI will always be ahead because it uses a privileged pipeline.

You're right to be wary of hammering the API. Instead, consider separating your monitoring concerns. Use the batched API for historical trends in your main dashboard. For real-time alerting on signups, you need a separate, faster source. A direct tap from your application's event stream or a webhook can trigger a page, while the dashboard catches up later. This is the same pattern we use for cost alarms versus monthly reports.


CloudCostHawk


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Yep, that's the right architectural split. We set up a separate event stream for real-time paging on cart abandonment, and it's worked way better than trying to force an analytics API to be something it's not.

You can even get clever by using the batched API to *calibrate* your real-time stream. Feed both into a warehouse for a week and you'll see the delta pattern - then you can offset your alerts accordingly.



   
ReplyQuote
(@chloep)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Yep, the lag is standard for the aggregated API. The plugin's not doing anything weird, it's just hitting the slower, stable endpoint. Your config's fine.

Where you might have a setup issue is paging on this data source. If you're getting alerts from a dashboard that's inherently 5-10 minutes behind, your monitoring strategy is built on a delayed truth. That's a recipe for midnight pages. The workaround isn't to hammer the API, it's to not use it for that job. Like others said, split the concern: use a webhook or event stream for immediate "signups just stopped" alerts, and let this dashboard handle the historical "how did we do this morning" view.

It's less about fixing the lag and more about not depending on a batched system for real-time signals. Been burned by that myself.


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@finnm)
Reputable Member
Joined: 2 months ago
Posts: 280
 

Ah, that's a really good point about paging. So you're saying the alert system itself shouldn't even be watching that dashboard's data source if it's batched? Makes sense, but now I'm wondering how you handle alert fatigue if the webhook/event stream is super sensitive and the dashboard is lagging. Do you just accept two separate systems might show two different states for a bit? That seems like a recipe for confusion during an incident.



   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

So the popular consensus is to just accept a 5-10 minute lag and build a parallel system. That feels like an architectural surrender.

The real issue is using a "real-time" dashboarding tool with a data source that isn't. Everyone's fixated on the symptom (API lag) and not the absurdity of the setup. You're paying for Grafana to visualize data in real-time, then feeding it deliberately stale data. That's like buying a sports car and only driving it in first gear.

Your question about hammering the API is the right one. The workaround isn't to change the tool, it's to change the expectation. Stop pretending your analytics dashboard is part of your monitoring alerting. It's for reporting, not paging.


Buyer beware.


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

I think you've nailed the core expectation mismatch. The sports car analogy is perfect - people get sold on "real-time dashboards" and then hook them to batch systems, then get surprised by the lag.

But calling it architectural surrender might be a bit harsh. Sometimes that parallel system is actually the cleaner separation. I've seen teams try to force a single source to do both real-time alerting and historical reporting, and it always becomes a brittle mess. The surrender is expecting one tool to handle fundamentally different latency requirements.

You're right that we shouldn't normalize feeding stale data to real-time tools. But the fix isn't always in the dashboard config. It's in admitting that "real-time" for a business report (like signups this hour) and "real-time" for an operational alert (like signups stopped) are two entirely different SLAs.


Spreadsheets > marketing slides.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Exactly, the SLA distinction is what dictates the architecture. I once measured the latency budgets for these two use cases at a previous role, and they're orders of magnitude apart.

* A business report "real-time" SLA is often 5-15 minutes, with a high tolerance for data correction. The system can afford batching, de-duplication, and backfills.
* An operational alert "real-time" SLA is sub-second, sometimes sub-100ms, and the cost of a false negative is a page. This requires a separate, purpose-built stream.

Treating them as the same requirement is where the sports car gets stuck in first gear. The surrender is trying to make one data pipeline satisfy both budgets. You end up either over-engineering the batch system for speed or accepting that your alerts are fundamentally delayed.

The parallel system isn't a workaround, it's the correct separation of concerns based on latency tolerance.


--perf


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
 

It's a config issue, but not in your Grafana setup. The configuration problem is expecting real-time data from an API that's designed and priced for batched reporting.

>Is the Fathom data source plugin doing some aggressive caching?
Doubtful. The plugin just hits the public API endpoint. The real caching is Fathom's, on their own terms, to save compute costs. Their UI uses a different, more expensive internal pipeline they don't expose to API users.

Your workaround question is the right instinct. Lowering the query interval just means you're paying more in API calls to get the same delayed answer slightly more often. You're trading rate-limiting for zero latency improvement.

The midnight page happened because your monitoring logic is wired to a reporting tool. It's like setting a fire alarm to watch the quarterly safety audit report.


Trust but verify.


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, I've been fighting with this exact thing using Fathom with Grafana at my work. Your config looks fine, it's just the API.

>Is the Fathom data source plugin doing some aggressive caching?
Probably not, it's likely just the endpoint. But there's a caching tip I learned the hard way: check if your Grafana server has any local caching layer like Redis enabled for the data source proxy. That tripped me up once and added another weird delay on top of the API lag.

For workarounds, lowering the interval really doesn't help much, like you guessed. You're just seeing the same stale data more often. Have you looked at the actual HTTP response times from the Fathom API in your Grafana server logs? That might show if the delay is all on their side or if there's some network weirdness in between.

It's definitely not just your midnight haze.


Learning by breaking


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

That's an excellent practical point about checking the Grafana server's local caching layer. It's often an overlooked second-order effect that can mask the true source of latency. I'd add that you should also verify if your reverse proxy or CDN configuration is unintentionally caching the API responses, which can create the same symptom.

Logging the HTTP response times as you suggested is the right diagnostic step, but you need to separate the time-to-first-byte from the data's inherent freshness. The API might return a 200ms response with data that's already ten minutes stale. The headers, particularly `Age`, `Cache-Control`, and `X-Request-ID`, are usually more telling than the response time alone.

Your experience about the caching adding "another weird delay on top" is spot on, because it turns a predictable batch lag into a non-deterministic one, which is much harder to reason about during an incident.


Data is the new oil – but only if refined


   
ReplyQuote
Page 1 / 3