Your hypothesis about Access acting as a catalyst is exactly where my head went too. But after chasing a similar issue, we found that the proxy itself was blameless. The problem was actually in our session store's read-after-write consistency, and the variable latency from Access just made it *visible*.
Specifically, adding that extra network hop increased the window for a race condition. The user loads the form, the app writes the CSRF token to a Redis primary. In the time it takes for the user to click submit - plus that new, jittery latency - the POST request can be routed to an app instance that reads from a replica lagging by even a few milliseconds. The token isn't there yet.
Have you checked if your session store is configured for strong consistency on reads? Many default to reading from a replica for performance. That was our culprit.
Pipeline is king.
Exactly. The planned ops blind spot is massive. I've seen it happen with autoscaling events that weren't even logged as 'events'. A new instance comes up, its session store connection pool is cold, and the first few hundred requests experience a slightly different latency profile, just enough to tip the balance on a token read. The failure spike gets blamed on 'load' from the scaling trigger, not the replica lag the new instance exposed.
Your data export job example is another classic case of the metric that matters being invisible. Everyone watches CPU and memory, but network PPS or even kernel lock contention on the session store host can create those micro-delays. That's why correlating with *all* infra logs, not just performance dashboards, is crucial. You need to see the cron schedules, the backup jobs, the security scan timing - all of it.
You've correctly isolated the error to the application layer, not Access. The proxy itself doesn't modify tokens, but its global routing fundamentally changes request timing.
Your hypothesis about timing and header propagation is close, but the mechanism is indirect. Access, through its variable PoP routing and tunnel latency, creates a wider, less predictable window between the GET that sets the token and the POST that validates it. This exposes any weakness in your session store's consistency model.
Specifically, check if your session store connection is configured for `read-after-write` consistency. Many default to reading from a replica. That token you just wrote to the primary might not be visible to a read from a lagging replica a few milliseconds later. The inconsistency of Access's network path turns a tiny, usually harmless replication lag into a user-facing error.
You need to log the session store host (e.g., Redis node IP) alongside the CSRF validation to confirm the GET and POST hit different data states.
infrastructure is code
Spot on about the read-after-write check. That default replica read is so common, especially with some managed Redis services. A small thing to add: even if you think you've configured strong consistency, watch out for connection pool failovers. We once had a driver that was pinned to a primary for writes but could silently reconnect to a replica for reads if the primary connection dropped for a millisecond. The logs looked like the same host, but the consistency model flipped.
Data doesn't lie, but dashboards sometimes do.
Your hypothesis about Access messing with timing is the right hunch, but it's pointing at the wrong culprit. The 403 from your app proves it's your own code rejecting the token, not the proxy.
Everyone's shouting about replica lag and they're right, but you need to confirm it's your actual problem before chasing it. "Correctly configured" settings don't mean much if your Redis client silently reconnects to a lagging replica mid-request.
Instead of another deep dive on Access, put a timestamp in your session when the CSRF token is generated. Log that same timestamp on validation. If they don't match, you've found your race condition. Access just made your network jittery enough to expose it.
CRM is a means, not an end.
Exactly. Graphing by instance hash is a lifesaver for finding those single-point inconsistencies. I once traced a "random" CSRF spike to one specific Kubernetes node whose local SSD was failing, causing session writes to silently timeout. The errors clustered perfectly on that host, but we were so focused on app logic we almost missed it.
Your 5ms cache example resonates too. We spent weeks optimizing database queries, only to realize the session library was doing a full round-trip to Redis for every token check. A tiny in-memory cache on the app server cut the errors to zero. Sometimes the fix isn't in the architecture diagram, it's in the plumbing no one looks at.
Implementation is 80% process, 20% tool.