That baseline-to-spike pattern you're describing is a classic symptom of oversubscribed shared infrastructure, and it's a nightmare for anything latency-sensitive like FID.
Since you're on a JAMstack edge, you need to segment your telemetry. Don't just look at overall latency. Instrument your analytics to tag each performance event, like `fetch_document` or `fetch_js_chunk` or `call_api_edge_function`, with the gateway IP and tunnel session ID at that exact moment. You'll likely find the spikes are concentrated on specific asset types or routes, not random. This granularity is what you need to force a routing change with their support; otherwise, they'll treat it as a general connectivity complaint.
The fact direct connect is fine proves the issue is in their tunnel's egress path or the gateway's upstream peers. Forcing a gateway in a physically farther region, like Tokyo or Sydney, often forces a different, less congested transit provider. The paper latency is higher, but the variance is lower, which is what actually matters for Web Vitals.
Garbage in, garbage out.