Hi everyone! 👋 I've been lurking here for a while, trying to absorb all the incredible info. Honestly, it's a bit dizzying trying to keep track of all the different model providers and their performance.
Our small team is finally experimenting with a few LLM APIs for some internal tools, and we quickly realized we needed a better way to see what was actually happening. We were just checking average speeds, but then we'd get these random, really slow responses that messed up the user experience. Someone on the team built a simple dashboard to track P99 latency specifically.
It's nothing fancyβbasically pulls data from our logging, groups it by provider and model, and shows the 99th percentile latency over the last 24 hours on a chart. We're tracking OpenAI, Anthropic, and one other provider right now. It's already eye-opening! The averages can look similar, but the P99 tells a very different story about which one feels more consistent under our (admittedly light) load.
I'm curiousβfor those of you who track this stuff, is P99 the right metric to focus on for latency in a user-facing app? What other key numbers do you put on your dashboards for provider reliability? And does anyone have a good, simple way to correlate latency spikes with specific types of prompts or times of day?
✌️ annie