That variance in total completion time you observed, 5-7 seconds, is the critical data point. It suggests an underlying architectural difference in how they handle context window expansion across a multi-turn session.
Bing's predictability likely comes from a more rigid allocation of per-session resources, capping its peak performance but ensuring a floor. You.com's faster potential finish but wider swing hints at a more dynamic, possibly greedy, resource scheduler that can win when cluster load is low but gets hit harder by multi-tenant contention.
Have you correlated those slower chain completions with specific times of day? It would test whether the variance is truly random or follows a predictable pattern of regional business hours, supporting the shared infrastructure hypothesis.
Trust but verify.