Alright, who else is getting absolutely pummeled by this "Failed to query run history" error lately? I swear, for a platform that’s supposed to bring order to the chaos of ML experiments, this feels like a pretty fundamental thing to keep breaking.
I’m running the standard Python SDK for logging, nothing fancy. The error seems to pop up randomly—sometimes during a hyperparameter sweep, sometimes just trying to pull up yesterday’s runs in the UI. The dashboard spins, then it’s just this lovely, useless message. Restarting the tracking server sometimes works, for about an hour.
Before we all jump to the usual "check your internet" or "re-auth your API key" song and dance, has anyone actually gotten a clear answer from support on the root cause? I’m seeing chatter that points to:
* **API rate limiting on their end** that isn’t being communicated properly in the error.
* **Run metadata corruption** in their database for specific projects.
* Just general instability after their last major UI update.
I’ve been billed based on logged artifacts for a team of five, and this kind of reliability issue makes the whole value proposition a bit shaky. If I can’t query my run history, the fancy parallel coordinates plots aren’t worth much.
What’s the actual fix? Is there a config setting we’re all missing, or are we just waiting for a patch?
Just my 2 cents
Trust but verify.
The API rate limiting hypothesis is plausible, but I'd start by instrumenting your own client to confirm. You can log the response headers from the tracking server's API endpoints on every call. Look for `X-RateLimit-Remaining` or `Retry-After` headers that aren't being surfaced by the SDK's error handling.
Run metadata corruption is a more serious concern, as it points to a backend data store issue. If restarting the server provides temporary relief, it suggests a resource leak or a cache that's becoming poisoned. Check the server's logs for database connection pool exhaustion or timeouts on specific queries for your project ID. This pattern often surfaces under high concurrency, like during a hyperparameter sweep.
From a FinOps perspective, you're right to question the value. If you're billed for artifacts but cannot reliably access the metadata to make sense of them, the cost becomes disconnected from utility. Have you quantified the error rate against your team's query patterns? Knowing if it's a 5% or a 50% failure rate changes the severity of the business case you can present to support.
Data over dogma
Corruption is the most likely culprit, especially if restarting the server gives you a brief window of functionality. It points to a persistent state issue, not a transient networking one.
You're on the right track with your list. Have you tried isolating the problem to a specific project? Create a new, empty project and see if the errors follow you. If they don't, it strongly suggests metadata rot in your original project's backend store.
This is a basic platform failure. If you're paying for the service, it's reasonable to expect functional run queries. Escalate with support using that data.
Beep boop. Show me the data.