The off-peak minute trick might work, but it's a brittle stopgap. If their pool is global, your "off-peak" is someone else's peak. The real issue is their silent failure. Any system that swallows API errors instead of surfacing them is fundamentally broken for debugging.
Your "standard setup" is their whole sales pitch. The fact it's failing means the pitch is wrong.
No logs isn't an oversight, it's a feature. If they showed you the 429 from TikTok's API, you'd know they're using a shared key pool. That would raise questions about their infrastructure costs versus what they're charging you.
Even if you re-authenticate a hundred times, you're just getting a new seat on the same overloaded bus. The driver still isn't telling you why it keeps breaking down.
Read the contract
The silent logs are the real problem. You're right, without the actual TikTok error code you can't fix it yourself.
I've gotten it stable-ish by avoiding their schedule queue entirely. I only use "Post Immediately" now, and even then I manually trigger during off-peak hours (late night PST). My success rate is maybe 80%, but that's still not good enough for a true hands-off workflow.
The shared API key theory others mentioned fits. You could try setting your scheduled posts for odd minutes (:12, :37), but it's a band-aid. The core issue is their lack of error transparency.
Automate the boring stuff.
That point about the "lottery system" is exactly what makes this so hard to evaluate for procurement. When a feature is this opaque, it's impossible to calculate the real TCO or SLA risk.
You're right that error pass-through is a basic requirement for any reliable integration. If we can't see the failure mode, we can't create a contingency plan or even properly document the downtime for our own reporting. It turns a technical issue into a business continuity one.
Has anyone found that the failure rate changes between Opus's pricing tiers? I'm wondering if the shared pool is a resource allocation strategy they're applying across the board.
Your setup is textbook, which isolates the failure to Opus's handling layer. The lack of actionable logs is the critical flaw; without the TikTok API error code, you can't determine if it's a rate limit, a permissions refresh, or a content validation issue on their end.
To your request for systematic data, the most consistent pattern I've observed is failure clustering around top-of-hour schedule triggers, supporting the shared key pool hypothesis. Have you attempted to correlate your failures with specific times, particularly UTC hour rollovers? That would indicate a quota exhaustion cycle.
One test you didn't mention: does manually triggering the exact same clip via "Post Immediately" after a schedule failure succeed? If so, the problem is almost certainly their queue management and not your asset or connection.