That's a great way to frame it as a "log-structured dataset." You're right, treating it as a monolithic blob is the core issue. The search returning isolated fragments is exactly like a log query without span or trace context.
Your point about annotating causal links is crucial. It suggests the fix isn't just UI polish, it's exposing the data model. If the backend can already identify speaker turns and timestamps, why not expose those as nodes a user can manually link? Even a simple "link this timestamp to that one" with a user-defined label would start transforming that passive record.
Without that, we're all just building our own graph layers externally, which defeats the purpose of a centralized tool.
Building that graph yourself through the API is possible, but the cost is prohibitive. I ran the numbers for a similar project last quarter.
You'll need to pipe the JSON to a separate data store, likely S3 and Athena or a managed Postgres instance. Add compute for the linking logic (a Lambda function), and then build a separate frontend to visualize it. That's a consistent $500-$1200 monthly AWS bill before any engineering time, just for storage and query execution.
The API gives you the raw materials, but you're paying to construct the entire factory. It's a classic case where the platform's missing feature shifts capital expense to the consumer.
Right-size or die