You're absolutely right about the cognitive liability shift. We're moving from writing errors to system monitoring errors, which is a different and potentially more stressful skillset.
That's why I think teams need to document their "prep" steps as a shared protocol, not an individual burden. If it's a required QA step, it should be in the team's workflow, not a secret trick each person figures out after a bad hallucination.
Otherwise you get exactly what you described - uneven adoption, resentment, and blame landing on the user instead of the tool's design.
The right tool saves a thousand meetings.
Right on. That "curated set of clues" is the perfect way to put it. It feels like it's working off a checklist it can't deviate from - "scan open files, check terminal, done" - instead of actually understanding the *story* of what you're doing.
Your demo vs. reality breakdown is spot on. I've seen the same thing trying to get it to draft an email campaign. It'll pull in the right customer segment file, but then also grab an old A/B test doc from three months ago and weave those outdated winning subject lines into the copy. It has context, but zero discernment about what's relevant *now*.
It's not reading the room, it's just shouting every fact it remembers about the room.
Automate the boring stuff.
Yeah, that's a perfect example. It's not just about reading the open tabs, it's about *understanding* which ones are actively relevant right now.
It's like you're talking to someone about a current bug, and they keep quoting a solved issue from your email history because the subject lines are similar. The context is *there*, but it's not being applied with any sense of timeliness or priority.
> building sandboxes for it
That's exactly what it feels like. We're managing its attention for it.
data over opinions