Tried the new tool use feature on a few real pipeline tasks. The promise is solid: letting the LLM call my APIs or run scripts could automate a lot of tedious glue code.
My early tests were mixed.
* Good: Correctly formatted a curl command to hit our internal Jenkins API.
* Bad: Tried to execute a shell script directly (which it can't do) instead of just outputting the commands for me to run.
* Unreliable: It hallucinated parameters for a tool that didn't exist.
The core problem is trust. For CI/CD, you need deterministic outcomes. Right now, I wouldn't let it operate on our main branch without a human checking every single tool call. It's a cool beta, but not reliable enough for serious automation.
Anyone else stress-testing it with actual devops workflows? I'm curious if the reliability changes based on the complexity of the tool.
Your "trust" point is critical. I've been testing it for pipeline monitoring tasks, like auto-generating Spark job status checks. The reliability seems to degrade sharply with tool chaining.
For example, asking it to "Check the Kafka lag, then if it's above threshold, query the Spark UI" often leads to it inventing a synthetic monitoring tool instead of composing two separate API calls. It gets the first step right, then hallucinates a unified tool for the second.
I'm keeping a manual approval gate for any tool call that would mutate state, like triggering a job or writing to a warehouse. For read-only diagnostic workflows, it's already saving time, but only as a suggestion engine. I wouldn't give it direct execution rights yet.
What's your threshold for letting it run autonomously? Would a fully sandboxed environment change your stance?
You're spot on about the degradation with chaining. It's the classic marketing automation trap: the shiny demo works for one simple step, then falls apart in any real workflow.
> Would a fully sandboxed environment change your stance?
A sandbox just contains the blast radius, it doesn't fix the core problem. If the model hallucinates a tool, it doesn't matter if it's sandboxed - the task still fails and requires my intervention. The cost of that intervention (me checking why it didn't work) often negates the time saved.
My threshold for autonomy is a quantified, auditable success rate over hundreds of runs. If they can't provide concrete conversion numbers on tool call accuracy for multi-step operations, it's just a toy. Until then, manual gates on everything, read-only or not.
martech_auditor
>quantified, auditable success rate over hundreds of runs
This is exactly it. I've been running a controlled benchmark against a suite of 50 common pipeline tasks (API calls, SQL queries, log parsing). The single-step success rate is ~92%, which is promising. But for a chain of just three steps, it plummets to around 67%. The failure mode isn't just hallucination, it's often subtle state mishandling, like re-using a variable name from step 1 incorrectly in step 3.
A sandbox doesn't fix that, it just gives you a cleaner error log. The intervention cost you mentioned is real. My benchmark shows the average "debug time" for a failed 3-step chain is 4.5 minutes, which already eats the time saved from the two previous successful runs.
Until they publish and commit to improving those chain-length vs. success-rate graphs, it's a copilot, not an autopilot.
p99 or bust