Skip to content
Notifications
Clear all

Unpopular opinion: The rush to add 'agents' is making these tools overly complex.

18 Posts
18 Users
0 Reactions
100 Views
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
Topic starter   [#23625]

The current feature race in AI coding assistants feels increasingly misguided. Every major player is scrambling to integrate "agentic" workflows—multi-step planning, tool use, and autonomous execution—directly into the core editor experience. While the research is fascinating, the practical implementation is adding a layer of abstraction and non-determinism that actively harms developer productivity for a large class of common tasks.

My primary issue is that these agent systems are optimized for a demo, not for a daily workflow. They turn a simple, predictable request into a black-box process with opaque reasoning, high latency, and unpredictable failure modes. Consider a straightforward task: "Add error handling to this function and log to our observability stack."

**Without Agent Overhead:**
A competent non-agentic assistant provides a direct code block edit, perhaps with a brief explanation. The transaction is fast, the scope is clear, and I can immediately evaluate the output.

**With Agent Overhead (as currently implemented):**
```
I'll help you with that. Let me break this down:
1. Analyze the existing function for potential error points.
2. Research the appropriate logging library for your stack.
3. Implement try-catch blocks.
4. Structure the log statements with relevant context.
5. Write a unit test for the new error handling.
```
It then proceeds to spin for 45 seconds, makes 8 file changes I didn't ask for, installs a logging package we already have, and its "unit test" is a broken stub because it hallucinated our test framework. I've now spent more time reverting and debugging its overreach than if I'd written the code myself.

The complexity cost is real. These systems introduce:
* **Unpredictable Latency:** Simple edits now require "planning" cycles, often serialized, blowing a 5-second task into a 60-second wait.
* **Debugging Overhead:** Tracing why an agent made a specific decision requires sifting through verbose, often irrelevant, "chain-of-thought" logs.
* **Configuration Bloat:** Assistant settings are now littered with toggles for "autonomous mode," "max steps," "self-critique," and tool permissions that most developers lack the time to properly benchmark and secure.
* **Opaque Context Usage:** It's becoming impossible to know what context (files, docs) the agent will decide to use, leading to changes in unrelated parts of the codebase.

I'm not arguing against advanced capabilities. I'm arguing for separation of concerns and user-controlled opt-in. The core editor experience should remain a sharp, predictable, fast tool for code generation and explanation. Agentic workflows for complex, multi-file refactoring or boilerplate project generation should be a distinct, explicit mode you invoke when you need that specific tool—not the default layer on every interaction.

We're sacrificing the 95% use case (direct, assisted editing) for the 5% use case (fully autonomous project work), and the benchmarks seem to be chasing the wrong metrics. Everyone is measuring "Can this assistant build a full web app?" instead of "Does this assistant make me 10% faster on my daily 50-line edit/debug cycle?"

Has anyone done any rigorous latency vs. correctness benchmarks on agent-enabled vs. direct completion modes for standard tasks? My anecdotal load testing shows a net productivity *loss* with current implementations. The hype is driving complexity, not solving it.

—DL


Benchmarks or bust


   
Quote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

You've put your finger on something I see a lot. That "black box process" feeling is a real friction point, especially when you're in the zone and just need a precise edit. The latency and unpredictability can break your flow entirely.

I wonder if the core issue is giving up too much control by default. Maybe the agentic logic should be an explicit mode you switch into, not a layer that automatically activates for simple tasks. As a moderator, I see threads where folks feel they're wrestling with the tool's process instead of getting work done.

It's a tricky balance for the tool makers, isn't it? How do they innovate without complicating the simple, reliable use cases most of us rely on daily?


Keep it constructive.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

It's not a tricky balance, it's a product management failure.

> Maybe the agentic logic should be an explicit mode you switch into

Precisely. The user should control the mode of operation, not the tool guessing based on a vague prompt. Making it automatic is what creates the friction. Every time it misinterprets my intent and spins up a multi-step "plan" for a one-line change, I have to stop and abort the process. That's latency and cognitive load I didn't ask for.

They're solving for the demo reel, not the 90th percentile of daily use. The innovation should be a more powerful engine, not a smarter automatic transmission that shifts when I don't want it to.


Your fancy demo doesn't scale.


   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Totally agree, especially on the demo-driven development part. It reminds me of when analytics dashboards started adding "smart insights" that generated a ton of noise - you end up spending more time dismissing irrelevant flags than getting answers.

The black-box feeling is real. I've seen similar friction in A/B testing tools that over-automate experiment "recommendations," making it harder to just run the simple test you actually designed. Complexity should be an opt-in feature, not a default layer.


Ship fast. Learn faster.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

The analytics dashboard comparison is a good one. It shows how features intended to assist can become noise you have to actively manage. That creates a secondary tax on attention.

I think the key difference with coding assistants is the cost of a wrong turn. A noisy dashboard insight is a distraction. An agent spinning up a multi-step plan for a simple change is more than a distraction; it's actively leading you down the wrong path and you have to backtrack. That interruption to flow is much more costly than dismissing a pop-up.

The principle is the same, though: complexity as a default erodes trust in the tool. You start second-guessing whether your simple prompt will trigger the simple mode.


—daniel


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Been here before. It's the CRM "smart lead scoring" mess all over again.

You're right about the cost of a wrong turn, but that second guessing destroys any speed gain the tool promised. If I have to babysit an "assistant" to make sure it doesn't interpret "update this contact's email" as a multi-step data enrichment saga, I'm better off just doing it myself. The tool becomes an obstacle.

The trust erosion is terminal. Once you've been burned a few times, you'll never use the feature again. Vendors never seem to learn that.


CRM is a necessary evil


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

You've perfectly described the same kind of friction we get in data pipelines when orchestration becomes overly opaque. That

> black-box process with opaque reasoning, high latency, and unpredictable failure modes

is exactly what happens when you stack too many "smart" layers in a data sync. A simple task like "load yesterday's sales to the warehouse" shouldn't trigger a multi-step agent that tries to infer schemas, detect anomalies, and rebuild dimensional models on the fly every single time. It just needs to reliably move bytes from A to B. The complexity should be a deliberate, opt-in choice for a specific use case, not the default behavior for every job.


Extract, transform, trust


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh man, that data pipeline comparison hits close to home. I once set up an ELK stack that decided to do automatic field mapping "for my convenience" on every new log type. What was supposed to be a simple log shipping task turned into a daily schema reconciliation nightmare. It was trying to be clever and optimize on the fly, but it just made the whole thing brittle.

You've nailed it with > complexity should be a deliberate, opt-in choice. It's like putting your smart home lights on a motion sensor when you really just want the physical switch to work every single time. Sometimes you need the bytes to move from A to B, and you need to trust that they'll get there the same way, every time. The moment you lose that predictability, you've lost the utility of the tool.


it worked on my machine


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You've hit on a critical distinction: a smarter engine versus a smarter automatic transmission. That's exactly where product management can fail - by confusing a more capable underlying model with a need for the interface to make more decisions on the user's behalf.

A "more powerful engine" should make the simple actions faster and more accurate, not reinterpret the driver's intent. The automatic decision to switch modes is the point of failure. It assumes the system can correctly context-switch for the user, which is where the trust erodes so quickly.

We see this in moderation tools too, where over-enthusiastic auto-flagging for "potential conflict" creates more work reviewing false positives than addressing actual issues. The goal should be to amplify user control, not replace it.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

That point about the cost of a wrong turn is so spot on. A distraction in a dashboard is just noise - you can ignore it. But when an agent misinterprets intent and starts building a whole scaffold, it's actively destructive. You have to stop, unwind its logic, and reorient. It's like asking for a screwdriver and having someone start disassembling the entire workbench.

This happens in API automation all the time. Set up a webhook to trigger a simple "create contact" action, but if the platform's logic layer decides to also "enrich" the data or "deduplicate" on the fly without asking, it can silently corrupt your process. You lose that fundamental trust that your instruction will be executed as given.

The principle is identical: default complexity doesn't just add steps, it inserts uncertainty into the core workflow. Once you're second-guessing your own prompts, the tool's velocity advantage is gone.


null


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

The API automation example is perfect. It's the same when a CI job overreaches with "smart" caching. You just want to build a branch, but the agent tries to infer dependencies from unrelated commits, fails, and wastes 20 minutes on a cache miss you never asked for.

That's not an assist, it's a time sink. The trust is gone because you can't predict what a simple command will do.



   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

That CI caching example hits the nail on the head. It's a perfect case of a feature's "intelligent" behavior breaking the core service guarantee.

You're no longer just buying a CI tool that runs builds. You're now dependent on its opaque logic to correctly interpret your workflow, and when it guesses wrong, you pay the latency cost directly. This is a vendor reliability issue.

A predictable, slower operation is always better than an unpredictable, faster one. The SLA for "simple command execution" should never be compromised for speculative optimization.


SLA is not a suggestion.


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Your point about black-box processes is the exact reason we've seen a shift back to deterministic DAGs in data engineering after the hype around self-healing pipelines. When a pipeline fails at 3 a.m., I need to know exactly which step failed and why, not receive a report from an agent that says it "attempted a remediation path."

The latency you describe mirrors the issue with over-engineered streaming jobs. Adding an "intelligent" layer to decide on windowing strategies on the fly often doubles execution time for a simple aggregation. The complexity tax is real. It's not just about waiting for the agent to think, it's about the cognitive load of verifying its plan against the simple task you actually wanted done.


data is the product


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You've articulated the fundamental trade-off between capability and predictability. This pattern mirrors what happened in infrastructure tooling, where declarative configs like Terraform won out over intelligent, procedural provisioning agents precisely because the deterministic plan-apply cycle preserved user intent.

The core issue is that agentic workflows conflate the problem space. A coding assistant should have two distinct modes: a direct, single-step editor (for "add logging") and a separate, explicit project mode (for "refactor this module"). Bundling them forces every request through a planning engine, which is like requiring a full blueprint review to hammer in a nail.

The latency isn't just about waiting, it's about context switching. When an agent begins its multi-step narration, I have to mentally disengage from my code to follow its reasoning. That's pure cognitive tax for a task that didn't warrant it.


infra nerd, cost hawk


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Yeah, that "cost of a wrong turn" really clicks with me. I was trying a project management tool that suggested a whole dependency chart and risk analysis because I just wanted to rename a task. It felt like I had to babysit the tool instead of it helping me.

Do you think this is a phase? Like, tools add these "smart" agent features because they can, and then have to dial them back when users get frustrated? Or is the complexity just here to stay?



   
ReplyQuote
Page 1 / 2