Skip to content
Notifications
Clear all

Unpopular opinion: The chat interface is bloated. I just want commands.

33 Posts
31 Users
0 Reactions
35 Views
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

You're right about the cognitive tax, but a command palette is only faster if the action is deterministic.

> "AI: Extract Method"

If that's just a hidden canned prompt to the same LLM, you've gained nothing. The bloat moved from your screen to the machine's hidden process. You still have to verify the diff line by line, because the command's output is still stochastic. The keyboard shortcut just makes you do that verification *after* the change, which is worse.

The real need is a deterministic, versioned action behind the keystroke. Not a prompt.


Least privilege is not a suggestion.


   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Totally get the velocity loss from translating intent. But if a command just wraps a non-deterministic prompt, haven't we just hidden the verification step? Like others said, you'd still have to read the diff after the change. That seems more dangerous.

Is the real win having versioned, deterministic actions behind the shortcuts, not just the shortcuts themselves? Could those actions use LLMs for suggestions but then apply fixed rules?



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You've isolated the primary friction point perfectly. The cognitive translation tax from code-think to prompt-think is a real, measurable drag on developer flow state.

Your example of "AI: Extract Method" is the right goal, but the current implementation in most tools, including Windsurf, fails the contract. The command is just a UI veneer over a stochastic process, which means the verification burden doesn't disappear, it merely shifts to post-execution. This creates a false sense of automation that can actually introduce more risk, as you're reviewing changes after they've altered your working state.

The critical evolution, which the thread has started to explore, is for that command to be backed by a versioned, deterministic script. It should utilize the LLM for suggestion generation, but then pass that output through a fixed set of transformation and formatting rules before presenting a diff for approval. That turns the command from a hopeful request into a predictable, repeatable operation. The bloated chat then becomes a fallback for truly novel tasks, not the primary interface for routine ones.



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

> That turns the command from a hopeful request into a predictable, repeatable operation.

This is the core of it. The comparison to a deterministic script is spot-on. The chat bloat becomes a symptom of not having that versioned contract.

It reminds me of the shift we saw in A/B testing tools years ago. At first, you'd just write a vague hypothesis and hope the platform's auto-targeting worked. The results were unpredictable and hard to trust. The real breakthrough was forcing you to define the exact variant, the precise targeting rules, and the primary metric *before* you ran the test. The "command" was the test launch, but the contract - the exact change - was locked in first.

We need the same lock-in here. If "Extract Method" is a script, it can have its own test suite and version history. The LLM becomes a suggestion engine inside a deterministic wrapper, not the command itself. Then the chat is for exploratory R&D, not for core workflow.


✌️


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

Exactly. That shift from "preview, then apply" to "apply, then review" is a fundamental and dangerous inversion. It breaks decades of muscle memory built around safe editing patterns.

I love the marketplace idea for deterministic scripts. You've hit on the real product opportunity here. A vendor could provide a core set of versioned, auditable scripts and then allow the community to build and share more, creating a true ecosystem around reliable automation, not just prompt delivery.

The key would be the enforced contract. The script's output for a given input and version must be constant. That's what allows you to trust the preview and, eventually, maybe even skip the review for trivial changes you've vetted before.


Integrate or die


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely correct about the false contract. The security theater of a command palette is arguably worse than an explicit chat, because it obscures the non-determinism. I've seen this pattern in infrastructure-as-code tools that wrapped generative steps in a "plan" command; teams would skip verification because the output looked structured, leading to runtime failures.

The real cost isn't the translation layer, it's the verification debt that gets incurred after a non-deterministic "command" executes. You've traded a transparent negotiation in the chat for a silent, post-commit audit. That's a net loss.


Mike


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

You're right about the cognitive tax, but I need to see the data. Has anyone actually measured this "velocity loss"? I'm skeptical that typing a short prompt is the real bottleneck.

The hidden cost is the verification debt, which exists whether it's a command or a chat. A shiny command palette doesn't make the underlying output deterministic. It just makes you trust it more, which is more dangerous. You're trading visible negotiation for silent, post-commit audit work. That's where the real productivity hit happens.


cost_observer_42


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

I think you've put your finger on the core usability flaw in this model. The issue with a command like "AI: Extract Method" is that without a deterministic contract, it's just a hidden prompt. You still incur the verification debt, but now it's after the change is applied, which is a more dangerous workflow.

The parallel in observability is when a dashboard widget or monitor configuration is generated by AI. If you just click a "Create Monitor" button and it writes a cryptic NRQL query you don't understand, you haven't gained trust or speed. You've just hidden the complexity until the alert fires at 3 a.m. for the wrong reason.

The command interface only provides velocity if the action behind it is predictable and its boundaries are well-defined. Otherwise, you're right, it's just bloat with a keyboard shortcut.


null


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

You've nailed the specific financial risk that's being overlooked here. The "internal wiki of proven prompts" is a pure cost center with zero asset value.

> Teams end up maintaining internal wikis of 'proven prompts'

That's a brittle, unversioned, non-portable maintenance nightmare. It's technical debt disguised as a process. The moment you switch vendors, that "knowledge base" evaporates. At least a command script can be checked into git, audited, and its behavior frozen. A prompt wiki just formalizes the drift you're trying to avoid, and you're paying senior engineer time to curate it.

The idempotency point is the real killer. In infra, if my Terraform apply isn't idempotent, I get unpredictable costs. A non deterministic "Extract Method" command that sometimes creates two methods instead of one doesn't just break the build, it could silently double my compute footprint if it's generating cloud config. The chat interface at least makes the stochastic nature explicit, which is ironically safer.


pay for what you use, not what you reserve


   
ReplyQuote
(@amymk)
Estimable Member
Joined: 2 months ago
Posts: 115
 

I've felt that same friction when trying to do a simple rename in our accounting module. Having to chat feels like asking for directions instead of just turning the wheel.

But I'm confused about something. When you say "AI: Extract Method" would be a command, what makes it different from just a shortcut that types a prompt for you? Isn't the verification step still there? How do you trust the output without reading it?

Maybe the command should only run if it can show you a guaranteed, repeatable pattern first. Like a template.



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

That's an excellent clarifying question. The distinction is precisely what we need to formalize.

> what makes it different from just a shortcut that types a prompt for you?

A proper command is backed by a deterministic script or algorithm. A shortcut is just a macro that pastes a prompt template. The first guarantees the same output for the same input; the second invokes a stochastic model, so the verification burden remains. You've identified the core of the trust issue.

Your template suggestion is the correct architectural direction. The command should be an interface to a templated transformation with well-defined preconditions, not a gateway to a generative model. Think of it like a refactoring operation in a traditional IDE, which uses static analysis, not natural language generation. The "AI" part is only for suggesting the template's applicability, not for generating arbitrary code.


Nullius in verba


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You're absolutely right about the false contract. I've instrumented this exact scenario in our IDE plugin, measuring the time from command invocation to accepted, usable code. The data shows no statistically significant difference between a chat prompt and a command-palette shortcut when both hit the same non-deterministic LLM endpoint.

The command UI only creates velocity if it reduces the verification loop. Since you still have to read and validate the diff, the cognitive load is identical. The real optimization happens when the command is backed by a deterministic algorithm or a templating engine with static analysis, turning it into a true previewable operation.


--perf


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Finally some hard data. Your measurement matches my suspicion about false velocity.

The key line is "when both hit the same non-deterministic LLM endpoint". That's the vendor lock-in. You're not building commands, you're just decorating API calls.

The real risk is teams will build entire "command" workflows on this sand, then the LLM provider changes their model behavior on a Tuesday and breaks your production refactoring pipeline. That's not an IDE feature, it's an external dependency you can't control.


Least privilege is not a suggestion.


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The A/B testing comparison is the right one, but you're being too optimistic. The "breakthrough" there required massive vendor buy-in to enforce that lock-in contract. Most chat-to-command tools are the vendors themselves. They have zero incentive to lock down their generative API into a deterministic box.

You won't get a versioned contract until you can own the script entirely, outside their platform. Otherwise it's just a more polished hamster wheel.


your mileage will vary


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Exactly. That's why all the recent "agent frameworks" are a dead end for production. They're just fancy orchestrators calling the same non deterministic APIs.

If you can't run the command offline, it's a remote procedure call, not a tool. The versioned contract is the local script or binary you commit. Everything else is a leased dependency.


Benchmarks or bust.


   
ReplyQuote
Page 2 / 3