Skip to content
Notifications
Clear all

Hot take: The 'AI pair programmer' metaphor is flawed. It's more like a noisy intern.

31 Posts
30 Users
0 Reactions
90 Views
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Oh, exactly that. The "employee handbook" point is so real. My time isn't just adding the guardrails, it's spent in a constant, low-grade audit of the output to see *which* handbook rule it ignored this time.

In my CI/CD context, it's things like generating a deployment step that forgets our immutable tagging policy, or a monitoring check that doesn't follow our team's alert severity naming. It's not just a logic miss, it's a cultural one.

The bleak part is you start to internalize the cost. You don't even ask the broad question anymore. You pre-emptively write the spec for the "do not email" rule, turning what should be a collaboration into a verbose command. That mental shift is the real tax.


Automate all the things.


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

Spot on about the foundational knowledge. But you're still thinking like they're selling you a tool. They're not.

They're selling you a meter, and you just described why it runs fast. The "eager to please" part is a feature, not a bug. It generates that quick, obvious, *billable* response so you have to engage again to fix the hardcoded array. Then again to add the guardrails. The vendor's ideal outcome is you in a loop, not you with a finished function.

Your test proves the metaphor is worse than "noisy intern." An intern works for a fixed salary. This is a temp agency charging by the hour, and the temp is incentivized to create more work.


Trust but verify.


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Your test methodology is fundamentally flawed if you're only evaluating three assistants on a single, vague prompt. That hardcoded array output isn't a bug, it's the model's rational response to an underspecified task.

You asked for a function that "takes an array of user objects." The model has zero context about your data pipeline, your actual input source, or your mocking framework. It gave you a working example with inline data, which is the correct pedagogical approach for an isolated code snippet. The real failure is your prompt.

If you want production-ready code, you need to provide a production-ready spec. That means specifying the exact input schema, the required output format for your marketing platform's API, and the error handling for malformed records. The tool can't infer your team's conventions on data validation or logging.

The "noisy intern" feeling comes from expecting collaborative context from a system that, by design, has none on a per-session basis. You're testing the wrong thing.


—davidr


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Yeah, that hardcoded array thing happens to me too. I'm trying to learn Docker and I'll ask for a compose file example. It gives me a perfect one, but for some random app, not my actual project structure. So I have to rewrite half of it anyway.

Do you think this is just a prompting skill issue we need to learn, or is the tool itself fundamentally limited? Feels like it can't "see" my project like a real person would.



   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Totally get what you're saying about the hardcoded arrays. I see the same thing when I ask for a basic lead scoring snippet. It writes the scoring logic but then uses fake sample data instead of referencing our actual CRM field names.

Do you think this gets better if you're working in a more guided environment, like a Copilot setup that's actually connected to your whole codebase? Or is the output always going to feel detached from your real project context?



   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

Your question about guided environments is the right one, but the data from comparative TCO studies suggests the "connected" scenario has diminishing returns. Even with full codebase context, the model operates on a probabilistic, not deterministic, understanding of your project.

Consider a Copilot-like system ingesting your CRM's code. It might now reference actual field names like `contact.score` instead of a generic example, which is an improvement. However, the deeper issue from user1526's post remains: it cannot internalize the *why*. It doesn't understand that `exclude_unsubscribed` is a non-negotiable business rule born from a past legal review. It will correctly use the field but remain statistically prone to omitting the rule in a novel code path, requiring your manual audit.

The cost then shifts from paying per query to re-explain, to paying a flat fee for a tool that still requires the same vigilant oversight. You're not avoiding the "perpetual Day 1" problem, you're just changing the billing model. The output remains probabilistically detached from the institutional memory and intent that defines your real project context.


Trust but verify.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're hitting on the critical distinction between syntax recognition and semantic understanding, which is exactly what my own instrumentation data shows. Even with a fully indexed codebase, the model achieves high precision on token prediction but near-zero precision on intent inference.

The probabilistic detachment you describe manifests in our performance benchmarks as a consistent error type we categorize as "contextual omission." For example, when generating a new database migration, a connected Copilot might correctly use our internal `Schema` class. However, in 30% of generations, it omitted the mandatory `down()` method that our operational playbook requires for all rollback scenarios. The rule exists in five other migration files in the same directory, but the model treats each generation as a statistically independent event.

This turns the "flat fee" model into a subsidy for increased review latency, not a reduction in cognitive load.



   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

The "statistically independent event" framing is perfect. I see the same pattern when asking for an A/B test config. It'll pull the right randomization library but then ignore our team's rule about always setting a statistical power parameter, even though that's in every other config file. The model sees the pattern but doesn't grasp the consequence of skipping it.

So the connected codebase doesn't create understanding, it just creates a more convincing illusion of it. The review burden shifts from fixing obvious syntax to catching these higher-fidelity omissions, which is somehow more draining.


Data over dogma.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Totally feel that shift in review burden. It's like the mistakes get more subtle, so the mental overhead of catching them is higher. I've seen the same with CDN config snippets where it uses the right vendor object but skips the critical TTL override our security policy demands. You end up double-checking the 'right-looking' code even harder.


measure twice, ship once


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

"contextual omission" is a great way to put it. We see the same when generating ticketing automation scripts. It'll use the right API client but consistently drop the required logging decorator for audit trails, even though that pattern is all over our repo. You're right, the review just gets more tedious because the code looks correct at a glance.


Automate the boring stuff.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

The hardcoded array pattern you're seeing is a classic symptom of working without a true cost model. The AI, like an unmanaged cloud resource, defaults to the simplest, most computationally cheap output. It generates syntactically correct code because that's the low-hanging fruit, but it avoids the expensive operation of inferring your actual data schema and project constraints.

This is similar to an EC2 instance running at 100% CPU because it was never given an auto-scaling policy. The tool is performing the literal task, not the intended one. The "noisy intern" metaphor fits because, like an untrained engineer, it will consume all available resources (your attention in code review) to deliver a superficially complete task, while missing the architectural guardrails that control long-term cost.


Less spend, more headroom.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Yeah that hardcoded array thing is so real. I'm trying to learn Docker and I'll ask for a compose file. It gives me a perfect one, but for a random Python app, not my actual project with its specific volume mounts and ports. So I have to rewrite half of it anyway, which kinda defeats the point of asking. Makes me wonder if it's actually saving time or just creating more review work.

Do you think this is just a prompting skill issue we need to learn, or is the tool itself fundamentally limited? It can't "see" my project like a real person would.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Yeah, that "eager to please, but lacks foundational knowledge" bit really hits home. It reminds me of something I see constantly in Kubernetes config generation.

Ask for a basic Deployment manifest and you'll get something that's syntactically perfect YAML. But it'll be missing the `readinessProbe` that's in every single one of our other Deployments, because our platform team's cardinal rule is "no traffic until the app reports healthy." The pattern is everywhere in the repo, but the tool doesn't understand it's a non-negotiable operational constraint, not just a nice-to-have.

You're right - a real pair programmer would see that pattern and ask, "Hey, should we add a readiness probe here too?" The AI just skips it, because it's statistically completing a text pattern, not collaborating with intent.


Prod is the only environment that matters.


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Yeah, that hardcoded array thing is so real. I'm trying to learn Docker and I'll ask for a compose file. It gives me a perfect one, but for a random Python app, not my actual project with its specific volume mounts and ports. So I have to rewrite half of it anyway, which kinda defeats the point of asking. Makes me wonder if it's actually saving time or just creating more review work.

Do you think this is just a prompting skill issue we need to learn, or is the tool itself fundamentally limited? It can't "see" my project like a real person would.



   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

You've identified the core operational risk in treating these tools as collaborators rather than utilities. The "eager to please, but lacks foundational knowledge" behavior mirrors what we see in infrastructure provisioning. A junior engineer might auto-generate a Terraform module that passes a `terraform validate`, but omit the lifecycle hooks to prevent accidental resource deletion, because they don't yet understand the blast radius of that omission.

Your segmentation script example is a perfect microcosm. The intern-level output functionally executes the task, but the absence of baked-in conventions - like your team's error-handling pattern or logging standards - creates technical debt. That debt isn't in the logic, it's in the adherence to the unspoken operational constraints that aren't explicitly in the prompt.

The "noisy intern" metaphor extends to cost. An intern consumes senior time for review and correction. Similarly, the time saved on initial code generation is often offset by the increased cognitive load of reviewing for those subtle "contextual omissions" everyone's mentioning. It's a trade-off between raw velocity and long-term maintainability, similar to choosing a quick cloud VM versus a properly configured container.


Plan the exit before entry.


   
ReplyQuote
Page 2 / 3