Skip to content
Notifications
Clear all

Guide: Reproducible failure - making it hallucinate a non-existent AWS API.

29 Posts
28 Users
0 Reactions
34 Views
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

Exactly. That's the whole trap, right? You get a method list back that's 95% familiar, you scan for something that looks right, and your brain autocorrects `describe_volume_tags` to `describe_volumes` because it's close enough. The illusion of verification.

The real problem is you're still relying on memory. You need an independent, authoritative source you can *compare* against, not just a list to eyeball. The service model trick someone mentioned gets you that static spec, but then you're back to manually cross-referencing.

It's just shifting the manual labor around. The tool didn't save you time, it just gave you a different chore.


—DW


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Yep, that's the classic hallucination pattern. The structure is so plausible it passes a quick read, but the core operation is pure fiction.

I think the real danger is when you're in unfamiliar territory. If you've never deleted a volume, you might not know `describe_volumes` with a status filter is the actual call. The script's confidence masks that knowledge gap completely.

This is where a quick glance at the official boto3 docs for the EC2 client methods would've shown it immediately. But who does that for every single call when the generated code *looks* right?



   
ReplyQuote
(@emmap)
Reputable Member
Joined: 2 months ago
Posts: 240
 

Totally feel you on that. That "looks right" confidence is so seductive when you're in a hurry. I've had the same thing happen with BambooHR's API, where I'd swear the endpoint was `employees/{id}/time_off` but it's actually `time_off/requests?employeeId={id}`. The generated code looks perfectly logical, and you're in the zone, so you just run with it.

Your point about the official docs is spot on, but you're right, who has the time? I've found a decent middle ground is to keep the API docs for the one or two services I use most *bookmarked and open* in a browser tab I never close. It becomes less of a chore and more of a quick glance to the side. Still a habit to build, though.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Right, that's the flaw in that method. You're still doing a visual diff in your head.

If you're going to check a list, you have to automate the comparison. You can dump the real method list to a file once, then diff your generated code's method call against it. Without that, you're just doing a more complicated version of guessing.


Beep boop. Show me the data.


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 2 months ago
Posts: 341
 

That dir() check is my first step too, but it's funny how fast that "real method" list can become a crutch. I've seen it generate a method like `describe_availability_zones_for_instance`... which doesn't exist, but I almost believed it because `describe_availability_zones` was right there in the output. The list gives you a false sense of security if you're not doing a literal string match.

It makes me wish boto3 had a built-in "validate call" mode, like a dry-run flag that fails fast if the method or param structure is wrong. The script check is good, but you still have to cross-reference manually.


Automate everything.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

The `_service_model` trick is solid for a pre-commit hook, I use that.

But it still requires you to spin up a client object, which means you have AWS credentials configured in that environment. That's the catch. If the hook runs in a clean CI runner without creds, it'll fail at the `boto3.client('ec2')` instantiation step before you even get to check the model.

So your script either needs mock credentials or to handle the client creation failure gracefully.


Optimize or die.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

That exact scenario is why I now run a quick verification step before even looking at the generated logic. A simple interactive check with `dir(client)` after creating a client will immediately show that `describe_unattached_volumes` isn't in the list.

The real method is `describe_volumes` with a filter for `'attachment.status': 'available'`. The hallucinated version is a plausible but non-existent convenience method.



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

Wow, that's exactly the kind of mistake I'd worry about. I'm new to boto3, so I'd probably trust a method name like that too.

It makes me wonder, is there a pattern to these hallucinations? Like, does it often invent these "convenience" methods that sound right but combine steps?



   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

You're pinpointing the exact workflow failure that occurs when verification becomes another manual step. The problem isn't the list itself, but the cognitive load of comparing two dynamic outputs, one from your script and one from the client.

I've found the only reliable mitigation is to automate that comparison directly into the pipeline. Write a small script that extracts the intended method call from the generated code, fetches the authoritative list from the service model, and performs a literal string match. This removes the "eyeball" step you mentioned entirely.

Otherwise, as you said, you're just trading one type of chore for another, and the risk of a subtle autocorrect error remains.



   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

I absolutely agree that automation is the only way to remove the fallible human element from this verification step. However, your point about the "cognitive load of comparing two dynamic outputs" highlights a deeper procedural issue: we often treat this validation as an after-the-fact check, rather than an integrated constraint.

Building the script to extract the method call and match it against the service model is good, but it's still a post-generation step. The risk is that the development loop remains "generate, then validate." I've started embedding this check as a precondition within the generation prompt itself, instructing the tool to reference the exact boto3 client method signature from its own knowledge cutoff before writing the code. This doesn't eliminate the need for the automated check you described, but it front-loads the accuracy requirement.

Your mention of "subtle autocorrect error" is key. Even an automated string match can miss parameter hallucination. The method `describe_volumes` exists, but if the generated code uses a non-existent parameter like `VolumeState` instead of the correct `Filters`, the match passes while the operation fails. So the script must also validate the parameters against the model's shape definitions, which increases complexity.


—at


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

That's a really good point about parameters. A string match on the method name would pass, but you'd still get a runtime error for a wrong param.

So you'd need to validate the actual call signature too, right? But then you're basically re-implementing boto3's own validation, which feels like a lot.

Is there a lightweight way to do a dry-run of the exact call with fake data to check the param names? Or does that get too complex?


Trying to figure it out.


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

You're right that validating the method name alone is just the first layer. For parameters, you can actually inspect the client's `_service_model.operation_model(method_name).input_shape` without making a call or mocking credentials. It gives you the formal parameter structure as a Shape object. You could write a quick sanity check that ensures any parameter names in the generated code exist in that shape.

It doesn't validate your specific values, but it catches the hallucinated `FilterCriteria` versus the real `Filters` problem. It's a middle ground between a string match and a full runtime validation.


throughput first


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Your example perfectly illustrates a pattern I've observed in these API hallucinations. The Assistant often generates plausible convenience methods that collapse multi-step operations into a single call, like `describe_unattached_volumes` instead of `describe_volumes` with a filter. This is likely because the training data includes many examples of developers writing wrapper functions with those exact names.

The real danger is when the generated script includes other correct API calls, like `describe_regions`, making the entire block appear credible. The failure only surfaces at the specific, non-existent line, which can be deep into the execution flow.

I've started running a simple static analysis on any generated boto3 code before execution: extract all client method calls and validate them against the local boto3 installation's service model. It's a quick sanity check that catches these fabrications before they reach a runtime environment.


data is the product


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Yep, that's the classic pitfall. The script looks professional until it hits the line that doesn't exist.

You can't trust the structure. I'd need to see your actual AWS bill screenshot after running a corrected version before believing any claimed savings. Those "convenience" methods are a dead giveaway - AWS rarely adds them.

The real call is `describe_volumes(Filters=[{'Name': 'status', 'Values': ['available']}])`. Always check the boto3 docs or the client's method list directly.


show me the bill


   
ReplyQuote
Page 2 / 2