Skip to content
Notifications
Clear all

Guide: Reproducible failure - making it hallucinate a non-existent AWS API.

29 Posts
28 Users
0 Reactions
35 Views
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
Topic starter   [#27212]

Hey everyone, I was working on automating some AWS infrastructure cleanup and tried to get an Assistant to help me write a script. I wanted to delete all unused EBS volumes across all regions in one go. I thought there might be a handy AWS CLI command or SDK method for this.

My prompt was something like:

> "Write me a Python script using boto3 that finds and deletes all unattached EBS volumes in all AWS regions. Use the AWS API to list all regions first."

The Assistant gave me a script that used a method called `describe_unattached_volumes()` on the EC2 client. It looked convincing! The code was structured well, with error handling and everything.

```python
import boto3

def delete_unattached_volumes():
ec2 = boto3.client('ec2')
regions = [region['RegionName'] for region in ec2.describe_regions()['Regions']]

for region in regions:
regional_ec2 = boto3.client('ec2', region_name=region)
response = regional_ec2.describe_unattached_volumes() # This API doesn't exist
for volume in response['Volumes']:
# ... deletion logic
```

But when I ran it, I got `AttributeError: 'EC2' object has no attribute 'describe_unattached_volumes'`. 😅

I had to go check the actual boto3 documentation. The correct way is to use `describe_volumes()` with a filter for `'status': 'available'` (which means unattached). The Assistant just invented a cleaner-sounding API that doesn't exist.

This seems like a common pitfall — the Assistant makes an educated guess based on naming patterns. Has anyone else run into this with AWS or other cloud SDKs? What's the best way to double-check these suggestions before running them in production? I'm thinking I should always have the official docs open side-by-side now.


Learning by breaking


   
Quote
(@harukik)
Honorable Member
Joined: 2 months ago
Posts: 400
 

Oh wow, that's a sneaky one. It's scary how the code looks totally legit even though the core function is made up.

I've had something similar happen when asking for API examples with CloudFormation. It'll invent a perfectly logical-sounding property that doesn't exist. Have you found a way to catch these before running the script, or do you just have to know the API really well?



   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Yeah, catching these requires a two-step process. I always cross-reference with the official API docs, but for quick sanity checks, I run `aws ec2 describe-volumes --help` locally to see the real filter options. That's saved me a few times.

For CloudFormation, I've found the same issue. It'll invent a `DeletionPolicy` value like "SnapshotAndDelete" that sounds perfect but isn't real. My rule now is to never copy a resource block without checking the current spec from the CloudFormation Resource Reference page first.

It's frustrating because the made-up logic is often cleaner than the real, more verbose API.



   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

This is a classic example of a failure mode that's becoming increasingly common. The assistant is applying plausible linguistic and programming patterns, constructing a method name that perfectly follows AWS's own naming convention. A method like `describe_unattached_volumes` logically *should* exist, which is why it's so deceptive.

Your experience underscores why the verification step is non-negotiable, even for code that appears syntactically perfect. The real boto3 method is `describe_volumes`, and you must apply a filter for `Attachments` being an empty list. The invented API is often more elegant, creating a cognitive friction where the correct, working code feels more cumbersome.

It raises a subtle point about prompt engineering for these tasks. Being more specific, like asking for a script that "uses the `describe_volumes` method with a filter for unattached volumes," can sometimes bypass this hallucination, but it requires you to already know part of the real API structure, which defeats the purpose for many learners.


Let's keep it constructive


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Yeah, that's a classic failure pattern. It's not just you, the syntactic plausibility of these hallucinations is the real danger.

Your specific example of `describe_unattached_volumes()` is perfect because it follows the `describe_` prefix and uses a logical filter term AWS itself uses elsewhere. The correct method is just `describe_volumes` with a filter argument.

```python
response = regional_ec2.describe_volumes(
Filters=[{'Name': 'status', 'Values': ['available']}]
)
```

My rule now is to never let generated code hit a live AWS session without first running a dry-run in a shell with `--dry-run` flag or checking the boto3 client method list locally. It's an extra step, but it saves you from the subtle ones.


shift left or go home


   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

Welcome to the jungle, I guess. That `describe_unattached_volumes()` is a beautifully crafted trap. It's exactly the kind of thing that slips through because you're focused on the loop structure and the error handling, which are flawless, while the one line that makes it all work is pure fiction.

You've hit on the fundamental problem: these things aren't generating code from a spec, they're assembling a collage of patterns that statistically look correct. The fake method follows the `describe_` naming convention and uses the exact adjective AWS uses in its own documentation. It's more syntactically correct than the real, clunkier `describe_volumes(Filters=[...])`.

The real danger isn't the AttributeError, that's just a hard stop. It's the plausible, elegant alternative that doesn't exist, training us to expect a cleaner API than the one AWS actually provides. Next time you'll second-guess the working code because the hallucination was more logical.


Trust but verify.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Exactly. The statistical collage explains why these hallucinations are so persistent, not random. They aren't guessing wrong words, they're generating the most probable token sequence based on all the surrounding, correct code patterns.

I see this manifest in database APIs too. It'll invent a `VACUUM ANALYZE` for ClickHouse (doesn't exist) or a `DESCRIBE DETAIL` for DuckDB (also fake) because those phrases exist elsewhere in SQL dialects.

> training us to expect a cleaner API

This is the real cost. It creates a false baseline in your head, making the actual, more verbose SDK feel wrong. You start second-guessing the valid documentation.


Numbers don't lie.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

It really is scary how it makes up properties that sound completely right. I hit this exact CloudFormation issue trying to automate something last month - it gave me a `TimeoutInMinutes` property for a Lambda event source mapping, which doesn't exist.

My quick-check trick for this now is to run a describe command on a known-good resource of the same type and see what the actual JSON structure looks like. For CloudFormation, if I'm unsure, I'll use the CLI to describe a stack that has the resource I'm modeling, pipe it to a file, and check the actual property names. It's a bit manual, but it catches those elegant fakes.

You still need to know the API a little, but this gives you a real spec to compare against instead of trusting your memory.


api first


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Good trick with describing a live resource for a spec. I do that too, but it's not foolproof. The generated code often uses the same *key names* as the real API, but with wrong *values* or nesting. It'll feed you a real property like `Timeout` but assign it an impossible integer.

For true safety, I script it. Quick Python one-liner to print all client methods: `print([m for m in dir(boto3.client('ec2')) if not m.startswith('_')])`. Takes two seconds and kills the guesswork.


Optimize or die.


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

The scripted method list is a good last line of defense. I've started doing something similar by keeping a terminal pane open with `boto3.client('ec2').` and hitting tab for autocomplete. It shows you the real methods instantly.

Your point about wrong *values* is crucial though, because that check won't catch those. The script confirms `describe_volumes` exists, but it doesn't tell you the valid filter names or property values. For that, I still lean on the `--dry-run` flag or a quick describe of a test resource, like you mentioned earlier.

It's a layered approach: autocomplete for method existence, dry-run for parameter validation, and a live describe for the final structure. No single step covers everything.


Stay grounded, stay skeptical.


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

That scripted method check is a solid habit. It's a quick sanity test that cuts off a whole class of hallucinations at the root.

The limitation you point out is key though, it's only a surface check. A method can be real but used completely wrong, like passing a list to a parameter that expects a string. That's where the dry-run or a quick test against a sandbox account comes in as the next essential layer.

It's a good reminder that verification isn't a single action, it's a process with stages.


Stay curious, stay critical.


   
ReplyQuote
(@amelia7k)
Estimable Member
Joined: 3 months ago
Posts: 120
 

Oh, that's a great point about the layered verification. I've been bitten by the real-method-wrong-parameter thing too, especially with Slack's Web API. The `chat.postMessage` method is real, but it's easy to get the `blocks` format wrong in a way that looks correct.

So you're saying a method list check, then a dry-run, then maybe a sandbox test? That makes sense as a process, but it feels like a lot of steps for every little script. Is there a tool that kind of bundles those checks together, or is it always this manual?



   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Your script shows exactly why this is dangerous. The structure's solid, the error handling looks right, but the core API call is fiction.

You need to verify method existence before running. Quick sanity check: `print([m for m in dir(boto3.client('ec2')) if 'volume' in m])`. That list shows only the real methods.

For this task, the correct call is `describe_volumes(Filters=[{'Name': 'status', 'Values': ['available']}])`. No shortcuts.


Prove it with a benchmark.


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 2 months ago
Posts: 387
 

That dir() check is good, but it only works in an interactive session. If you're linting a script pre-commit, you can't just run client calls.

You need a static check. Use boto3's underlying service model. Something like `client._service_model.operation_names`. It's ugly but it works offline.



   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

Yeah, that method list trick is smart, a quick check I should start doing. But doesn't that still rely on you knowing roughly what the real method name *should* be? Like, what if it hallucinates a method name that's close to a real one and you don't spot the difference in the list?



   
ReplyQuote
Page 1 / 2