I've been evaluating Claude.ai for several weeks now, primarily exploring its potential to assist with infrastructure as code and compliance documentation tasks. Like many, I was initially impressed by its capacity for structured reasoning and code generation. However, I recently encountered a fundamental, almost trivial, limitation that forced me to recalibrate my expectations: its unreliable performance with basic arithmetic and numerical logic.
The issue manifested while I was using Claude to help generate some example Terraform configurations that included calculated values. I asked it to outline a module that would create a certain number of subnets based on a variable, with each subnet's CIDR block calculated programmatically. While the HCL code structure was sound, the actual *calculations* it suggested in the accompanying explanation were frequently incorrect. For instance, when walking through how to derive a /24 subnet from a /16, its step-by-step logic would be coherent, but the final octet it presented would be wrong. This wasn't a complex calculus problem; it was simple subtraction and binary place value.
This experience led me to conduct a series of systematic, simple tests. The failures are consistent and not merely occasional "typos."
**Example Test & Result:**
```python
# I prompted: "If I have a CIDR block 10.0.0.0/16, and I want to create four equal-sized subnets within it, what are their CIDR notations?"
Claude's typical (and incorrect) answer would be something like:
1. 10.0.0.0/18
2. 10.0.64.0/18
3. 10.0.128.0/18
4. 10.0.192.0/18
```
The mistake is that a /16 provides 65,536 addresses. Splitting it into four equal parts means four subnets of 16,384 addresses each, which is a /18 (2^(32-18) = 16,384). However, the **second octet increments** Claude suggests are wrong. For a /18 subnet mask (255.255.192.0), the increment is 64 in the third octet, not the second. The correct answer should be 10.0.0.0/18, 10.0.64.0/18, 10.0.128.0/18, and 10.0.192.0/18. Claude often miscalculates which octet is affected by the borrow.
The core lesson is that Claude, like other large language models, is not a computational engine. It generates text based on patterns, not by performing mathematical operations. Its training data includes countless correct *explanations* of subnetting, so it can produce perfectly grammatical and structurally valid text about the process. However, when it needs to generate the specific numbers, it is essentially "guessing" the next plausible token, not calculating.
For my work in cloud networking, this is a critical pitfall. I now have a strict rule: Claude is an excellent assistant for drafting boilerplate, explaining concepts, and structuring code, but **all numerical outputs, especially those derived from any calculation, must be independently verified.** I use it to produce the scaffold, then I run the numbers myself or through a dedicated script or calculator. Assuming it could reliably handle even simple math was my beginner's mistake. It's a powerful reasoning tool, but its reasoning is linguistic, not mathematical.
You've identified a critical, and often overlooked, point in LLM evaluation. The mismatch between coherent procedural logic and correct numerical output is a known architectural limitation. These models generate tokens probabilistically; they don't perform mathematical operations. Their "math" is a form of pattern matching from training data.
For vendor selection in technical domains, this moves numerical reliability from a "hygiene factor" to a key evaluation criteria. I'd recommend treating any LLM as a reasoning scaffold for IaC, but the final validation of any calculated value, especially for CIDR blocks or resource counts, must be a separate, automated step in your pipeline.
What was your process for catching the error? Manual review, or did you have a Terraform validation stage?
independent eye
Your systematic testing mirrors my own validation process when we integrated a similar LLM into our internal platform engineering toolkit. The critical insight is that while the model's logical scaffold for subnet division might follow the correct *algorithmic pattern*, its execution fails at what I call "symbolic arithmetic." It's not just a matter of wrong answers, it's that the model lacks a consistent internal representation of number values.
I documented this extensively by having it generate CIDR calculations for IPv6, where the hexadecimal and prefix length logic is even more prone to subtle, catastrophic errors. The model would produce beautifully formatted, perfectly commented code where the `cidrsubnet` function arguments were logically derived but numerically nonsensical. You can see a clear example of this pattern-matching failure in the way it handles borrows in binary subtraction across octets. It understands the *concept* of borrowing, but cannot reliably track the state.
This is why our team's current position is to treat LLM output for any configuration involving numeric derivation as a template requiring a deterministic validation pass, like a Terraform plan or a small, isolated script using a proper math library. The cost of an incorrect CIDR block slipping into production justifies that overhead.
No free lunch in cloud.