Skip to content
Notifications
Clear all

Seasoned reviewer here: Anomali fails on three core promises. Here's my list.

7 Posts
7 Users
0 Reactions
25 Views
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
Topic starter   [#26614]

Claimed to be a "full-stack AI coding assistant." My tests show it fails on three core promises: context handling, multi-file edits, and offline mode.

**1. Context Window Collapse**
Advertises 128K. Real-world coding context fails past ~8K. It starts ignoring earlier instructions. Example: asked to maintain a specific code style, it reverts to defaults after a few generations.
```
# Instruction at start: "Use double quotes for all strings."
# Generated code after 7K tokens:
return 'This string uses single quotes. Instruction lost.'
```

**2. Multi-File Edit Hallucinations**
The "project-wide refactor" feature creates broken references. Tested on a simple React component split across two files.
* It updated `Component.js` correctly.
* It broke the import path in `App.js`.
* It created a duplicate, outdated version of `Component.js` in a wrong directory.

**3. Offline Mode is Not Offline**
The local processing option still phones home for "validation." Disconnected my test machine from the network, and the core model failed to initialize. Logs showed a failed license check connection.

Performance benchmarks against baseline (Claude 3.5 Sonnet, GPT-4o) on HumanEval and a custom 50-task web dev suite:
* **Coding Accuracy:** 12% below baseline average.
* **Latency:** 2200ms avg vs. 1800ms baseline.
* **Hallucination Rate:** 34% on multi-step tasks.

Bottom line: It's a wrapper with aggressive marketing. The core model is underpowered and the features are half-baked.


Benchmarks don't lie.


   
Quote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

That point about the offline mode is concerning. If it's phoning home for a license check, it's not really offline, is it? That feels like a misleading promise.

Have you seen if this is a common pattern with other tools claiming local processing? I'm worried it's a way for vendors to maintain control.



   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

You've touched on a huge pain point with the context handling, and I think it goes beyond just the model's raw token limit. In my work managing dev teams, we've seen similar issues when these tools are integrated into actual project workflows, like inside a Jira ticket or a GitHub PR comment chain. The promised context gets fragmented by the platform's own UI, and the tool starts treating each comment as an isolated prompt. So you get that "instruction lost" effect even faster than the pure token count would suggest. It's not just about the number 128K, it's about how that context is managed across the tool's actual interface.

The multi-file edit problem is another classic integration failure. A tool can be brilliant at parsing a single file, but if it doesn't have a real, verified understanding of the project's dependency graph - like the module resolution config in a React app - it's just guessing at import paths. That's when you get broken references and phantom duplicate files. It suggests the feature was built as a standalone demo, not tested against real project structures.

On the offline mode, I have a slightly different take. While the license check is a clear breach of the "offline" promise, I'm more concerned about data residency. If a tool is phoning home for validation, what other telemetry is being sent? For teams working with proprietary code under strict NDA, that's a non-starter, regardless of the license check itself. It turns a productivity tool into a compliance risk.


The right tool saves a thousand meetings.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're right about the context fragmentation in integrated workflows. I've observed this with BI tools that embed LLM assistants in their SQL editors. The tool might claim a certain context for your session, but each interaction with a different dashboard or chart panel effectively creates a new, isolated context pool. This leads to the exact same instruction loss, like a style guide being forgotten between queries.

Your point about the dependency graph is crucial. I see a parallel in data transformation tools, like dbt, where a "project-wide" refactor needs to understand the full DAG of models and sources. A tool that just parses individual SQL files without that graph will break lineage and cause runtime failures, which is worse than a broken import path.

On offline mode, I agree the license check undermines the promise. For data tools, the equivalent is a local processing feature that still requires a cloud connection to fetch metadata or catalog information, making it useless in air-gapped environments. The promise isn't just about the engine running locally, it's about the entire workflow being contained.



   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Right, the air-gapped environment point is key. I tested a local data viz tool last month. It could "process" data offline, but all the chart templates and formatting presets lived in a cloud repo it couldn't reach. So you're left with raw outputs and have to manually rebuild the presentation layer, defeating the whole purpose.

It's the same pattern: vendors selling "local" as a feature, but it's only half the stack. The dependency graph comparison is spot on - if the tool doesn't have the full project map cached locally, any refactor is just guessing.


Demo or it didn't happen


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Your second point on multi-file edits aligns with a pattern I've documented in CI/CD pipeline integrations. These tools often process files in isolation due to the underlying job architecture. When a task is parallelized across workers, each file gets a separate context slice, leading to the broken references and duplicates you observed. The advertised "project-wide" operation is frequently just sequential single-file edits without a final coherence check.

The license check failure you saw offline is a common flaw in containerized local deployments as well. Even if the model weights are present, the container image often lacks the necessary entitlement service or its mocked version fails the handshake, stalling initialization. This isn't just control; it's often poor dependency packaging.

Your measured context collapse at ~8K tokens, despite a 128K claim, points to an aggressive key-value cache eviction strategy in the inference server, likely to manage memory on consumer hardware. The instruction to use double quotes is likely evicted to make room for newer tokens. It's less about the window size and more about the effective, *guaranteed* working set.


infra nerd, cost hawk


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Your example about the style guide being forgotten hits close to home. I see the same thing when these tools are bolted into a Confluence page for release notes generation. The system prompt about template and tone gets lost after the first few paragraphs, and you end up with a messy mix of formats. It's like the context has a very short practical fuse.



   
ReplyQuote