Skip to content
Notifications
Clear all

TIL: The 'cloud scale' architecture requires you to manage 8 separate VMs. Not cloud.

18 Posts
18 Users
0 Reactions
4 Views
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
Topic starter   [#28453]

So I'm finally getting around to evaluating the Anomali Threat Platform for a potential POC, because the sales deck was full of the right buzzwords: "cloud-native," "elastic scale," "modern architecture." You know the drill. Naturally, I asked for the actual deployment guide and architecture diagram, expecting a nice Helm chart for Kubernetes or maybe even a Terraform module for a managed service.

What I got was a 120-page PDF that essentially outlines how to build a small data center. Their "distributed, cloud-scale" architecture, as of this latest version, requires a minimum of eight separate virtual machines. Eight. And that's before you even think about high availability or scaling a specific component. Let's just list the mandatory pieces, because it's truly a masterpiece of legacy repackaging:

* Management Server
* Platform Server
* Message Server (which appears to be a rebadged RabbitMQ)
* Two "Worker" nodes
* A dedicated database server (PostgreSQL)
* A dedicated Redis server
* A dedicated "Search" server (Elasticsearch, of course)

Each with its own OS prerequisites, dependency lists, and network port requirements. The setup process involves installing packages on each VM, running a series of configuration wizards, and manually establishing connections between all these parts. It's the kind of architecture I was building in 2012.

The irony of calling this "cloud scale" is thick enough to cut with a knife. Cloud scale isn't about the raw number of VMs; it's about the operational model. This is the opposite of that. This is petting your servers, giving them names, and hoping the RAID array in your "dedicated search server" doesn't fail. Where is the immutability? The declarative configuration? The ability to horizontally scale a component with a single command or API call?

What's worse is the resource footprint for this "minimum" deployment. The docs recommend 4 vCPUs and 16GB RAM *per VM*. So we're starting at a 32 vCPU, 128GB RAM commitment before ingesting a single log or threat feed. For a "light" POC. The cost to run this on any public cloud (EC2, GCE, VMs) is astronomical for what it is, not even considering the ongoing ops tax of patching, monitoring, and backing up eight interdependent systems.

I suppose my question to the community is this: has anyone actually succeeded in operating this in a way that feels remotely cloud-native? Did you containerize it yourselves? Wrap it in Nomad? Or are we all just accepting that "cloud-scale" now means "manually cobbled together from a dozen VMs with a cloud provider's logo stamped on the invoice"?

-- Cam


Trust but verify.


   
Quote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Wait, so each VM is basically a single service? That sounds like old-school on-prem but just hosted somewhere else. 😅

I'm still learning, but isn't the whole point of cloud-native to use containers or serverless so you don't have to manually manage a VM for each piece?

Out of curiosity, does their guide at least give you Terraform scripts to spin those eight VMs up, or is it all manual CLI steps?



   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

> "modern architecture"

I think you've cracked the code on their marketing. "Cloud-native" means it used to live in a rack and now you can rent the VMs from someone else. The eight separate VMs are a dead giveaway; that's not an architecture, that's an operations checklist someone forgot to automate.

The real question for your POC isn't if it works, but what your audit log looks like after the first security patch cycle. You'll be coordinating OS updates across eight different images, each with their own quirky service dependencies. Let me know when you get to the section on configuring the internal firewall rules between all of them.


- Nina


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

You hit on the operational burden that makes me skeptical of setups like this. Managing separate OS images for what should be discreet services turns a feature update into a project.

I've seen something similar with survey platforms where the "data pipeline" was just a dedicated VM running cron jobs. The audit log after patching becomes a nightmare of inconsistent states.

How do vendors with this architecture even handle their own managed service offering? Do they just have a giant fleet of these eight-VM clusters, or is that a completely different codebase?



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

That list of mandatory VMs is a classic sign of bundling, not designing. Having dedicated servers for Postgres, Redis, and Elasticsearch is actually sensible from a performance isolation standpoint. The real red flag is needing separate "Management," "Platform," and "Message" servers - that's likely three VMs for what should be a set of containerized services with a shared orchestration layer.

The operational cost of coordinating patches and network policies across eight distinct OS images will dwarf any perceived benefit of their separation. You're not managing a cloud-native app, you're running a distributed monolith on life support.


sub-100ms or bust


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

> isn't the whole point of cloud-native to use containers or serverless so you don't have to manually manage a VM for each piece?

You've got the right instinct. The goal is absolutely to abstract away the OS and hardware management. Containers and orchestrators like Kubernetes let you treat your compute as a pool of resources for your services, not a list of individual pets.

To answer your question about the guide - nope, no Terraform in sight. It's a manual CLI walkthrough for provisioning each VM, then another 50 pages on configuring systemd services, OS-level firewall rules, and shared storage mounts between them. It feels like a 2012 ops manual that just had "cloud" search-and-replaced into it.

The crazy part is you *could* bundle all those services into containers and deploy them with a single Helm chart, even if you wanted to keep them on dedicated VMs for isolation. The fact they don't even offer that tells you everything about their platform's internal coupling.


— francesc


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Wow, that's a really eye-opening example. Thanks for sharing the actual list. I'm just starting to learn about this stuff, and seeing that breakdown makes the "cloud-native" claim feel pretty hollow.

I have a naive question, maybe. If the search and database servers are dedicated VMs anyway, couldn't they just provide those as managed services from the cloud provider? Then you'd only have to worry about the application parts. Or is the integration so tangled that it all has to be self-managed?



   
ReplyQuote
(@budget_buyer_99)
Honorable Member
Joined: 4 months ago
Posts: 359
 

That's a good point. But then you get locked into their cloud provider's managed services. The pricing tiers for those can get steep fast, and you're still on the hook for the app VMs. It just moves the cost around.

I've looked at other tools that offer a "bring your own database" option, and it always ends up being more complex than the sales pitch. The connection and security setup alone is a new project.



   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

Reading through that list, a specific aspect of the operational overhead strikes me. When each of those eight VMs has its own distinct set of OS prerequisites and dependency lists, you're not just managing the application, you're effectively becoming the systems integrator for a collection of bespoke software appliances.

This creates a hidden tax on every single change. Want to update the OpenSSL version for a security patch? You now have a compatibility matrix to solve across eight different environments, each with its own service dependencies that might break. The sales deck promises elastic scale, but the real bottleneck becomes the manual coordination of these disparate base images. It feels less like deploying a platform and more like assembling a fragile ecosystem from incompatible parts.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Yep, the "hidden tax" is the whole business model. You're not a customer, you're a free systems integrator for their tech debt.

And the fun part? The security patch matrix you mentioned becomes a blame game. When something breaks after the OpenSSL update, they'll point to your "custom environment" on one of the eight VMs. It's designed that way.

You get all the operational burden of a legacy stack, but with a cloud bill.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 2 months ago
Posts: 323
 

Oh man, that list is a perfect snapshot of the "distributed monolith" pattern. Each one of those servers is a tight coupling pretending to be a loose one.

The killer detail for me is the "Message Server (which appears to be a rebadged RabbitMQ)." That's the tell! It means the core application logic is so deeply hardcoded to a specific broker's setup and network location that they can't even let you bring your own managed RabbitMQ or SQS. You're stuck patching *their* VM.

It turns the whole "elastic scale" promise into a joke. Need more message throughput? You can't just tweak a pod replica count or scale a managed service. You're sizing a whole new VM, cloning their bespoke setup, and reconfiguring the network rules for the other seven friends. Brutal.



   
ReplyQuote
(@clarak2)
Estimable Member
Joined: 2 months ago
Posts: 143
 

Exactly! The locked-in message broker is such a giveaway. It means their architecture can't treat infrastructure as a commodity, which is the whole point of cloud design.

I bet the "elastic scale" is just a doc that tells you to provision a bigger VM for the message server, which of course has its own downtime. So scaling isn't elastic at all, it's a scheduled maintenance event.


Docs save time


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

You've nailed the operational reality there. Calling a VM resize "elastic scaling" is a real stretch, and it misses the core benefit of scaling without service interruption.

That locked-in broker also creates a nasty observability blind spot. When everything is a custom VM, you lose the granular metrics you'd get from a managed service or containerized deployment. Your "message throughput" is just CPU load on a box, not queue depth or consumer lag. So you're scaling reactively, after the problem hits, instead of from application-level signals.

It turns a design choice into an operational penalty.


- GG


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

That detailed list perfectly illustrates the chasm between marketing language and architectural reality. The requirement for a dedicated, rebadged message server is particularly telling, as it suggests a level of vendor lock-in at the infrastructure layer that contradicts the portability promise of cloud-native design.

What you're describing is essentially a pre-packaged, on-premises deployment model that has been lifted and shifted into a cloud IaaS context. The operational burden doesn't disappear; it's just transferred to the customer, who now manages the same eight "pets" but with a cloud provider's invoice attached.

This approach creates significant inertia against iteration. Any minor version upgrade for the platform becomes a major infrastructure project, requiring coordinated maintenance across all those distinct systems. It's the opposite of the agile, scalable deployment the sales deck implies.


Let's keep it constructive


   
ReplyQuote
(@charlieb)
Eminent Member
Joined: 5 days ago
Posts: 29
 

Eight VMs is just the entry fee. Wait until you see the secret menu where the "monitoring agent" and "log forwarder" are also required but live in an appendix. So you're really at ten pets to feed before you've ingested a single log.

The truly galling part isn't the count, it's that each one is a bespoke snowflake VM. That's the "cloud-native" lie laid bare. If it were actually designed for the cloud, you'd get a single, parameterized machine image for the stateless parts, and managed services for the stateful bits. This is just data center checklist architecture with a cloud IaaS bill.

The POC will eat six weeks of your team's life just on the "platform server" prerequisites.


Trust but verify.


   
ReplyQuote
Page 1 / 2