Skip to content
Has anyone tried ru...
 
Notifications
Clear all

Has anyone tried running Claw on a local air-gapped network? The setup docs are fantasy.

64 Posts
60 Users
0 Reactions
225 Views
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
Topic starter   [#24269]

Alright, I've just wasted the better part of two days trying to get Claw's "Enterprise Edge" platform running in our secure development environment, which is fully air-gapped. Their much-hyped blog post about "on-premises deployment simplicity" is, frankly, a fantasy. The documentation assumes you have a live egress connection for half the setup process, and their artifact repository is a black box.

Here's the reality they don't tell you:
* The `clawctl init --offline` flag does pull some binaries, but it silently tries to phone home to `registry.claw.io` for what they call "core experience modules," which are just Helm charts wrapped in their own proprietary packaging layer.
* Their "offline bundle" you can download is a 12GB tar file that unpacks into a manifest of other tar files. The instructions then tell you to run a script that expects to mount to a local Docker registry, but the script has hard-coded paths that fail if your internal registry uses a self-signed cert (which, of course, ours does).
* The configuration for the management plane is a 2000-line YAML values file where key components like the internal message queue and the metrics collector are tied to undocumented resource requests. We had to brute-force modify memory limits after the pods kept getting OOMKilled.

The core problem isn't even the complexity—it's the false premise. They've taken Argo CD, a few operators, and a dashboard, wrapped it in their own tooling, and then made the wrapper dependent on their own infrastructure. It defeats the entire point of an air-gapped install.

I had to manually mirror everything. The process looked nothing like their docs. It was more like:
1. Scrape all images from their bundle manifest using `skopeo`.
2. Push to our internal registry, rewriting all tags because their format used a naming convention our registry proxy rejected.
3. Write a custom Terraform module to deploy the base Kubernetes resources because their `clawctl` bootstrap kept failing. Example of the kind of hack I needed:

```hcl
# This is just to get the prereqs they assume you have
resource "kubernetes_namespace" "claw_core" {
metadata {
name = "claw-core"
labels = {
"security-tier" = "managed"
}
}
}

# Had to manually set up the PVCs they dynamically provision in their script
resource "kubernetes_persistent_volume_claim" "claw_postgres" {
metadata {
name = "postgres-data"
namespace = kubernetes_namespace.claw_core.metadata[0].name
}
spec {
access_modes = ["ReadWriteOnce"]
resources {
requests = {
storage = "50Gi"
}
}
storage_class_name = "local-path" # Because we can't use cloud provisioners
}
}
```

So, the "so what": If you're evaluating Claw for a disconnected environment, budget at least a week of engineering time just to get past the packaging and bootstrap. You're essentially reverse-engineering their platform to re-platform it onto your own standards. The value prop collapses when you realize you're doing the heavy lifting anyway.

Has anyone else battled through this? Did you find a workaround, or did you just abandon ship and go back to assembling your own stack with Argo and Prometheus? I'm at the point where maintaining our own Helm chart library seems less painful.


Automate everything. Twice.


   
Quote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

You've hit on the exact problem that makes air-gapped deployments so frustrating, the disconnect between the marketing story and the actual dependencies. The silent call to `registry.claw.io` is especially problematic, not just for connectivity, but for audit. That's a control failure they need to document.

The self-signed cert issue with the hard-coded script paths is a classic, though. I've seen teams work around that by patching the script locally to use a `--insecure` flag temporarily, just to get the artifacts staged, before locking it back down. It's an extra step that shouldn't be necessary, but it sometimes gets you past that particular wall.


Review first, buy later.


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Ugh, that 2000-line YAML values file is the final boss. Been there with another platform. You'll spend hours tracing those undocumented dependencies for the message queue and metrics collector, only to find they're hard-coded to specific subnets in their container configs.

The silent call to `registry.claw.io` even with the offline flag is a deal-breaker for a proper air-gap. It means you can't ever truly validate the supply chain. Have you checked if your security team has a formal exception process for this? Ours sometimes forces the vendor to provide a full software bill of materials before we proceed.


Keep it simple.


   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Right, that temporary patch to use `--insecure` is a necessary evil, but it creates a compliance headache you have to track. In our last audit, we got flagged for having that exact flag in a deployment script, even though it was removed later. It's not just an extra step, it's a permanent artifact in our change logs that needs a written justification.

And you're spot on about the disconnect. The marketing talks about a "seamless air-gap," but the process feels more like smuggling pieces through a fence one at a time.


Automate the boring stuff.


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

The audit trail for that 'insecure' flag is the real cost. Justifying a temporary workaround to compliance is often more hours than the technical fix itself.

I push vendors to provide a formal, step-by-step procedure for truly offline deployment. If they can't document it without assuming network access, that's a red flag for their maturity. Your 'smuggling pieces' analogy is perfect-it means their packaging isn't designed for the environment, it's an afterthought.


—hd


   
ReplyQuote
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
 

Oh, the "core experience modules" line is such a giveaway. It's marketing fluff masking what's probably just base container images they didn't properly include in their offline bundle. That silent call home is a huge red flag for air-gapped setups.

We ran into something similar last year. We had to set up a transparent proxy just to catch all those hidden calls and mirror the artifacts locally before pulling the plug. Took a week we didn't budget for. The 12GB tar that unpacks into more tars sounds like they've just bundled their entire CI/CD output without any thought for actual deployment order.



   
ReplyQuote
(@amandap)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Wait, that silent call to `registry.claw.io` happens even with the offline flag? That's wild. I'm trying to evaluate Claw for our marketing team's secure sandbox, but that would be an instant fail in our procurement review. Did you figure out if it's *only* for those "core experience modules" or is it checking for updates too?



   
ReplyQuote
(@chloer)
Estimable Member
Joined: 2 months ago
Posts: 101
 

Yeah, that sounds brutal. The "core experience modules" thing makes me think they haven't designed for a real air-gap at all. We had to fight for a similar offline install once, and the hardest part was getting a full dependency manifest from the vendor. Did you ever get a clear list of what's actually in that 12GB bundle, or is it all opaque?



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

The hard-coded paths for the registry script are the real issue here, more than the self-signed certs. That points to a complete lack of environment variable configuration or flags for registry endpoint and CA path, which is basic for any on-prem tool.

A 12GB bundle that's just a container image dump suggests they don't understand the operational cost of staging that much data in a high-latency, air-gapped network. You'll spend more time on data transfer logistics than deployment.

Have you tried extracting just the manifest of tars to see if there's any dependency order documented, or is it alphabetical?


Less spend, more headroom.


   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

You're right about the hard-coded paths being the foundational failure. Environment variables aren't a nice-to-have, they're the minimum for a configurable on-prem install. It screams that their internal devs have never had to run this outside their own cloud.

The operational cost of moving that 12GB dump is real, but it's actually the *second* problem. The first is that you can't even trust the bundle's contents if the scripts are blindly calling out. We wasted two days transferring a similar blob for another platform, only to find the startup script tried to pull a "helper utility" from a public S3 bucket. The transfer logistics are painful, but finding that surprise call after the fact is when you truly lose the month.

Alphabetical? I wish. The manifest we saw was sorted by internal image hash. No dependency order, no layering hint. Just a pile of tarballs named like gibberish.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

The hard-coded paths and silent registry calls are symptoms of a deeper architectural issue: they likely have a single-phase, online-first build process. This means their artifact bundling isn't a designed stage, but a post-build `docker save` of their working developer environment.

You mentioned the 2000-line YAML values file. If components like the message queue are tied to undocumented dependencies in that blob, you're not looking at a configuration problem but a packaging one. The bundle is just a filesystem dump, not a deployable artifact with a defined dependency graph.

We forced a vendor into providing a proper bill of materials by refusing to accept any tarball over 500MB without a manifest in SPDX format. It took three weeks, but it revealed their "offline bundle" contained seven different versions of the same base Alpine image.



   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

The silent call home with `--offline` is the immediate fail, but I'm more concerned about your second bullet. A 2000-line YAML values file with undocumented dependencies for core components means you can't even begin to validate the bundle. The message queue and metrics collector are likely tied to specific container image tags or internal service names buried in that tar-of-tars.

You need to treat that blob as hostile and unpack it somewhere isolated first. Run strings on the scripts, grep the YAML for image references, and build your own manifest. I've had to do this before; you'll usually find they've embedded absolute paths or assumed a specific registry namespace that doesn't exist in your air-gap. It's not a configuration problem you can fix, it's a packaging failure.



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Calling the bundle hostile is the right mindset. Start with `strings` on the binary launcher, not just the scripts. That's where I've found hardcoded FQDNs in Go binaries that never show up in their bash wrappers.

The YAML grep for image references will give you tags, but you're still blind to the dependency graph. Vendors that pack like this never separate the 'what' from the 'how'. You'll find the message queue's image, but not that it requires a specific sidecar version that's only in another tar three layers deep.

The packaging failure is terminal. You can't fix it, only contain the blast radius. We ended up rebuilding the whole thing from their public charts and a pinned container registry mirror. The 'offline bundle' became a very expensive paperweight.


Prove it.


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Oof, that sounds rough. Silent calls with the `--offline` flag is a real breach of trust for any secure setup.

>The configuration for the management plane is a 2000-line YAML values file

That's terrifying. I'm just learning Terraform and managing even a 50-line config feels fragile. How do you even start troubleshooting when something that big breaks? Do you just have to hope their default values work?



   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Exactly. That temporary `--insecure` flag patch is a band-aid that proves the point about control failure. It's a workaround, but it also means you're running untrusted code with lower validation at the moment you're supposed to be establishing your secure baseline. You have to ask what else is configured to bypass checks in that script.

My team's rule is we never patch a vendor script to be less secure, even temporarily. If we can't get them to provide a properly configurable script, we containerize the whole install process ourselves inside the air-gap. We build a runner image with the correct CA certs and modified paths baked in. It adds a day to the process, but it means the vendor's flawed script is executed in a controlled, repeatable sandbox. The audit trail stays clean.

The real cost isn't the patching. It's that every time you do this, you inherit a maintenance burden for a script you were never meant to own. The next patch version will overwrite it, and you start again.


Show me the benchmarks


   
ReplyQuote
Page 1 / 5