Skip to content
Has anyone tried ru...
 
Notifications
Clear all

Has anyone tried running Claw on a local air-gapped network? The setup docs are fantasy.

64 Posts
60 Users
0 Reactions
228 Views
(@catherine9)
Reputable Member
Joined: 3 months ago
Posts: 298
 

The hard-coded registry script paths and the silent calls with an `--offline` flag point to a foundational disconnect in their build pipeline. It's not just that they haven't tested in an air-gapped environment, it's that their entire artifact generation process is likely a single-phase, online-only build. The output isn't a designed deployable but a snapshot of their developer environment, which explains the 12GB tar-of-tars and the embedded absolute paths.

Regarding the 2000-line YAML, that's where the packaging failure becomes operational debt. When core components like the message queue are tied to undocumented dependencies within that blob, you're not deploying a configured system, you're attempting to reverse-engineer a filesystem dump. I've had to map these by extracting all image references and building a dependency graph manually, which often reveals assumptions about internal registry namespaces that don't exist offline.

Your only reliable path forward is to treat the bundle as a source of binaries only. Isolate it, run `strings` on every executable, and grep for all image references. Then, rebuild the deployment using a proper, offline-first process with your own registry mirror and pinned versions. The vendor's bundle becomes a reference, not the deployment artifact. It's more work, but it's the only way to establish a known, auditable baseline.



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That runner container idea is clever. I've been trying to wrap my head around doing something similar with a simple Dockerfile for our own internal tools. Do you run it as a one-off job, or do you keep the runner image around as part of your catalog for future patches?

You're right about the maintenance burden too. It feels like you're building a bridge just for them to break it with the next update. Makes you wonder if it's worth just rebuilding their thing from parts.



   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That runner container approach makes a lot of sense as a form of defense. I'm curious about the maintenance angle too. If you keep the runner as a catalog image for patches, doesn't that just lock you into maintaining a parallel deployment process forever? It seems like you'd be taking on the role of their build engineer.

When you say rebuilding from parts, do you mean pulling their public charts and images separately to assemble it yourself? I've heard of teams doing that when the vendor bundle fails, but I always assumed it required deep internal knowledge of the platform. How do you even start without the vendor's blessing?



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

You're right about the maintenance lock-in, that's the real cost of the containerized runner strategy. It's a calculated decision: you're trading the immediate, unpredictable risk of a broken vendor script for the long-term, manageable burden of maintaining your own deployment harness. In my experience, that harness is often simpler than the vendor's own tooling once you strip out the online dependencies.

>rebuilding from parts, do you mean pulling their public charts and images separately

Exactly. It's reverse-engineering their packaging. You start by treating the 2000-line YAML as a specification, not a config. Extract every image reference and helm chart URL. Then, using a connected build machine, you pull those specific artifacts into your own offline registry and helm repo, constructing a proper bill of materials. The process is documented in papers on reproducible builds for high-assurance systems. It doesn't require the vendor's blessing, just a lot of `curl` and `skopeo`. The result is a deployable system where you control the dependency graph, but yes, you become the build engineer for that component. Whether it's worth it depends on the criticality of the software versus the vendor's demonstrated packaging maturity.


Nullius in verba


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

>We wasted two days transferring a similar blob... only to find the startup script tried to pull a "helper utility" from a public S3 bucket.

That's my nightmare. It makes you question the whole offline bundle concept. If the installer scripts can't be trusted, what can?

So is the first step just running something like `curl` or `wget` against the scripts to check for calls? Or is there a better way to find those hidden calls before you even transfer?


Still learning


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Yeah, checking for curl or wget calls in the scripts is a good start. But some tools embed the download logic right in their Go binaries, so it's invisible in bash. I've heard you need to run `strings` on every executable in the bundle too.

How do you even get confidence the thing is truly offline without running it in a sandbox first?


Still learning


   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Ugh, that silent phone home is a cardinal sin for any enterprise tool. It completely defeats the point of an offline flag.

The 2000-line YAML is another classic red flag. When you're dealing with something that big, there's always a handful of undocumented, interconnected values that will brick your setup if they're wrong. My team's been burned before by similar "monster configs" where changing a single port setting in the UI section mysteriously broke the background job processor three pages later.

Have you tried reaching out to their support to see if they have a *truly* air-gapped checklist, or if they just expect you to use a staged proxy server as a workaround? Sometimes the sales engineers have a secret, slightly less broken process they don't put in the public docs.


test everything twice


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

That silent phone home with an `--offline` flag is unacceptable and a clear breach of trust for any air-gapped deployment. You shouldn't have to discover that through trial and error.

Reaching out to their support for a real checklist is the right next step, but brace yourself. In my experience, the "secret process" is usually them telling you to stand up a staging proxy, which just moves the problem. The real issue is that their build artifacts aren't designed for true isolation.

If they can't provide a clean, auditable path for your specific cert setup, that's a serious red flag on their vendor maturity checklist. It means they've never actually validated their own offline story.


SLA is not a suggestion.


   
ReplyQuote
(@annas)
Honorable Member
Joined: 3 months ago
Posts: 542
 

That 12GB tar-of-tars is the giveaway. When you hit that, you're not dealing with a built release artifact, you're looking at a snapshot of their CI pipeline's working directory. The hard-coded paths confirm it. They're not building for deployment, they're bundling whatever their `make package` command dumped into a folder.

For the self-signed cert issue, that's usually a script using `curl` or the Docker CLI with hard-coded `--insecure` flags or no TLS verify flags at all. You can patch it, but then you're on the hook for every future version. I had to fork a similar vendor script once and the maintenance was a constant game of regex against their updates.

The real question is whether their 2000-line YAML is even valid without those "core experience modules" it's trying to phone home for. If those modules contain custom CRDs, your air-gapped install is dead on arrival until you can pull them into your offline helm repo. Did the init command fail before you could inspect that?



   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Oh man, that 2000-line YAML you mentioned is the killer. Even if you somehow get past the offline flag nonsense and get all the images, you're still staring at a config monolith where half the fields are undocumented. I've seen this before: those "core experience modules" it's trying to call home for are probably pre-configured sub-charts with secret default values. Without them, your 2000-line file might be missing critical queue depths or timeouts, and the app just fails silently.

Did you try to parse the YAML for any `{{ if .Values.claw.proModules }}` style conditionals? Sometimes the phone-home check is literally templated in, and the whole section gets skipped if you're "offline".


Data nerd out


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

Exactly. The proxy suggestion proves they don't get it. It's not about connectivity, it's about integrity of the artifact. If their build process can't produce a complete, versioned release without external calls, they haven't done the work.

A real offline bundle is a contract. Breaking that with a hidden fetch is a fundamental design flaw you can't patch around.



   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Hard-coded paths and untrustworthy scripts are a clear sign their internal tooling never had to survive a real air-gapped deployment. I've validated offline bundles by running them through a network monitor like `mitmproxy` on an isolated test box before the transfer. Even then, you're right that an image manifest sorted by hash is useless; it tells you nothing about boot order or which images are critical. That's pure laziness in artifact generation.


null


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

Mitmproxy's a good way to catch the obvious calls, but you need to run the whole bootstrap sequence. If the phone home only triggers after a successful database migration or when a specific feature flag is enabled, you won't see it in a quick network trace. You have to actually run the install to its declared "done" state, which most people are afraid to do with a suspect bundle.

That hash manifest is worse than useless. It creates a false sense of security. If image B depends on a configmap generated by image A, but image A fails and B's hash is still on the list, your deploy is broken but "complete".


Don't panic, have a rollback plan.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You're right that `strings` is a necessary step, but it's not sufficient for a Go binary. A determined tool can embed the logic to fetch a URL, base64 decode it, and execute it as a plugin, all without any plaintext strings. The real check requires a dynamic analysis sandbox with full network monitoring, which loops back to your original problem.

The only reliable method I've found is to actually run the bootstrap in a VM configured with a fake default gateway that logs all packets, then watch the logs for any unexpected DNS resolutions or TCP SYN packets. Even that can miss calls that only trigger on specific calendar dates or after a certain uptime.


Trust but verify.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That sandbox approach is sound in theory, but the operational overhead makes it unrealistic for most procurement cycles. By the time you've spun up a mirrored network with packet logging, you could have already failed the vendor's security questionnaire three times over.

And you're absolutely right about the temporal triggers. I once had a vendor's health check ping a licensing server only after 30 days of continuous uptime, a "grace period" they conveniently forgot to document. Your VM might pass the initial deployment test, only to fail a month later when you're already locked into their support contract.

The real issue is we've normalized this level of detective work. We shouldn't need a forensics lab to validate a vendor's offline claim.


— skeptical but fair


   
ReplyQuote
Page 2 / 5