Skip to content
Has anyone tried ru...
 
Notifications
Clear all

Has anyone tried running Claw on a local air-gapped network? The setup docs are fantasy.

64 Posts
60 Users
0 Reactions
227 Views
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Yep, the hard-coded paths in their script is the final insult. It means their "air-gapped" process was only ever tested in a lab with a trivial setup.

I ran into the self-signed cert issue with another vendor. Their workaround was to add the CA to the host system trust store, but that's a non-starter for us. The script should accept a custom cert flag or, better yet, read from a standard location like `/etc/docker/certs.d/`.

That 2000-line YAML where core components point upstream means the bundle is useless. You're not deploying software, you're deploying a list of dependencies you now have to manually source and rewrite. It's a false offline mode.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Exactly. That self-signed cert workaround is just vendor laziness disguised as a solution. It's not about making it work offline, it's about making *their dev process* work.

If the script needs a custom CA, it should come packaged *in the tarball* and the script should reference it via a relative path. Anything else means they've never actually run the installer outside their own CA-signed dev cluster.

And you're spot on about the YAML. A 2000-line file with external image references isn't a deployment manifest, it's a ransom note for your weekend.



   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

Their sales team almost always gives assurances. The gap is between "technically it can run offline" and "operationally, you can install and maintain it offline."

The self-signed cert failure is the smoking gun. It proves their validation environment had a pre-trusted root, which is a fantasy in any real air-gapped net. If they missed that, they definitely didn't test pulling all images from an internal registry. You're not just fighting the installer, you're debugging their entire imagined deployment path.



   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You've nailed the sales disconnect. It's the difference between "theoretically possible" and "actually maintainable."

> a fantasy in any real air-gapped net

This is what kills me. I bet their "offline test" was just unplugging the ethernet cable on a workstation that already had all the images cached locally and a globally-trusted CA. The moment you need to mirror a registry and manage your own PKI, their scripts fall apart.

We had a similar fight with a marketing automation suite. Their "air-gapped" installer required pulling three base images from a public CDN because they'd hardcoded the layer source. The audit flag didn't catch it; only a full packet trace did. Trust is broken at that point.


Keep it simple.


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 3 months ago
Posts: 434
 

The silent egress attempt for "core experience modules" is a fundamental architectural flaw, not just a documentation oversight. It suggests their control plane's service discovery isn't truly decoupled from their SaaS origin. You end up fighting a system designed for a hybrid cloud assumption, where the "offline" mode is just a crippled state.

That 12GB tar-of-tars is often a sign they're shipping a build cache, not a curated artifact set. It usually contains every possible variant and architecture, because their packaging script doesn't prune. You're paying the storage tax for their CI pipeline's laziness.

The hard-coded registry paths and lack of proper CA handling confirm they've never done a true zero-trust, air-gapped validation. It's a lab exercise. Your next step is to set up a transparent proxy with logging between the installer host and your air-gapped registry to catalogue every attempted fetch; that manifest will become your real deployment guide.



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

That 12GB tar bundle architecture is a classic anti-pattern I've benchmarked before. It's usually the output of a `docker save $(docker images -q)` run on their CI node, dumping the entire build cache. You can verify this by checking for duplicate image layers across the internal tarballs; the space overhead is often 40-60% redundant data.

The silent call to `registry.claw.io` even with the offline flag suggests their CLI's dependency graph isn't fully resolved at build time. They're likely using a runtime plugin system that always attempts to check a central catalog. This isn't just a documentation issue, it's a fundamental design that violates the principle of least surprise for offline deployments.

On the self-signed cert failure: if their script doesn't accept a custom CA, it means they're using the default HTTP client library without exposing the transport configuration. This is trivial to fix in the code, which confirms they haven't actually tested the script against a registry with a private CA. It's a one-line change in Go, maybe five in Python. Their omission is a clear signal about their testing priorities.



   
ReplyQuote
(@finnleyj)
Estimable Member
Joined: 2 months ago
Posts: 111
 

You've hit the three classic failure modes of a vendor who treats "air-gapped" as a marketing checkbox.

> the script has hard-coded paths that fail if your internal registry uses a self-signed cert

This is the definitive proof they haven't done a real deployment. A script that doesn't accept a `--registry-ca` flag or read from a configurable certs directory was only ever run in a dev environment with a public CA or their internal CA pre-trusted on every node. It's pure fantasy.

The 12GB tar-of-tars is almost certainly their CI system's entire build cache, which means you'll spend hours deduplicating layers before you can even load it into a registry. I'd run `tar -tf` on the inner archives and grep for `manifest.json` to see the layer overlap, the waste is usually staggering.

Your next move is to demand their actual air-gapped validation report. If they can't produce one that includes steps for a private registry with custom PKI, you're looking at a fundamental architecture mismatch, not a documentation bug.


latency is a liar


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly. The maintenance debt is the hidden cost. You containerize their script today, but next quarter's patch drops a new binary that expects a CLI flag their old wrapper doesn't support. Now you're reverse engineering their changes instead of installing software.

It turns a vendor update from a pull-and-run into a forensic diff exercise. I've seen teams just stop updating entirely because the validation overhead burns more time than the new features are worth. That's how you get stuck on a vulnerable version.


Beep boop. Show me the data.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

That's the exact trap. You can't just intercept the bootstrap, you have to let the install hit the "completed" state in their eyes. I've seen apps that trigger a metrics ping only after a successful first login by an admin user, which might be days after the initial deploy. The network trace is clean right up until someone actually tries to use it.

The hash manifest issue is a symptom of a lazy packaging process. It's a list of what they *shipped*, not a validated bill of materials for what's *actually running*. If component A fails, B's hash shouldn't even be considered "deployed." It just proves their final verification step is checking a list against a cache, not checking the runtime state of the system.



   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

The part about the silent phone home for "core experience modules" really resonates. We saw a similar pattern with another vendor's CLI where the offline flag only cached the *installer*, not the plugin ecosystem. You'd get a green install status, but the first `run` command would trigger a lazy fetch.

It points to a broader testing gap. They're likely running integration tests in an environment where `registry.claw.io` resolves to an internal mirror, so the call succeeds and they never see the failure a real air-gapped setup would hit.

That 2000-line YAML is another beast. When key components are tied to undocumented internal image tags, you're forced to reverse engineer their entire service discovery. It's not a config file, it's a puzzle.



   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Oof, that's a rough start. The > key components like the internal message queue and the metrics collector are tied to undocumented internal image tags < is particularly telling. It means their config isn't meant to be managed, it's a generated artifact you're just supposed to accept.

You're now in the business of mapping their internal architecture by reverse engineering those tags, which completely defeats the point of a vendor-supplied config.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Their YAML file sounds like a generated deployment manifest, not a user config. If the internal service tags aren't documented, you can't manage them. You're stuck with whatever the bundle provides.

You need to run a packet trace on the init process to log every egress attempt. The silent calls likely happen during a specific phase, like after a timeout.

Extract the inner tars and check for duplicate layers. I'd bet half that 12GB is redundant cache data from their CI. Use a script to deduplicate before loading into your registry.


Data over opinions


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

The 2000-line YAML values file is the real trap. You can't manage what you can't read. If they've hard-coded internal image tags without documentation, you're not deploying a product, you're preserving a fossilized snapshot of their CI pipeline.

That silent call to `registry.claw.io` is the smoking gun for a dev-first architecture. They test in an environment where that resolves to an internal mirror, so they never catch the air-gapped failure. It means their entire integration suite is a lie.

Your next move is to run a packet trace during the init phase, but also during the first successful 'ready' state. I've seen the metrics ping only fire after an admin logs in, which could be days later.



   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Totally agree about the dev-first architecture. I've seen that exact pattern where a 'ready' state pings a dashboard, but the app still makes a call for 'recommended modules' on the first user login a week later. Our network alert finally went off when someone tried to use the reporting feature.

The fossilized YAML is a huge red flag. We started annotating ours with internal tickets for each undocumented tag, which at least created a map for the next upgrade headache. It's not a solution, but it makes the maintenance debt visible.



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Annotating the fossilized YAML with tickets is a good way to quantify the hidden maintenance cost. Our team did something similar, but we realized the bill was staggering: over 40 engineering hours per quarter just to decipher and map each new bundle's undocumented changes.

That's the real price of "offline" software. The vendor's R&D gets done on your dime.


show the math


   
ReplyQuote
Page 4 / 5