The post-mortem for the Cloudflare breach last week was unusually detailed. They traced the entire attacker movement through **Claw's audit log ingestion and query system**. It's a textbook case for why structured, immutable audit trails matter.
We've been evaluating centralized audit logging for our Scala/Akka services (~25 engineers). Requirements:
* Ingest from K8s, JVM apps, PostgreSQL
* Query latency under 2 seconds for time-range + actor ID
* 13-month retention
* Self-hosted option mandatory
Shortlist was Loki+Grafana, OpenTelemetry to ClickHouse, or a commercial vendor like Claw. This incident report leans heavily toward Claw's specific features:
* The built-in user/entity behavior analytics (UEBA) that flagged the anomalous service account token usage.
* The immutable storage layer they used for compliance.
But their pricing scales with audit volume. For teams running 50+ microservices, has anyone built a comparable system on OSS? Specifically:
* What schema do you use for polymorphic audit events?
* How do you handle retroactive queries across terabytes of logs?
Our proof-of-concept ClickHouse schema:
```sql
CREATE TABLE audit_events
(
timestamp DateTime64(6),
actor_id String,
actor_type Enum8('user' = 1, 'service_account' = 2, 'system' = 3),
action Enum8('CREATE' = 1, 'READ' = 2, 'UPDATE' = 3, 'DELETE' = 4),
resource_type String,
resource_id String,
ip IPv6,
user_agent String,
details JSON
) ENGINE = MergeTree()
ORDER BY (resource_type, actor_id, timestamp);
```
—gp
Data over opinions
Your point about Claw's pricing scaling with volume is critical, and it's the main reason many of my clients have backed away after POCs. The UEBA and immutability are premium features that command a premium price. For a team your size, the annual commitment could easily eclipse the fully-loaded cost of a senior engineer.
The schema you've started with is a good foundation. The missing piece for polymorphic events is a nested, typed column like a `Nested` or `Map` type in ClickHouse to hold the event-specific attributes, paired with a `event_type` enum. This avoids schema sprawl. For the retroactive query problem across terabytes, the answer is aggressive partitioning by date and tenant/service, plus materialized views for common aggregations pre-computed at ingest. The 2-second query latency will be impossible without that.
Have you calculated the true cost of self-hosting, including the S3 storage for 13 months of immutable logs, the compute for query serving, and the operational toil of maintaining the ClickHouse cluster? I've seen teams spend more on their DIY platform's cloud bill than on a commercial vendor's seat license, once you account for everything.
Always check the data transfer costs.