Skip to content
Notifications
Clear all

Am I the only one who thinks Claw's benchmark reports cherry-pick easy tasks?

1 Posts
1 Users
0 Reactions
35 Views
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
Topic starter   [#7149]

I have been conducting a comprehensive evaluation of several analytics platforms over the past quarter, with a specific focus on their capabilities for A/B testing and statistical analysis. As part of this process, I have rigorously examined Claw's publicly available benchmark reports and technical documentation. My conclusion, which I wish to present for community scrutiny, is that Claw's performance benchmarks appear to systematically favor a subset of less computationally demanding operations, thereby presenting a skewed view of their system's capabilities in a real-world, mixed-workload environment.

To substantiate this claim, I performed a side-by-side analysis of the tasks highlighted in their "Performance Benchmarks: Claw vs. Competitors Q3" report versus the tasks required in a standard, complex conversion funnel analysis with cohort segmentation. The discrepancy is pronounced.

**Primary Tasks in Claw's Benchmark:**
* Simple count aggregation over a predefined time window.
* Single-dimension filtering on a high-cardinality user ID field (a lookup, not a complex join).
* Calculation of a basic average for a single, pre-aggregated metric column.
* "Funnel" reports that are, upon inspection of the provided query examples, linear and stateless (each step is a simple count of users who performed action B after A, without considering sessions, drop-offs within the same session, or time-bound constraints).

**Typical Tasks in a Multi-Touch Funnel Analysis (not emphasized in benchmarks):**
* Multi-join queries across event tables, user attribute tables, and session tables.
* Sequential, stateful event analysis where user progression must be tracked across a series of steps with permissible gaps and backtracking.
* Time-bound calculations between steps (e.g., "step 2 must occur within 24 hours of step 1").
* Concurrent cohort comparison within the same funnel, requiring sub-queries or complex conditional aggregation.

The benchmark queries, while valid, represent what I would classify as "easy tasks" for a modern columnar database. They avoid the more expensive operations that truly stress an analytics engine. For instance, none of their showcased queries involve:

* Window functions for calculating running totals or session rankings.
* Complex `CASE` statements inside aggregations for multi-branch cohort logic.
* Self-joins or recursive CTEs for pathing analysis.
* Geospatial calculations or array unnesting.

When I attempted to replicate a more complex analysis—specifically, a cohort-based funnel with time decay—using Claw's query interface, the performance degradation was significant. A query that completed in ~1.2 seconds for their benchmark-style count aggregation took over 28 seconds for the complex cohort funnel. The vendor's support response was that "complex, multi-step analytics are inherently more resource-intensive," which is a truism, but it underscores my point: their benchmarks do not prepare you for this reality.

My question to the community is multifaceted:
* Has anyone else conducted similar comparative performance testing with Claw under complex, real-world query loads?
* Are there specific optimization techniques or data modeling requirements (like extensive pre-aggregation) that Claw expects users to adopt, which their benchmarks implicitly assume?
* More broadly, should we as practitioners demand more representative benchmarks from vendors, perhaps a standard suite of complex operations mirroring common product analytics workflows?

I am currently compiling a more detailed table comparing query structures and execution plans, which I will share in a follow-up post if there is interest.

— Amanda


Data > opinions


   
Quote