Skip to content
Notifications
Clear all

Did you see GCP's new Dataplex? Is it just another data lake wrapper?

1 Posts
1 Users
0 Reactions
39 Views
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
Topic starter   [#19184]

Spotted the announcement in my feeds this morning. Another week, another "unified data platform" promising to end the eternal struggle between the data lake and the warehouse. GCP's Dataplex is being pitched as the intelligent fabric that ties BigQuery, Cloud Storage, and Dataflow together with governance and discovery. My immediate reaction, after two decades of watching this cycle repeat, is profound skepticism. It looks for all the world like a management console and a metadata layer wrapped in some new jargon.

Let's break down what they're actually selling versus what you could already do, albeit with more manual effort. They're offering:
* **A "lakehouse" metaphor:** This is just a logical grouping of GCS buckets and BigQuery datasets under a single label. You could already point these services at each other.
* **"Intelligent" data discovery:** Automated metadata and classification. Sounds like a managed Data Catalog with some new rules engine bolted on. Useful, but not novel.
* **Unified governance:** Centralized IAM and policy tags. This is arguably the core value, as managing fine-grained access across BigQuery, GCS, and Pub/Sub is a headache. But it's a policy manager, not a technical breakthrough.
* **Integrated analytics:** The promise of a single pane to run Spark, query via BigQuery, etc. This is the fanciest wrapper of allβ€”it's orchestration glue for services that already exist.

The real question for this forum is about the cost-per-unit and lock-in. They aren't giving you a new engine; they're charging you for the privilege of using their existing engines in a specific, orchestrated way. My fear is the pricing model: you'll pay for Dataplex processing units on top of the underlying storage (GCS) and compute (BigQuery, Dataproc) charges. It's the classic cloud vendor play: solve complexity by adding another managed layer with its own opaque billing line item. Has anyone run a TCO comparison yet between:
1. A manually stitched setup using Terraform for infra, a separate Data Catalog, and careful IAM roles?
2. The full Dataplex "managed experience"?

I'm also deeply curious about the practical limits. Show me the config for a non-trivial, production-grade pipeline that handles late-arriving data and schema evolution. Does their "automated data quality" handle the mess of a real CDC stream from a legacy PostgreSQL source, or does it fall over the moment it sees a JSONB field with nested arrays? I'll believe it when I see a `dataplex_task` configuration that isn't a toy example.

```yaml
# This is the kind of gritty detail I want to see.
# How do you define a 'data domain' as code? Is it just a YAML front-end for a gcloud command?
apiVersion: dataplex.googleapis.com/v1
kind: Task
metadata:
name: production-cdc-ingestion
spec:
spark:
pythonScriptFile: gs://my-bucket/scripts/debezium_spark_processor.py
# How are the dependencies handled? Container image? Inflexible requirements.txt?
trigger:
type: EVENT
pubsubTopic: projects/my-project/topics/postgres-cdc
# Where are the retry policies? The dead-letter queues? The alerting integrations?
```

So, has anyone kicked the tires? I'm looking for war stories, latency numbers for metadata propagation, and most importantly, the first unexpected bill. Is this a legitimate simplification for complex multi-service governance, or just another coat of paint on the same old wall?

-- old salt



   
Quote