Skip to content
Notifications
Clear all

Anyone using Kling with Snowflake? How's the connector in practice?

11 Posts
11 Users
0 Reactions
18 Views
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
Topic starter   [#27785]

So they finally shipped the official Kling connector for Snowflake. About time. I've been jury-rigging exports and using the S3 stage as a middleman for months, which felt like using a horse and cart to deliver a USB stick.

I've got it wired up now, but the "real-time" sync feels like a generous term. The initial historical pull was fine, but the incremental batches seem to have a mind of their own. I'm seeing latency spikes of several hours during peak load, which kinda defeats the purpose of using Snowflake as our single source of truth for session replay triggers. Also, the schema mapping for nested event properties is... opinionated. Let's just say you'll be getting familiar with lateral flatten() if your event data is even slightly complex.

Is anyone else running this in production? I'm particularly curious about:
1. How you're handling the merge logic for updated records. The default seems to be append-only, which is a problem for user attribute updates.
2. Whether you've hit any concurrency limits or warehouse scaling weirdness when Kling fires up a sync.
3. If the cost attribution is clear, or if Kling's queries are getting lost in the sea of other warehouse costs.

The promise is there, but the implementation feels like a v1. I'm hoping I've just configured it suboptimally. What's your verdict?


Data over dogma.


   
Quote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Your point about incremental latency mirrors what we saw during our evaluation. The connector uses a polling mechanism for change data capture, which falls apart when Kling's event throughput peaks. We ended up creating a secondary "heartbeat" table in Snowflake that the connector also syncs, just to monitor the lag. It's a workaround, but it gives us visibility.

On your specific points:
- For merge logic, we abandoned the built-in options and wrote a simple post-sync stored procedure that handles SCD Type 2 logic. It runs after each batch, using the transaction timestamp Kling provides.
- Cost attribution was murky. We had to tag the virtual warehouse specifically for the Kling syncs and set up a separate resource monitor. Otherwise, its sporadic, heavy queries were impossible to track in our central dashboard.

Have you tried adjusting the batch frequency to see if it smooths out the load, or did that just make the latency spikes less frequent?


Every dollar counts.


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That initial historical pull being fine is interesting. In our manufacturing context, we had the opposite problem because our event volumes are so high and uneven. The full load of session history crippled our medium warehouse for almost a full day. We had to provision a separate, much larger warehouse just to run the initial sync, then scale it back down for the incremental polling. It makes me wonder if the connector's resource estimation is based on a simpler use case than what we're throwing at it.

You mentioned the latency spikes during peak load. We've been trying to correlate that with Snowflake's auto-suspend setting. There's a noticeable lag on the very first incremental pull after the warehouse resumes, almost like the connector has to re-establish its checkpoint. Have you seen anything similar, or is your warehouse dedicated and always running?



   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

You're spot on about the schema mapping being opinionated. We had the same flatten() party. I ended up writing a view to handle the transformation so our analysts wouldn't have to think about it.

> incremental batches seem to have a mind of their own
We saw this too. What helped was adjusting the `batch_frequency_minutes` parameter down from the default and, weirdly, making the batches *smaller*. It seemed to stabilize the lag, though it did increase the number of queries slightly.

On your Q1 about merge logic: we also ditched the built-in options. The append-only default was a non-starter for us. We built a small dbt model that runs after each sync to deduplicate and apply soft deletes, using the `_synced_at` timestamp. It's a few dozen lines of SQL, but it works reliably.


Clean code, happy life


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

The view to hide the flatten mess is a band-aid. You're now maintaining a view for a connector that's supposed to simplify your life.

And increasing query frequency to stabilize lag is classic. You're paying for their latency problem twice: once in compute, and again when the Snowflake bill shows all those small, inefficient queries.


Your stack is too complicated.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

It's not a band-aid if you treat it as the interface layer. The connector dumps raw data, your view defines the usable schema. That's a classic separation of raw and curated layers.

But you're dead on about the cost. Cranking up batch frequency creates a death spiral of small, inefficient warehouse loads. The connector's resource profile is just wrong for Snowflake's billing model. You end up paying for warehouse start-up overhead every few minutes instead of a few efficient long-running queries.


Automate everything. Twice.


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Treating the connector's output as a raw layer is a valid architectural choice, I'll grant you that. But calling it classic feels like praising the shovel for digging the ditch you didn't want.

You're still left maintaining that interface layer for a paid product that fundamentally misjudges its target platform. The real problem isn't the pattern, it's that the pattern is your only option because the connector's design makes the curated layer practically mandatory. That's not architecture, that's damage control.

And your point about the billing death spiral is exactly why these "native" connectors are so often a false promise. They're built for a checkbox on a feature list, not for the economic reality of the platform they're plugging into. When your vendor's solution creates a need for constant babysitting and cost engineering, maybe the problem is the solution.


cg


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Your "horse and cart" comment is spot on. The official connector just replaces one S3 stage with another internal one. It's still a batch-polling hack with lipstick.

On your points:
1. Default merge logic is useless. We went straight to a post-sync procedure. Don't waste time with the config options.
2. Concurrency is fine, scaling is not. The connector can't handle warehouse suspension. If it wakes up an XS warehouse to poll a massive queue, the sync will time out and retry, causing a cascade. You need a dedicated, always-running warehouse, which kills the cost argument.
3. Cost attribution is impossible without a separate warehouse. Its queries look like any other ad-hoc load. That's the real red flag.


slow pipelines make me cranky


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Your approach of building a post-sync dbt model for deduplication is solid, but it introduces a state management problem the connector should handle. The `_synced_at` timestamp is only reliable if the connector's batches are truly sequential and complete, which we've seen they aren't under load. A missed or retried batch can corrupt your soft delete logic.

Similarly, reducing `batch_frequency_minutes` stabilized your lag because it reduces the queue depth the connector has to process per poll. But as others noted, that's a tax on compute efficiency. You're trading one form of latency for another, more expensive one. Have you measured the warehouse credit burn from those smaller, more frequent queries against the value of reduced lag?



   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

Your comment about the "horse and cart" made me laugh, because it's true - sometimes the official connector just gives you a shinier cart.

You're not alone on those latency spikes. We saw the same pattern, and it turned out the connector's internal checkpointing couldn't keep pace with Snowflake's query queue under heavy load. It wasn't just Kling's throughput; it was how the connector issued its SQL.

On your specific questions:
1. We gave up on the built-in merge logic after two days. Like others, we run a manual upsert procedure after each batch, keyed on the event ID and the Kling-provided sequence number.
2. The scaling weirdness is real. If your warehouse suspends, the first poll after resume is a gamble. We set a resource monitor with a very high credit limit to prevent auto-suspend, which isn't ideal but stopped the timeout cascade.
3. Cost attribution is a black box unless you isolate it. We tagged a dedicated virtual warehouse for the sync. Without that, the sporadic, compute-heavy queries are impossible to separate from other ad-hoc loads.

Have you tried setting a fixed, larger warehouse size just for the sync? It helped stabilize our polling, but at the expense of a constant baseline cost.


~Harry


   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

> It wasn't just Kling's throughput; it was how the connector issued its SQL.

This is the key. The checkpoint query pattern is naive. It often runs a COUNT(*) or MAX(sequence) on the entire staging table before fetching the batch, which murders performance when volumes are high. The lag isn't just data transfer, it's self-inflicted query latency.

A dedicated warehouse solves the suspension cascade but validates the cost failure. You're now paying for a always-on resource to run a connector that promised efficiency.

Have you looked at the query profile for that initial post-resume poll? I bet it's a single-threaded monster.



   
ReplyQuote