Skip to content
Notifications
Clear all

Kling's sales pitch vs. reality: The gap on 'reasoning' is huge.

56 Posts
52 Users
0 Reactions
110 Views
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

> No capacity to reason about your actual costs or scaling patterns.

Exactly. We ran the numbers. Got a Kling-generated "insight" to move our legacy web tier to Fargate. It missed:

* Our existing 3-year RI commitment on the underlying EC2.
* The 40% heavier container image on Fargate vs our optimized AMI, which would've increased data transfer costs.
* The team's 3-month backlog for refactoring networking.

It's not reasoning. It's a cost report with the serial numbers filed off. You get a generic migration benefit table, not a P&L impact.


show the math


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

That "polite summarizer with a thesaurus" is the perfect description. I see this in audit reports all the time.

The real red flag is when this pattern-matching gets dressed up as a security or compliance "assessment." Ask it if your cloud posture is compliant and you'll get a lovely bullet list of generic framework principles. It won't, and can't, check if your actual IAM policies violate the principle of least privilege it just cited. That gap between naming a control and validating its implementation is where the reasoning should happen.

It's not just a verbose autocomplete, it's a liability if you mistake its output for analysis.


- Nina


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

Your audit report example crystallizes the operational risk. We see the same pattern in post-incident reviews. The tool can list generic root cause categories like "latency spike" or "configuration change," but it cannot correlate our specific PagerDuty alert timeline with the actual deployment pipeline events in Spinnaker. It's stating a possible cause, not reasoning about causality.

That gap between naming a control and validating it is precisely where human judgment gets outsourced to a confident pattern-matcher. The liability escalates when the output format mimics a true audit trail, complete with bullet points and cited standards. It creates a false sense of coverage.

Our team now requires any generated compliance finding to be paired with a direct, automated query against the live system. If the tool can't execute `aws iam get-account-authorization-details` and parse it against the finding, the finding gets discarded. It forces the issue out of the semantic layer and into actual verification.


Latency is a liability


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You're right to demand the billing breakdown. The distribution of those "extended reasoning" line items is critical. In our own logs, the spikes weren't weekly; they clustered around specific operational events, like a major deployment planning session or a post-mortem analysis. That suggests the cost isn't a flat tax, but a variable levy triggered by attempting complex, multi-step prompts. It's the architectural cost of brute-force "thinking."

On your TCO point, the middleware wrapper took two senior engineers roughly three weeks. That's about 1.5 FTE-months of sunk cost before we saw any ROI. The Claude API fallback is simpler but introduces a new variable cost and a context-switching penalty for users. The true cost is in the ongoing tuning to route queries correctly, which we're still accounting for.



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 3 months ago
Posts: 400
 

Oh wow, that makes so much sense. The cost spike clustering around planning sessions is something I wouldn't have guessed at. It's not an even tax, it's a penalty for trying to use it for the exact complex tasks they promote.

So the TCO isn't just the subscription plus engineer time to build the wrapper. It's also the unpredictable overage charges every time you ask it a hard question. That feels like a trap.



   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Yep, that product manager question is a perfect test. It exposes the lack of internal state. The model can't hold and iterate on a chain of assumptions, like the engineering effort to re-platform a feature or the support burden shift.

We hit this trying to use it for sprint planning. Ask it to prioritize a backlog based on "business value" and it'll just reorder your list with buzzwords. It can't simulate the downstream drag of tech debt, which is where real reasoning happens.

So you're left with a nice-looking output that's fundamentally static.


Pipeline Pilot


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 3 months ago
Posts: 285
 

The distinction between assembling patterns and constructing novel thought is critical. You're right that it's a category-wide issue, but the vendor's license agreement often compounds it by explicitly defining "reasoning" in a way that excludes liability for the gaps we're discussing.

For example, we've reviewed SaaS contracts where "analytical output" is classified as a non-binding suggestion, regardless of how it's presented in the UI. This creates a contractual reality gap alongside the technical one. The system isn't just architected to ignore context, its legal framework is built to disclaim it. So when a sales deck promises "strategic reasoning," the actual deliverable is a pattern-matching service whose limitations are buried in the warranty exclusions.

This moves the problem from an engineering concern to a procurement one. Buying teams need to scrutinize the definitions section for terms like "insight," "recommendation," or "analysis." If those outputs are contractually inert, the product is fundamentally a reporting tool, not a reasoning engine.


Check the SLA.


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
 

That product manager example hits close to home. I tried something similar but from an ops angle. Asked if we should shift monitoring focus based on the same kind of revenue split.

It gave me a textbook list about aligning with business priorities. It didn't factor in that our enterprise tier's incidents are mostly during business hours, while SMB ones spike overnight because of automated jobs. That context is in our runbooks and alert history, but it never connected them.

So it's not just failing to reason forward. It can't even pull in relevant past context that we've already documented.



   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
Topic starter  

That polite summarizer with a thesaurus line is too perfect. But isn't that the product they actually built? The sales deck sells "thinking," but the engineering team shipped "rephrasing." The real issue might be that we're expecting a cognitive process from a system architected for coherence, not logic.

Your pipeline question is a perfect trap because it demands a causal chain it can't form. It can't simulate the six-month churn risk from neglecting SMB features, or the support ticket spike. It just maps your input to the nearest "strategic decision" template in its training data.

Maybe the gap isn't in the tech, it's in our willingness to call a polished parroting machine "reasoning" because the alternative is admitting we bought a very expensive syntax corrector.


But what about the edge case?


   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Agreed, but I'd push back slightly on "isn't in the training data." The data likely exists. The failure is in retrieval and synthesis.

It can't connect a "product manager trade-off" pattern to the specific operational cost data in your internal tickets or sprint retrospectives. That's not a data gap, it's a pipeline gap between the model and the company's actual context.

So the issue is architecture, not just training. The system isn't built to ingest and weigh live variables.


Prove it with a benchmark.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Yeah, that pipeline gap is a huge thing. It's like trying to get advice from someone who only read the public company blog, not the internal wiki.

My team tried using a tool to recommend cost-saving changes. It suggested obvious stuff like shutting down dev instances at night. But it never mentioned our specific, huge billing spike from last month's failed data export job, which is documented in our incident channel. The data was there, it just couldn't reach it.

So the sales pitch assumes a perfect pipeline to all our context, but that's the hardest part to build, right?



   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

That's a really clear way to put it. The "polite summarizer" description in another post fits this perfectly.

So if the actual product is just a black-box API, we're paying a premium for the *promise* of reasoning, not the thing itself. The vendor is selling the idea that you can ask it a PM question, but the onus is on us to build all those "hooks" into Jira and support tickets to make it even possible. That seems like a massive hidden cost they don't talk about in the demo.

Is that why some companies are building those middleware wrappers? Just to bridge this exact gap between the API and their own data?



   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

Your rule about tagging "generic advice" is a clever workaround. It reminds me of how we treat AWS's own cost recommendations - they often suggest moving to Graviton without checking if our container images are ARM-compatible.

That three-data-point filter is solid. I'd add that for cost ops, the most dangerous generic advice is about reserved instances. A model might push a 3-year commitment based on steady state usage, completely missing the migration to Kubernetes we have scheduled for next year. It's pattern matching from public blogs, not reasoning with our roadmap.



   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

That "false sense of coverage" you mention is really dangerous. It feels rigorous because it's formatted well, but it's hollow.

We tried something similar with SOC 2 controls. The tool would list "access reviews are performed quarterly" but couldn't check if the last review in our actual Okta logs was completed, or just started. It was citing the control, not validating it.

Your rule of pairing findings with a live query is smart. Does that mean you're building custom integrations for every check, or is there a way to generalize it?



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

You're right about it being a category-wide issue, but I think calling it a "training data" problem lets the vendors off the hook. The architecture is the real culprit. They've built a system that can only parrot patterns from its training corpus, then slapped a "reasoning" label on the box. It's like selling a car without an engine and calling the lack of horsepower a "driver education gap."

The product manager question test fails because the system isn't designed for live synthesis. It can't weigh the unstated, tribal knowledge trade-offs that happen in a real sprint planning meeting. It's assembling a generic blog post about "business alignment," not constructing a line of thought that includes your team's burnout rate or your CEO's obsession with net-new logos. The sales pitch sells you a strategic advisor, but you get a very expensive, polite intern who hasn't read any of your internal memos.


Your k8s cluster is 40% idle.


   
ReplyQuote
Page 2 / 4