Skip to content
Notifications
Clear all

Am I the only one who finds the 'continue' button after a cut-off response annoying?

33 Posts
32 Users
0 Reactions
68 Views
(@emilyc)
Reputable Member
Joined: 3 months ago
Posts: 161
Topic starter   [#25543]

Okay, maybe it's just me being a newbie, but this drives me a little nuts 😅. I'm using Poe to brainstorm some basic landing page copy, and the bot will give a great suggestion, then just... stop mid-sentence. And there's that 'continue' button.

It feels like waiting for a second part of a text message. My brain is already running with the idea, and then I have to click and wait again. Sometimes the 'continued' part feels like a totally separate thought and it breaks my flow.

Is there a trick to avoid this? Or do you all just get used to it?



   
Quote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

It definitely isn't just you, and I feel that same interruption. What I've noticed is that the break often happens right when the response is getting to the practical part, the actual 'how-to' or the concrete example. You're left with the theory and have to prompt again for the application.

In my case, I was asking about reconciling inventory valuation methods between systems, and it stopped right before listing the specific fields to map. That pause completely derailed my train of thought as I was preparing to take notes.

Have you found that rephrasing your initial question to be more specific about wanting a single, complete response helps at all, or does the cut-off seem arbitrary?



   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Right? The cut before the concrete example is the worst part. I've had it happen when asking for step by step automation triggers, and it stops just as it's about to list the third app. That specific detail is exactly why I asked!

I don't think rephrasing for a "complete response" helps much. The cut-off seems pretty arbitrary and tied to length. My workaround is just to immediately hit 'continue' while the initial idea is still fresh, but it's still a friction point.


dk


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Totally feel that about the automation triggers! Like, you're mentally ready to write the cron job or set up the webhook, and then... click.

Your point about hitting 'continue' immediately is smart, I'll try that. I wonder if it's a resource thing on their end, limiting the initial output burst.



   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

It's not just you, the flow break is real. I hit this constantly when getting AWS service config suggestions. The useful conditional logic or the specific CLI flag is always after the button.

My workaround is asking for the output format upfront, like "give me the complete Terraform snippet in one go." It cuts the initial reply a bit shorter, but you usually get the full context without the extra click. Does that work for landing page copy, or is the structure too loose?


Ask me about hidden egress costs.


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

Asking for a specific format might shift the cut-off, but it doesn't solve the underlying problem. You're just pre-paying the cognitive tax by structuring your prompt around the tool's limitation instead of your task.

The real issue is accepting that we have to work around this at all. If I'm paying for a service, I shouldn't need 'workarounds' for basic conversational flow. It feels like getting a truncated spec from a vendor and having to file a ticket for the appendix.


Show me the unit economics.


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Yeah, the cut before the practical part is exactly it. I had it stop right before giving the specific CIDR for a VPC subnet example. My note-taking app was open and everything 😅

>rephrasing your initial question to be more specific
I tried "please include a full terraform example in one block," but it still sometimes cuts off. It does feel arbitrary, like it's just counting tokens.



   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Ugh, the CIDR block example hits home. I was just asking about a simple three-tier architecture and it stopped after defining the public subnet. The actual IP range I needed was in the next click.

It really does seem like a token count, like you said. Even asking for a "complete example" feels like a gamble. Maybe they're trying to save on processing for the initial response? Not sure if that makes sense.



   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

That inventory valuation example is a perfect one - it's exactly where the practical, actionable details live, and being cut off there is more than just annoying, it's disruptive to the work process.

I've found the cut-off to be far from arbitrary; it seems quite deterministic based on token count or response time, which ironically makes it *more* frustrating because you can almost predict when you'll lose the useful bit. Your strategy of hitting 'continue' immediately is the closest thing to a workable habit I've developed, too.

Have you tried a more direct prompt right at the start, like "Please walk through the entire field mapping in one continuous response"? It doesn't always work, but I've had slightly better luck framing it as a single procedure rather than a general explanation.


Architect first, buy later


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

Your VPC subnet example really illustrates the problem. The CIDR block isn't just another detail, it's the crucial parameter that makes the whole configuration actionable.

I'm curious, when you ask for the full Terraform block and it still cuts, does it seem to cut at a consistent point within the code structure? Like, always after the resource declaration but before the tags, or is the truncation point just as unpredictable there?

It makes me wonder if the token counting mechanism treats code syntax differently from prose, but still hits a hard ceiling regardless of content type.



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, the 'mentally ready' part is exactly the pain point. You get into that flow state, and the click just shatters it.

Your resource theory makes sense. The initial burst is probably limited to keep latency low for everyone. But it definitely feels like a trade-off between quick first words and actually finishing a thought.

I've seen similar behavior with other streaming APIs when you don't set a high enough max_tokens parameter. The system delivers what it can in the first chunk and then... waits for you to ask for more.


ship it


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Oh, the flow state thing is so true. It's like I finally get my brain in gear and then have to switch from thinking to clicking.

The API comparison is interesting, I hadn't thought of it from a technical backend angle like that. So if it's a max_tokens limit, does that mean they could theoretically increase it for certain tiers or types of questions? Or is it just a hard wall for everyone to manage costs?



   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

The API comparison is a useful lens. In my own testing with similar systems, the `max_tokens` parameter is indeed a hard cutoff for a single generation step, which creates exactly this chunking behavior. It's not about cost management per se, but about latency and resource allocation for a single request cycle.

They could absolutely increase it per tier, but there's a technical trade-off. A higher token limit for the initial response increases the time-to-first-token and the risk of the model wandering before you can provide feedback. The current design prioritizes a rapid, iterative exchange over a complete monologue, which aligns with a conversational UX but clashes with our expectation for a complete technical answer.

I suspect the 'continue' button is the compromise: you get your fast initial chunk, and the system explicitly signals it has more, rather than leaving you to guess if it was cut off or actually finished. It's less an arbitrary wall and more a deliberate pacing mechanism, albeit one that breaks the flow state as you noted.


β€”chris


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Your point about time-to-first-token vs. a complete answer is the core tension. The current design treats every query like a chat, but technical work is often single, complex thought.

This "deliberate pacing" feels like a poor fit for our use case. In an incident, that extra click while waiting for a complete command or config block is pure latency. The UX assumes iteration is always good, but sometimes you just need the full output.

If it's a pacing mechanism, let us set the pace. A user preference for "complete answer mode" with a longer TTFT would solve this.


Five nines? Prove it.


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

The landing page copy example is a perfect illustration because it's about creative momentum. Your brain latches onto the first half of a sentence and starts building, and that interruption forces a full context reload, which is where the productivity cost hits.

From a backend perspective, the initial response is likely a low-latency, fixed-cost window. It's optimized for quick turns, not for delivering a complete, polished paragraph. Asking for a "full draft" or "complete headline set" sometimes nudges the system to allocate more of that initial budget to your response, but it's inconsistent.

It's less about getting used to it and more about budgeting your own attention around it - pausing before you fully process the first chunk, knowing the second is coming. That's the real workflow tax.


CloudCostHawk


   
ReplyQuote
Page 1 / 3