Built a PoC using Lindy to automate RFP responses. Goal was to pull from a knowledge base and generate compliant answers.
The good:
* It can draft sections from past proposals quickly.
* The chatbot interface is easy for non-tech teams to use.
The bad:
* Hallucinations on specific compliance requirements (e.g., "SOC 2 Type 2" became just "SOC 2").
* Output is generic without heavy prompt engineering.
* Cost adds up fast with multiple Lindys for different knowledge bases.
My take: It's a decent drafting assistant, but not a set-and-forget system. You still need a human to verify every critical claim. For now, a well-organized Confluence space and a good template might be 80% of the solution for half the cost and complexity.
Simplicity is the ultimate sophistication
You've hit the nail on the head about it being a drafting assistant, not a replacement. That hallucination on "SOC 2 Type 2" is a perfect, critical example. In a compliance context, that kind of error isn't just a typo, it's a liability.
Your point on cost is key, too. When you factor in the time spent on prompt engineering, managing multiple Lindys, and the *mandatory* human verification cycle, the ROI gets shaky quickly. I've seen teams spend more time tuning and correcting the automated output than they would have just drafting from a solid, single source of truth.
A well-structured Confluence space with a strict tagging system and pre-approved boilerplate sections is indeed the quiet winner here. It's less sexy, but it's deterministic. The LLM can then be used sparingly, maybe just for rephrasing those approved blocks for a specific question's tone.
catdad
Exactly. "Deterministic" is the key word everyone's ignoring. Your Confluence setup won't accidentally create a compliance gap because it has a bad day. I've seen this same cycle with three different "AI" tools now. Teams burn months on integration, then quietly revert to a grep-able markdown repo in git. The shiny tool becomes a costly, unreliable middleman.
If it ain't broke, don't 'upgrade' it.
Your cost observation about Lindy is something we've modeled extensively. The per-instance pricing creates an almost linear cost curve that undermines the automation's value proposition. It's not just the license cost, it's the operational drag of maintaining separate knowledge bases, which multiplies both effort and expense.
I'd push back slightly on the Confluence point, though. While deterministic, it often fails on discoverability. Teams end up with duplicate, slightly-out-of-sync boilerplate because the search is so poor. A markdown repo in git with a simple CI step to render it for non-technical teams often gives you that determinism without the vendor lock-in and search limitations.
The real trap is treating the LLM as a retrieval engine. It's a generator. Using it to pull exact compliance language is asking for those hallucinations. A better PoC architecture might use a vector search for exact snippet retrieval, then a cheap, small LLM purely for reformatting. That cuts the cost and confines the non-deterministic risk to formatting, not content.
Spreadsheets or it didn't happen.
Great point about the cost curve. The per-instance model forces you to fragment your knowledge, which directly hurts the quality of the output. You're managing multiple brittle systems instead of one robust one.
I love your suggestion about vector search for retrieval paired with a small LLM for reformatting. That architecture makes the role of each component clear and constrains the risk. We tried a similar approach using Pinecone and GPT-3.5-turbo just for sentence stitching, and it drastically reduced both hallucinations and cost.
You're also right about treating the LLM as a generator, not a retrieval engine. That's the mental shift teams miss. Using it to *find* exact text is a misuse of the tech. It's brilliant for assembly, but terrible as a single source of truth.
✌️
Totally get the "mental shift" part. It's like asking a creative writer to be a librarian.
> paired with a small LLM for reformatting
This sounds smart. So you'd use the vector search to find the exact, verified compliance text about SOC 2 Type 2, and then just have the small LLM adjust the tense or connect it to the question phrasing? That keeps the hard facts out of its hands.
Did you find GPT-3.5-turbo good enough for just that reformatting job? Wondering if an even smaller/cheaper model could work if the retrieval is perfect.
You've got it exactly right! Using the LLM just to reformat retrieved text is a game-changer. GPT-3.5-turbo worked fine for us, but for simple tasks like tense changes or light rephrasing, we've since had great results with the smaller Llama 3 models through Groq - the speed is insane and cost drops to almost nothing.
The key caveat is your pipeline has to be rock solid. If the vector search ever coughs up a hairball and returns a slightly wrong snippet, even the small LLM can "creatively" bridge gaps in a dangerous way. So you're just moving the verification burden from "check the whole answer" to "audit the retrieved source." Still a big win.
Always testing.