Skip to content
Notifications
Clear all

Switching from LlamaIndex to Marvin - initial impressions.

16 Posts
16 Users
0 Reactions
66 Views
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
Topic starter   [#22401]

Alright, so I finally hit my limit with LlamaIndex. I know, I know, it's the darling of the RAG scene right now. But after wrestling with it for a few months on a sales content automation project, the "batteries-included" promise started to feel more like "here's a box of parts, good luck."

My team needed a straightforward pipeline: ingest a ton of old case studies and blog posts, chunk them, embed them, and answer questions for the sales team. Simple, right? LlamaIndex can do it, but the cognitive overhead was insane. The abstractions kept *leaking*. I'd spend more time debugging why a query router wasn't firing correctly or why a specific node parser was eating my formatting than actually building anything.

Here's what pushed me over the edge:
* **The "Kitchen Sink" Problem:** Every new feature feels bolted on. Vector stores, agents, "tools," evaluators... it's becoming a monolith. I just wanted a clean data pipeline, not an entire LLM orchestration framework.
* **Opaque Costs:** Tracking token usage across all their nested abstractions is a nightmare. When a sales ops person asks "why is this costing so much?", you better have a PhD in LlamaIndex internals to explain it.
* **Development Speed ≠ Production Simplicity:** Prototyping was quick. Making it stable, observable, and maintainable? That was a different story. The gap felt huge.

So I switched to [Marvin]( https://www.askmarvin.ai/). It's a much narrower, more declarative framework. The core difference? Marvin feels like it's designed for *software engineers* who want to ship products, not *researchers* who want to explore every possible RAG configuration.

The immediate win was clarity. Instead of `LLM`, `ServiceContext`, `Index`, `QueryEngine`, `RetrieverQueryEngine`... you basically define your data sources, your embedding model, your LLM, and your query logic in a very straightforward, Pythonic way. It's boring. And I mean that as the highest compliment.

Biggest practical difference so far? **Debugging.** When a query returns a weird result, I can actually trace through my Marvin code and see what happened. With LlamaIndex, I'd often end up in the library's own guts, trying to figure out which component was mangling my input.

I'm not saying Marvin is perfect or for everyone. If you need the absolute cutting-edge, hyper-optimized multi-agent stuff, you're still in LlamaIndex/Haystack/LangChain territory. But if you're in B2B sales tech like me, and you just need a reliable, understandable RAG pipeline that answers questions about your product docs or sales enablement materials without a PhD... it's a breath of fresh air.

Anyone else made a similar jump? Or am I just being overly cynical about the complexity?

Just my 2 cents


Trust but verify.


   
Quote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

The cost tracking issue you mentioned is a critical, often overlooked operational hazard. While the framework provides hooks, the actual consumption is distributed across layers you don't directly control: the embedding model calls, the LLM calls for query decomposition, re-ranking, and response synthesis. Each can have its own tokenization pattern. We instrumented it and found that a single "simple query" through a default pipeline triggered 4 separate OpenAI calls, each with its own context window overhead. The bill came from places we didn't expect.

For a sales content pipeline, that opacity is a deal-breaker. You need predictable, per-query cost attribution, not a black box. Simpler architectures, where you manually chain discrete steps, are more painful to build initially but give you direct line-item visibility. You trade abstraction for accountability.

Have you looked at the actual network and logging output to map where the tokens are going, or was the debugging burden itself too high to justify the investigation?



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

Your point about opacity is precisely why we conduct a granular line-item analysis during any vendor evaluation for mission-critical pipelines. That "PhD in LlamaIndex internals" isn't an exaggeration. To forecast operational expenditure, you're forced to reverse-engineer the framework's execution plan, mapping each abstract component to its associated external API call.

This creates a fundamental procurement issue. You're no longer evaluating a simple service with a clear price per transaction. You're taking on a complex, bundled product where the internal allocation of costs is deliberately obscured. This makes negotiating a consumption-based contract with your LLM provider nearly impossible, as you cannot provide the usage breakdowns they require for volume discounts or custom pricing tiers.



   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

That "batteries-included" promise started to feel more like "here's a box of parts, good luck." This resonates so much. I've been through a similar evaluation for a vendor-facing Q&A system.

You mentioned the cognitive overhead and leaking abstractions. A practical consequence we ran into was around versioning. The framework moves fast, and a minor version bump would sometimes silently change the default behavior of a core component, like a retriever's similarity calculation. Our performance would drift, and we'd spend days tracing it back to an implicit default we didn't even know we were relying on. The "batteries" changed voltage without a changelog entry.

Your point about just wanting a clean data pipeline is key. It's the difference between buying a pre-assembled tool and buying a workshop. For a focused use case, you don't want the workshop; you want the tool to work predictably every time, with clear consumable parts.


buyer beware, but buy smart


   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

That's a really practical point I hadn't considered. I work in customer success, so I'm used to trying to explain our product's billing, but this is a whole different layer.

You're saying the procurement issue happens because the cost is bundled inside the framework's abstractions, right? So even if you wanted to negotiate with OpenAI for better rates, you can't give them a clean breakdown of what calls are for embeddings versus queries.

Does this mean frameworks like Marvin, by being simpler, might actually make those vendor negotiations easier? Because you'd have more direct control and visibility into each step?



   
ReplyQuote
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Exactly. That "clean data pipeline" feeling is what I chased, and it's why I landed on Marvin. It's basically a lightweight orchestrator for your prompts, letting you define clear, discrete steps.

Instead of a monolithic RAG framework, you can build each part of your pipeline as a simple function or class with typed inputs and outputs. You still get the AI magic, but you can *see* the data flow. When a sales ops person asks about costs, you can point directly to the single `ai_call` for the final answer synthesis, not some nested abstraction.

It does mean writing a bit more code for things like chunking, but honestly, a 50-line script I understand is better than 10 lines of framework magic that breaks on every version bump.


Prompt engineering is the new debugging


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

Your frustration with the "kitchen sink" problem is exactly why I shifted to a more compositional approach. While a framework promises to handle everything, it often ends up creating a dependency web where you're forced to adopt patterns you don't need just to get one feature working.

The key for a sales content pipeline, as you've found, is clarity and auditability. When you manually define each step, you inherently create the observability hooks that frameworks abstract away. For instance, you can log the exact token count and cost right after your embedding call and again after your final synthesis, because they're explicit functions in your code, not hidden inside a `QueryEngine`'s private methods.

That said, the trade off is you now own the orchestration logic. Things like retries, error handling, and parallelization become your responsibility. In my experience, that's a fair price for the cost transparency and control, especially when dealing with internal stakeholders who demand predictability.


null


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

"that's a fair price for the cost transparency and control" is the whole argument. Frameworks promise to handle orchestration for you, but they just make the complexity opaque. Owning the logic means you can instrument it. You can't have both complete abstraction and real observability.


Beep boop. Show me the data.


   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

That point about instrumentation is exactly right. With LlamaIndex, I tried adding cost logging to our pipeline and hit a wall because the API calls were buried under three layers of decorators. You can't instrument what you can't see.

When you build it yourself with a simpler tool like Marvin, you're forced to write the cost tracking into each step. It becomes part of the design, not an afterthought. You end up with a spreadsheet-friendly log for every query: so many embedding tokens here, so many completion tokens there.

The funny part? That manual logging often catches logic errors the framework would have silently swallowed.



   
ReplyQuote
(@brian7)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Yeah, I think that's a big part of it. If you can show your vendor a log where every call to embed a document is a separate, labeled line item, you're negotiating from a clearer position. It's not just about rates, it's about proving your usage patterns.

But doesn't this just push the complexity onto your team? You now have to design and maintain that clear logging yourself. Maybe that's okay if you have the engineering time.

Does Marvin have built-in helpers for that kind of cost tracking, or is it all manual from the ground up?



   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You're right to ask about built-in helpers. Marvin doesn't have a built-in auditing module, and that's actually a feature from a compliance standpoint. It forces you to design your observability layer explicitly, which aligns with control objectives in standards like SOC 2.

When you manually instrument each step, you're not just logging costs. You're creating an audit trail that maps directly to your data flow, which is crucial for demonstrating data lineage in a security review. A framework's automated logging often aggregates in ways that break this chain of custody.

The initial setup is more work, yes. But you trade that for a system where you can answer an auditor's question about where a specific piece of data was transformed by an AI model. That's not just procurement leverage, it's compliance evidence.


—at


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Yep, that spreadsheet-friendly log is the killer app. It's how I caught a recursive prompt bug that was quietly doubling our GPT-4 bill. The framework's aggregated "total tokens" metric would have just shown a high number, but seeing the line-by-line breakdown showed the same context being sent twice.

The flip side is you have to actually *look* at the spreadsheet. If your team isn't conditioned to check those logs weekly, the transparency is just overhead. I've seen teams build beautiful, granular cost logging and then never open the dashboard.



   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

Your line-item analysis approach hits on a core problem with these abstractions. It's not just procurement - it's capacity planning. When a framework obscures its API call patterns, you can't accurately size your rate limit quotas or predict latency spikes. You end up over-provisioning just to be safe, which itself becomes a hidden cost.

Moving to a simpler orchestrator means you can model the load profile of each discrete step. That lets you negotiate not just price, but also commit tiers and burst allowances with your LLM vendor, because you can show the exact sequence and volume of calls.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That "batteries-included" to "box of parts" feeling is real. I've been there with other frameworks. You start building and suddenly you're not working on your product, you're reverse-engineering someone else's abstraction.

Your cost point hits home for me. I was trying to estimate a budget for a similar project last week. With a complex framework, you can't predict the bill because you don't control the call patterns. How do you even explain a spike to a non-technical stakeholder? You can't point at your code, you have to point at the framework's magic and shrug.

So when you say you just wanted a clean data pipeline, did you find that starting from scratch with something simpler actually forced you to design a more logical flow? Or did you miss any of the pre-built components?



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

The "box of parts" analogy perfectly captures the initial friction. You're right that it forces a more logical design, precisely because you can't rely on a pre-built `QueryEngine`'s hidden flow. In my case, building the pipeline from discrete steps surfaced an unnecessary retrieval call in a summarization task that was adding latency and cost, a flow I had blindly accepted in the previous framework.

I didn't miss the pre-built components, but I did miss the *documentation* that implicitly came with them. When you use a framework's high-level component, its tutorials often provide a de facto system design pattern. With Marvin, you're responsible for that architectural pattern yourself. The trade-off is that your resulting design is documented in your own code and diagrams, not in a third-party tutorial that might drift from the actual implementation.



   
ReplyQuote
Page 1 / 2