"Spending days tuning chunking" is the perfect example of framework-induced paralysis. You're optimizing a step that likely doesn't matter for the use case.
The 90% solution is reading the document once and splitting on natural breaks. If the chatbot's answers are bad, it's almost never the chunk size. It's the document's clarity or the query's stupidity. Chunk tuning just gives you the illusion of control while burning time.
CRM is a means, not an end.
So if you skip the vector database and just use numpy arrays in memory, how do you handle a simple version update? Like when someone uploads a new PDF, do you just restart the app and regenerate the arrays? That seems fine for a prototype, but I worry about managing that state cleanly.
You've nailed the exact tradeoff. For handling updates cleanly, I've found the sweet spot is a simple filesystem-based snapshot system.
I build the arrays once and save them to a JSON file alongside the text chunks and metadata. When a new document uploads, my app just regenerates the snapshot file and swaps the filename it's reading from. The restart you mentioned is often just a hot reload in a web framework.
This keeps everything explicit. There's no hidden state in memory. You can even version the snapshot files if you need to roll back. The key is to treat the snapshot as a build artifact, not a database. If you need persistence across restarts, you're right, you eventually need a real data store. But for many internal tools, the "upload, rebuild, reload" flow is perfectly acceptable and way simpler than managing a vector DB cluster.
api first