Hi everyone! I've been trying to learn more about RAG (Retrieval-Augmented Generation) by building a small project. Since I'm already using Le Chat for daily questions, I thought, why not use it to help me build a RAG pipeline with it?
The goal was simple: make Le Chat answer questions using my own local documents. I used a local embedding model (all-MiniLM-L6-v2) with Sentence Transformers and FAISS for the vector store, running in a Docker container. Le Chat's API then gets the relevant chunks from my docs before generating an answer.
It worked pretty well for a first attempt! The tricky part was getting the chunk size right and formatting the context for the prompt. Has anyone else tried something like this? I'd love to hear if you used different local embedding models or faced similar challenges. This was a great learning experience for bridging basic Docker stuff with actual AI workflows.
Hey, nice project! I had a similar struggle with chunking. Using all-MiniLM-L6-v2 gave me decent results, but I found switching to a slightly larger model like all-MiniLM-L12-v2 really improved accuracy for technical docs, even if it was a bit slower.
Was your Docker setup just for the embeddings and FAISS? I ended up running everything, including Le Chat's API endpoint, in a single compose file to keep the networking simple. It got messy when I tried to split it up.
Also, prompt formatting! I kept hitting issues where Le Chat would "echo" the context too literally. Adding a clear instruction like "Use the following context to answer the question, but do not repeat it verbatim" in the system prompt made a huge difference.
Ship fast, measure faster.