I've been prototyping a lead scoring system for our sales team using the Gemini API. The goal was to move beyond basic keyword matching to something that actually reflects how our experienced reps evaluate a prospect. Our old rule-based system had a 0.3 correlation with manual rep scoring – basically useless.
I used Gemini 1.5 Flash for its speed and lower cost, perfect for iterating on prompt logic. The key was structuring the prompt to output a strict JSON schema, not just a score. We feed it enriched CRM data (website tech stack, funding news, job postings) and ask for:
* A numerical score from 1-10
* The primary reason for that score (e.g., "Strong signal: they're hiring for 3+ roles in our domain")
* A confidence level (High/Medium/Low) based on data completeness
* Three qualifying questions a rep should ask
The initial latency for this structured output was ~1200ms, which was fine for batch processing overnight. But when we tried a real-time variant for our Chrome extension, that was a non-starter. We switched to a two-phase approach: a fast, simple keyword check using a local model to filter out clear mismatches, and only the passing leads get the full Gemini analysis.
Some early pitfalls we found:
* The model was initially too optimistic, scoring every funded startup highly. We had to add negative scoring cues explicitly into the prompt.
* Without the confidence flag, reps would waste time on high-scoring leads with poor data.
* The "qualifying questions" output has been the biggest win – it gives reps an immediate conversation starter.
Has anyone else tried building a scoring or classification layer with Gemini that needs to align with a specific internal process? I'm especially curious about prompt patterns for balanced scoring and if you've measured the latency vs. accuracy trade-off in a live environment.
ms matters