Alright, let's cut through the marketing fluff. Every other week there's a new "AI-powered" transcription service that promises the moon, then falls over the second you feed it a meeting with three people talking over a poor Zoom connection and a fan in the background. I need something that works for a real engineering org.
We're a 200-person shop, everything's on AWS (mostly EC2, some ECS, RDS for Postgres). We have a mix of internal standups, client calls, and architecture reviews. The requirements aren't exotic, but they're non-negotiable:
* Must handle 50-100 meetings daily, some concurrent. Peak load matters.
* Must ingest directly from Zoom/Teams/Google Meet APIs. I'm not having people manually upload files.
* Output needs to be usable – a simple SRT/VTT for starters, but also a structured JSON with speaker diarization and timestamps so we can pipe it into our own analytics and search later.
* It has to live in our AWS environment, or at least in a VPC we control. I'm not shipping sensitive client dialogue to some random third-party's ungoverned blob storage. GDPR and all that.
* Budget is a concern, but reliability is paramount. I'd rather pay for something that works than get a "free" tier that fails silently.
I've evaluated the usual suspects. AssemblyAI's API looks clean but gets pricey at scale. Deepgram's accuracy is decent, but the pipeline setup felt like more work than advertised. Tried building our own with Whisper on GPU instances, but the operational overhead for diarization and scaling wasn't worth the engineer hours—turned into a full-time job babysitting inference queues.
What I'm looking for is a review from someone who's run something like **Sembly** (or a comparable tool) in a similar environment. Not a proof-of-concept, but in production, with real load.
I need the gritty details:
* How did you handle ingestion? Are there webhooks that write to an SQS queue, or is it a periodic poll?
* What does the actual deployment look like? A CloudFormation stack? Helm chart for EKS?
* Where does the processed audio land? S3? What's the permission model?
* How's the error handling when a meeting provider API has a hiccup?
* Most importantly, what's the *actual* accuracy like on a technical discussion with jargon, and what's the latency from meeting end to transcript availability?
If your answer is "just use the hosted SaaS," save your keystrokes. I need the configs, the gotchas, and the monthly bill for a couple hundred users. Show me your Terraform or your CloudWatch dashboards, and I'll show you my gratitude.
-- old salt