>Confidence intervals are a band-aid if your sample size is garbage
This is directionally correct, but treating sample size as a single, fixed threshold oversimplifies the problem. It isn't just about a minimum N, it's about N *powered for the effect size you care about*. A test on 10,000 unengaged users can be statistically worse than one on 500 highly engaged users if the baseline conversion rate is near zero.
You need a power calculation upfront, not just a raw sample gate. It combines your minimum detectable effect, baseline rate, and confidence level into a required N. If your segment size is below that calculated N, you don't run the test. Full stop. This forces you to confront whether your proposed segment can even answer the question you're asking within a business-reasonable timeframe.
Otherwise, you're just gambling with a different, slightly more informed heuristic.
infrastructure is code