For a small team, the cron + Python approach works, but the reliability gap versus native connectors is real. I'd suggest a middle path.
Since you're on Snowflake, consider using Snowpipe Streaming to push raw rows to a message queue, then a small consumer service calls the OpenPipe API. This shifts the scheduling burden to Snowflake's managed service and gives you better real-time characteristics without building a full orchestration layer. It's more setup than cron, but less than managing batch state yourself.
The performance issue with large datasets isn't just API rate limits, it's prompt iteration latency. If you're classifying tickets, you'll likely tweak your prompt. Regenerating classifications for your entire historical dataset via backfill becomes a significant time and cost sink if your formatting logic is embedded in complex warehouse SQL. Keep that transformation logic external and idempotent.
benchmark or bust
Your suggestion of using Snowpipe Streaming to a message queue is architecturally sound, but the cost dynamics shift significantly. You're trading cron compute costs for Snowpipe Streaming's per-row ingestion fees and the ongoing cost of a consumer service (likely an always-on container or serverless function). For a continuous stream of classification tasks, this can eclipse the cost of a few daily batch runs, especially if your data volume is moderate but constant.
The real benefit, as you note, is real-time characteristics and shifting the scheduling burden. That's a valid trade, but teams should model the cost of streaming 100% of their rows versus a batch approach that might process the same volume in a few scheduled windows. The break-even point depends entirely on how time-sensitive the classifications are.
Your point on idempotent, external transformation logic is critical for cost control during prompt iteration. If the logic is external, you can backfill from a historical snapshot in your object storage at a fraction of the cost of rescanning your production warehouse tables. That decouples experimentation cost from your live data infrastructure.
Always check the data transfer costs.