Their marketing claims "near-local" FS performance for batch jobs. I've yet to see a managed service deliver that without asterisks the size of a phone book.
What's the real latency on the first read in a job? Does it fall apart with 10k+ small files? And what's the egress bill look like when you inevitably need to pull data back out? Considering it as an alternative to a self-managed Spark cluster, but the pricing page is suspiciously light on details.
Your stack is too complicated.
> near-local FS performance for batch jobs
Tested it last month. Latency on first read is fine once the cache warms, but that first cold start can add 30+ seconds if your data isn't pinned. It's the 10k small files scenario where it really bogs down.
The metadata ops kill you. It's fine for bulk reads on a few large files, but directory listing performance isn't linear.
Egress is brutal, like all of them. You end up staging results back to your own S3 anyway.
YAML all the things.