Format-preserving crypto's the easy part. The hard part is getting vendors to show you the actual key calls, not their marketing architecture diagram.
> ask both vendors for a detailed architecture diagram
Those are polished to perfection. You want the *failure mode* diagram instead. What happens on a network partition between proxy and vault? Does it fail open or closed? That tells you more about the guts than any happy-path flow.
For key rotation, don't just ask if they support it. Ask them how *old* keys are purged from memory in the proxies. I've seen systems where rotated keys lingered in app-layer caches for days.
-- old school
Performance defaults always have a cost angle. That caching TTL means you're paying for KMS calls you don't need. I'd bet their "performance" tier pricing assumes you leave it on.
Asking for a CLI command to verify KMS hits is smart. Did Glide's command show you the cost per operation? Because ten rapid calls hitting KMS is ten line items on your bill.
show me the bill
Totally hear you on the checkbox feature fatigue. I've seen that "lookup table" approach, and worse, some try to pass off reversible encryption *without* a separate KMS as "tokenization" 😬
For your zero-trust PAM project, you're right to focus on the crypto guts. I'd push both vendors on their SDK or proxy's default behavior right out of the box. The default config often reveals their real priority: security or performance. Ask to see the exact line in their configuration (or better, a config file) that disables any local caching or in-memory key storage. If they hesitate, that's your answer.
Also, request their key rotation runbook, not just the feature list. How is a rotated key truly, verifiably, invalidated across all running instances? A "soft delete" in the KMS isn't enough if the proxy holds onto it.
Clean code is not an option, it's a sanity measure.
Yep, that persistent cache is a classic example of a secure design being undermined by a performance shortcut. It effectively brings the plaintext back into the application layer, which defeats the whole purpose of externalizing keys.
Your point about audit trail complexity is spot on. I've seen teams get so tangled in key version metadata that they can't easily answer "who accessed this data last month?" without a dedicated query. The policy sprawl is real, especially if you're automating rotations frequently.
Raise the signal, lower the noise.
The migration cost spike you observed is a critical FinOps consideration often omitted from the comparison. It exposes that both vendors treat token generation as a revenue event, which shifts the cost optimization burden entirely to the customer's capacity planning.
> you might be looking at the wrong layer altogether
This is the key insight. The architectural debate between these two services often ignores that tokenization as a service inherently adds latency and cost. For user-facing applications, we've had better results implementing format-preserving encryption at the database level with column-level encryption, using a local HSM for key operations. This keeps the cryptographic boundary tight and eliminates the network hop entirely. The trade-off is managing key rotation yourself, but the latency drops to single-digit microseconds.
Your stress test data on network hops versus FPE processing is telling. It confirms that the overhead isn't in the cryptography, but in the service boundary. If a team's requirement is truly zero-trust with a separate security domain, then that overhead is justified. But if the goal is simply PCI compliance or data masking, a simpler, closer-to-the-data model often wins.
No free lunch in cloud.
Exactly right about verifying logs yourself. I'd add that even when you do see those KMS logs, you need to check the *content* of the audit trail.
We once found logs showing calls from the expected service account, but the "reason" field was generic. We had to push for a custom audit log configuration to include the specific token ID or request context to make it useful for actual incident response. Without that, you can't tie a specific suspicious access to a detokenization event.
The "plug any backend" model means that audit log detail becomes even more critical, because you're responsible for validating that the backend you plug in provides the granularity you need.
Ask me about my RFP template
The audit log content issue is the difference between having a surveillance system and having one that records in a usable format. We built a similar requirement into our vendor evaluation matrix, demanding that every cryptographic operation log must include at least three contextual fields: the requesting principal's full IAM path, the exact data identifier (like a primary key or token), and the business reason code from the application.
A generic "Detokenize" event in CloudTrail or a vault audit log is useless for forensics. You need to be able to reconstruct the full data access narrative, especially when investigating lateral movement after an initial breach.
This becomes a systems integration problem. The vendor's proxy might log the token, but your application must pass the business context. We implemented a mandatory HTTP header for the business reason that the proxy validates and appends to its audit stream. If the header is missing, the request is still processed, but it triggers a high-severity security alert for an unexplained access attempt. This shifts the compliance burden back to the development teams in a visible way.
—BJ
That mandatory header approach is smart. It forces developers to think about the reason at the point of access.
But how do you handle that for automated, backend batch jobs? Do you create a generic "nightly-report" reason code, or does each report need its own? If it's too generic, you lose the forensic detail. If it's too specific, you create a management headache.
Look past the crypto spec sheets. Ask for their failure runbook.
Show me a cold start scenario where the vault cluster is down and a PAM session needs to connect. Does the proxy block, or does it fall back to a cached key? That's your real architecture.
You're dead right about the latency being a dealbreaker for user-facing apps. We saw the same thing when we tested Token for a checkout flow. That extra 100ms ballooned our cart abandonment.
>Have you looked at where that latency is actually spent
We did. For us, it was mostly round trips. Each operation, even with a presumably efficient SDK, required multiple handshakes with their control plane for policy checks before hitting the actual KMS. Their architecture is "clean" because the keys never leave their boundary, but that purity comes at the cost of extra network calls. Glide's model, with its local proxy caching the mapping, trades a bit of that purity for a massive latency win.
Your point about Glide's pricing spiking on migrations is a great catch. It turns a "pay for what you use" model into a punishing tax for the one-time event of moving your data. It forces you to do the capacity planning they should be providing.