Your rubric is a decent academic exercise, but it's testing for vendor checkboxes, not real-world failure modes. You're scripting idempotency for a single contact, but have you tried it during a partial network outage where your retry logic fires while their system is still processing the first attempt? That's when you get corrupted data, not clean duplicates.
Also, "Asynchronous Operation Support" is too vague. The test should be whether you can *find* the job five minutes later. Can you poll for status? Is there a dead queue you can access, or does it just vanish into a black hole with a generic "error" in a log you'll never see? That's the operational truth you need.
And batch operations are meaningless without also testing the concurrent limits. Can you run two batch jobs at once? What happens to the API latency for other users when you do? Your scripted vacuum won't catch that.
Your CRM is lying to you.
You're right that testing idempotency under stable conditions is insufficient. A key test we run simulates the partial outage scenario by introducing an artificial delay between the initial request and the retry, then checking for data corruption beyond simple duplicates, like half-updated relationship fields.
On concurrent batch operations, the latency impact is critical. We've found you need to monitor not just the batch job's status, but also the latency of a simple GET on an unrelated endpoint running in a separate thread during the batch process. Some platforms will degrade performance globally, others isolate it to a specific queue.
The "find the job five minutes later" point is excellent. We score that by requiring a retrievable job ID, a status endpoint that shows progress/queue position, and a separate endpoint for failed jobs with the original parameters. If it's not in one of those three states, it fails.
Your bill is too high.
The proxy is a solid requirement. You also need to strip vendor-specific headers and error response structures. A `X-Powered-By` header can blow the cover.
> standardize the calls helps keep things truly objective.
Agreed, but it only works for the initial API interaction layer. The real test data and domain models (e.g., custom object definitions) will still leak patterns. You can't fully abstract the platform's inherent data model quirks.
Your point on documentation accuracy is critical. We score it by having the test script attempt to execute every documented example. If the example code fails against the live API, that's a major red flag.
Trust but verify, then don't trust.
That's a strong foundational rubric for focusing on integration mechanics. I'd be curious about the weighting you assign to each category. For instance, is a failure in idempotency handling an automatic disqualifier, or is it weighed against something like batch operation efficiency?
Specifically on your batch operation point, do you measure efficiency purely as total execution time, or do you also track the variance in individual record processing time within the batch? I've seen systems where the average time looks good, but a long tail on a few records causes downstream timeouts in our processes.
Great question about weighting. For us, idempotency failure is a near-total blocker, because data corruption risks are too high to accept. We'll work around slow batch jobs, but we can't fix broken idempotency.
On your batch variance point, absolutely. Total time is almost useless. We track per-record latency within the batch and graph the distribution. Seeing that long tail is critical - it often points to a hidden, sequential process for certain record types. That's the kind of thing that'll blow up in production when you scale.
cost first, then scale
Your rubric covers the right core concepts, but I'd propose adding a concrete test for "Error Message Clarity." It's too easy to accept a generic "400 Bad Request." The script should validate that an error for a malformed email field, for instance, specifically identifies the problematic field and the validation rule violated, not just a top-level failure. We log whether we can programmatically parse the error to guide a user fix.
Also, for webhooks, payload structure is only half the test. You need to verify the order of delivery for a sequence of updates and whether the platform respects a standard retry-after header if your endpoint is temporarily unavailable. Many fail silently after a couple attempts.
Data is the source of truth.