Skip to content
Notifications
Clear all

Pitfall: Watch out for duplicate evidence if you use the API and UI

14 Posts
13 Users
0 Reactions
12 Views
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
Topic starter   [#28025]

Just a heads up for anyone integrating the Hyperproof API with their workflows. I hit a snag that created some duplicate evidence entries and extra cleanup work.

If you're uploading evidence via the API (say, from an automated monitoring system) and someone on the team manually uploads the same file through the UI for the same control, it can create two separate evidence records. The system doesn't seem to deduplicate based on file hash or name in that scenario. It's not a huge deal, but it can clutter the evidence list and make audits a bit messy. My advice is to pick one method per control type and stick to it!


measure twice, ship once


   
Quote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your experience highlights a critical failure point in evidence collection pipelines. The duplication isn't just a UI clutter issue; it introduces a data consistency problem that can skew metrics. For instance, if your control dashboard counts evidence records to show compliance progress, duplicates would falsely inflate the completion rate.

We observed similar behavior during our integration last quarter. The deduplication logic appears to be isolated to each ingestion path. The API handler and the UI controller likely insert records within separate transactional boundaries without a cross-request consistency check, like a database constraint on a composite key of control_id and file_hash.

A workaround we implemented was to expose a small internal API endpoint that returns the SHA-256 of all evidence linked to a control. Our automation script calls this before upload to check for existing files, acting as a poor man's distributed lock. It adds a few milliseconds of latency but prevents the cleanup tax you mentioned.


--perf


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Oh, good catch! I can see how that would get messy, especially if you're trying to keep a clean audit trail.

It's tricky because the manual upload might come from a team member who just isn't aware of the automated feed. Maybe a simple process note on the control itself could help, so people check before they upload manually?



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

"Simple process note" is about as effective as a "wet floor" sign in a hurricane. You're relying on human diligence to patch a system design flaw, which is a classic failure vector.

The real issue is that this creates plausible deniability for process gaps. An auditor sees duplicate evidence and asks why. "Oh, Bob didn't see the note" isn't a valid corrective action. The system should enforce the single source of truth, not a sticky note.


cg


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

"It's not a huge deal" is where you lose me. Maybe not today, but wait until you're six months into an audit and have to manually validate and reconcile dozens of these duplicates across hundreds of controls. That's not cleanup work, that's rework caused by a system that shouldn't allow it.

Picking one method per control and sticking to it is a fine theory, but it ignores organizational reality. Someone always forgets, someone new joins the team, or a panic manual upload happens during an incident. The system's job is to prevent the mess, not rely on a perfect human process to avoid it.


Anecdotes aren't data.


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Yeah, that's a strong point about auditors. A sticky note process won't hold up there.

But honestly, waiting for the system to be fixed could take forever. I'm wondering if there's a middle ground? Like, could the API call be set up to fail with a clear error if a file with the same hash already exists for that control? That way the system *does* enforce it, at least from one direction.

Probably easier said than done, I know.



   
ReplyQuote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

That's a practical suggestion for a defensive design pattern. You're right that having the API call fail on a duplicate hash would at least prevent automated systems from creating the duplicate, which is half the battle.

The challenge is that you'd need to maintain an index of evidence file hashes at the API layer, which might not exist. If the current schema doesn't have a `file_hash` column or a unique constraint, you'd have to query all existing evidence for a control to compute and compare hashes on every upload, which gets expensive.

A simpler, though less elegant, middleware approach could be to have your automated script perform a GET request first to check for existing evidence file names or metadata before the POST. It's not atomic, but it would catch most cases.


Your data is only as good as your pipeline.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

The middleware suggestion for a pre-check GET is indeed the most pragmatic immediate fix. However, it's important to note this creates a race condition window between the check and the subsequent POST. In high-concurrency environments, or even just with a slow user manually uploading during the script's run, duplicates can still slip through.

A more deterministic, though slightly more complex, pattern is to use a content-addressable storage layer upstream. Your automation could first push the file to a bucket (like S3) using the hash as the key. The upload call to the API would then pass this immutable storage URI as the source. Any subsequent upload, manual or automated, referencing the same URI would be inherently idempotent. This moves the deduplication responsibility to the storage layer, which is typically built for it.



   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

"It's not a huge deal" is where the cost spiral starts. Duplicate evidence means duplicate storage costs in your object store, plus the compute cycles wasted processing and indexing the same file twice. Those pennies add up fast across thousands of controls.


show me the bill


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Exactly, the operational waste is the real concern. Beyond storage and compute, there's the hidden management cost. Each duplicate becomes an item requiring review, potential manual deletion, and explanation in reporting, which consumes engineering time. For a large FinOps program, that's billable hours being spent on data cleanup instead of analysis.

You also have to consider the impact on any downstream systems consuming this evidence data through the API. If they're not built to deduplicate, your analytics on control coverage or automation triggers could be fundamentally flawed.


Your bill is too high.


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

I can definitely see how that could happen. When you say it's not a huge deal, I'm curious if you've noticed this affecting any of the reporting metrics within Hyperproof itself, like evidence coverage or control health scores? I work with similar systems, and sometimes duplicates can inflate counts and give a false sense of compliance completeness, which is its own kind of audit risk.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

You're right that picking a method is the straightforward approach! It works well for small, established teams. I've seen it succeed when the method is documented right in the control's description field as a reminder for everyone.

The caveat is that this breaks down if your automation fails. If the scheduled script hits a timeout and doesn't upload, a team member might panic and manually upload the file through the UI to meet a deadline, thinking they're helping. That's usually how these duplicates sneak in despite the rule. Maybe adding a clear status check for the last automated upload could help?


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Exactly, and that downstream corruption is worse than the waste. You build a dashboard pulling from the API, see "50 pieces of evidence," and think you're golden. Really, it's 25 files uploaded twice. Management makes decisions on bogus numbers.

Been there with Jenkins plugins that double-count build artifacts. The data looks fine until you try to automate a deployment gate based on "evidence count > threshold." Then it either fails for no reason or, worse, deploys based on a false positive. Garbage in, gospel out.


-- old school


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Yeah, that "pick one method and stick to it" advice is the right instinct for avoiding the mess. But I've found it creates a single point of failure for your process. If your team relies solely on the API upload for a critical control and the pipeline breaks, you're stuck between missing a deadline or breaking the rule and creating a duplicate. It's a process trap.


pipeline all the things


   
ReplyQuote