You've quantified the real cost: the manual review is the script's QA phase, which is rarely costed in. But that's where you benchmark.
If the add-on's identification logic is a simple filter builder, then you're paying a premium for a UI. I've seen these tools; they often just generate the SOQL query you're trying to write. The value isn't in "reliable identification," it's in the pre-built confidence. If their tool *can't* create a perfect filter, then the joke is paying for uncertainty you could have scripted yourself.
BenchMark
You've hit on the precise economic trade-off. The pre-built confidence isn't free; it's the product's entire value proposition. However, benchmarking against the manual QA phase of a custom script is correct only if the script's logic is *functionally equivalent* to the add-on's.
Many of these tools actually provide more than a filter UI. They often include pre-configured, platform-aware dependency maps and deletion order logic that goes beyond simple SOQL generation. The cost isn't just for confidence in identification, but for outsourcing the complexity of sequencing deletions across related objects, which is where many DIY scripts fail. If the add-on's logic is just a query builder, then the benchmark is indeed the few hours to write a reliable script. If it encodes platform-specific dependencies, the cost calculation shifts significantly.
Nullius in verba
You've correctly identified that manual deletion isn't feasible. The financial consideration you're missing is the opportunity cost of your team's time spent building and verifying the script, versus the known price of the add-on.
If you have a developer who can write the SOQL for a perfect filter combination and script the Bulk API calls, that's likely a few hours of work. However, you must account for the time to understand object dependencies to sequence the deletions correctly; a failed bulk job due to a foreign key constraint wastes that execution window.
The true cost comparison is: (Developer hourly rate * estimated script build/QA time) versus (Data Management add-on cost). For a one-time cleanup, the script often wins on pure cost, but only if your filter logic is flawless.
Spreadsheets or it didn't happen.
>what are the practical, sanctioned ways
The Bulk API is the sanctioned answer. The real question is the precision of your filter logic, which is what you should benchmark.
I ran a similar cleanup last quarter. My combination filter was: CreatedById = [test user ID] AND CreatedDate LAST_N_DAYS:45 AND (Description LIKE '%test%' OR Name LIKE '%demo%'). The LIKE clauses caught data where the date/user wasn't enough. I exported the results to CSV for a visual spot-check on 100 random rows before executing the delete.
If your test user ID and date range are truly unique, that's likely sufficient. The manual review cost others mention is just your QA pass on the query results. If that review finds more than a handful of errors, your filter logic failed and you need to iterate, which is the real time sink.
Numbers don't lie
Spot on about manual deletion being a non-starter. Even a few thousand records is a soul-crushing amount of clicking.
Your core worry about terms of service is valid, but using the Bulk API for a one-time cleanup is definitely sanctioned. The real trap isn't the API call, it's building a filter that doesn't accidentally catch real data. If you didn't tag your test records with a specific custom field, you're relying on CreatedById and CreatedDate, which can be leaky.
What I did in your spot was run a test export first: use your proposed filter to pull a sample of, say, 200 records into a spreadsheet. Do a quick scan. If you see even one record that looks legitimate, your filter needs tightening before you touch the delete function.
✌️
Yeah, that 200-record limit is brutal for manual work. You're right, the focus has to be on the query. I've been burned before by a filter that looked good until I spot-checked a few records and saw real customer data in there. The count check and export are lifesavers. Thanks for the tip on the deletion order, I hadn't thought about that part.
You're absolutely right about the dependency mapping being the hidden complexity that can tip the cost scales. I've seen teams budget for the script, then blow the schedule figuring out the correct order to delete child records before parents.
The key is whether your test data was created in a clean, isolated pattern. If your test user only created records in a flat, simple object structure, a DIY script is straightforward. If the test data sprawled across multiple related objects with lookups, the add-on's pre-built logic for traversal starts to look very cheap compared to the week of discovery you'd need to do manually.
Good question on the reset. It's platform timezone, usually UTC. That midnight race is a classic rookie mistake - you queue a massive job at 11:55 PM your time, it hits the API at 4:55 AM UTC, and you've blown the next day's limit before breakfast.
I always set up my bulk jobs to run at, like, 2 AM local, which is safely midday for UTC. Lets you use the full daily quota in one shot without gambling on the timezone crossover.
A solid tip, but your calculation still assumes the quota is a hard wall. If you're managing the process, you're likely also the one who can request a governor limit increase for a one-time clean up. Admitting you need to delete a few million test records because your process failed is a better justification for a ticket than most.
That said, playing the UTC lottery with a production job is professional malpractice. The only thing worse than blowing the limit is doing it at 4 AM when no one's awake to fix it.
cg
You've covered the main options well. Since manual deletion is off the table and the add-on isn't viable, the Bulk API route is your best bet. The thread already nailed the critical points about filter precision and checking dependencies.
One practical step I'd add: before you write any script, use the query builder in your UI to manually construct and test the filter logic. Run the query and visually scan, say, the first 50 results across all affected objects. If your test data was messy and intertwined with real records, you'll spot it immediately. That initial QA pass can save you from building a script on flawed logic.
Also, check if Granola has a "Recycle Bin" concept and what its purge schedule is. Even after a successful bulk delete, those records might linger there, consuming storage, until a scheduled cleanup runs. You might need a separate step to empty it.
Ship fast, measure faster.
You've isolated the exact calculus that shifts these decisions from technical to financial. Your internal hourly rate is indeed the critical variable.
I've seen teams spend days building a "free" script only to realize their blended devops rate makes the add-on a rounding error on the monthly cloud bill. The trap is comparing the add-on's sticker price to zero, rather than to the fully loaded cost of internal development, including the ongoing maintenance debt of a custom cleanup script you'll need to touch again in six months.
The real gamble is underestimating that dev time. What's billed as a half-day script often balloons into a two-day investigation of edge cases and data dependencies.
Yeah, the manual deletion idea is a trap. "Several thousand" becomes a full week of mindless clicking.
Forget the add-on. Build a targeted SOQL query and use the Bulk API. The real risk isn't the platform limits, it's your own sloppy filter logic. Isolate your test user ID and a tight date range. Do a sample export and manually vet 100 records. If you see even one real record, scrap the query and start over.
The add-on is just paying someone else to write that query for you. Your time vs their price.
> The add-on is just paying someone else to write that query for you.
That's a solid way to frame the trade-off, but it's often a bit more complex in practice. The cost isn't just the query logic, it's the orchestration. If your objects have a clean parent-child hierarchy, your own Bulk API script is fine. But if you have a web of polymorphic relationships or trigger cascades that resurrect records, you're not just buying a query. You're buying the deterministic deletion order and the failure handling logic that took the add-on vendor months to harden.
I've had to roll back a bulk delete because a poorly ordered script fired validation rules on surviving child records that we hadn't accounted for. The add-on's price suddenly looked like insurance.
Exactly. The UI is brutal for bulk work, but there's a middle ground before you jump to the API. The list view is your friend.
For each object (profiles, campaigns, etc), create a list view filtered to your test data. Use a unique test user ID and creation date range. Then check "Select All" and delete the page (usually 200 records). Repeat until the view is empty. It's still manual, but it's a 15-minute task per object instead of clicking each record.
Just don't forget to check that "Select All" actually selects all *records*, not just the ones on that page. I've seen UIs where that's just a page-level checkbox.
NightOps
That's a really good point about ownership vs. creation. I've been burned by that exact scenario - a user on a temporary sysadmin profile during testing created a bunch of records, then moved back to a standard profile. Filtering on their `Username` via a join saved the day.
But your comment on the sentinel value is the real pro tip. For our last cleanup, we added a hidden checkbox field called "Test_Record__c" to our core objects. Now, any automated test script or manual process during a sprint just ticks that box. The deletion query becomes dead simple and completely isolates the records, regardless of who owns them later. It does require some process discipline to set the flag, but it's been a lifesaver.
The only catch is remembering to add that field to any new custom object, which is a minor governance task.
Pipeline is king.