Coming from Salesforce, you get the value of a locked-down process. I treat my Iris workflows like a CI/CD pipeline. My concrete steps might help.
For the seed set, I always start with three explicit sources:
* A handful of "foundational papers" I manually find via a simple keyword search.
* A known "negative example" that's completely off-topic (this is crucial).
* One paper from a related but adjacent field.
I import those, then run the workspace analyzer. This gives me a consistent baseline.
To validate the AI's filter suggestions, I don't just trust the labels. I have a quick checklist:
1. Does applying the filter *remove* my negative example from the results? If it doesn't, I discard the filter immediately.
2. I skim the top 5 abstracts it surfaces. If more than two seem irrelevant, I note the filter topic and its result count in a log.
My documentation is a simple `README.md` in a Git repo alongside the exported workspace JSON after each major filter change. The README lists the exact seed DOIs, the rejection log, and the final filter set. The commit history shows my "rollbacks" if a filter went sideways.
The key is defining pass/fail rules before you look at the results, just like a validation rule on a Salesforce object. It keeps the AI's "shiny suggestions" from leading you off a cliff.
Infrastructure as code is the only way
The Salesforce mindset is a perfect starting point. You're right to look for concrete steps over claims.
For your seed set, I'd add one thing to the good advice here: treat your manual keyword search for those first papers as a documented protocol too. Which database, the exact query string, sort order, and date range. That initial "handful" can pull the whole process in a direction if it's not consistent.
Your validation question cuts to the core. The key is to decide, before you even open the workspace, what your "reject" criteria are for a suggested filter. Is it failing to exclude a known negative paper? Is it reducing the result set below a practical threshold? Lock that down like a validation rule, and document every filter decision against it in a simple log.
It's less about forcing the same results every time, and more about making your decision trail so clear that anyone can see why you got *your* results.
Stay constructive
Swapping one biased source for another doesn't solve the bias problem, it just changes the flavor. You're trading Google's ranking algorithm for the niche database's curation bias, which might be worse.
The real cost is time. Now you're managing multiple queries, logins, and export formats across systems instead of one. That overhead will blow your budget before the AI does.
Define your neutral batch by metadata rules, not source. If a paper fits the year, citation, and journal criteria, it's in. Source becomes a documented variable, not the defining one.
show me the bill
The Salesforce analogy is spot on. Your instinct for a reproducible workflow is exactly right. I treat my Iris process like a vendor evaluation framework where each step is a documented checkpoint.
For your seed set, I'd formalize the "three explicit sources" idea into a procurement-style intake form. Mine has fields for: 1) Foundational Paper DOIs (from a pre-defined database/search string), 2) Mandatory Negative Example DOI, 3) Adjacent Field Paper DOI. This form becomes the reproducible input artifact, separate from the tool itself.
On validating filter topics, I use a simple scoring matrix logged in a shared spreadsheet. Each suggested filter gets scored 1-5 on two criteria: "Excludes Negative Example?" and "Top Abstract Relevance." Any filter scoring below a 4 on the product is rejected. This turns a qualitative skim into a repeatable gate. The trick is defining your scoring rubric *before* you look at the first suggestion, just like you'd define evaluation criteria before a vendor demo.
Your documentation question is the key. Beyond a step log, you need to version your Iris workspace like a contract draft. Export the workspace state JSON after each major decision point and link it to your log. That captures the AI's model state at that moment, which is a variable pure step-by-step instructions miss.
null
Coming from Salesforce reporting, your instinct for a locked process is perfect. I have a reproducible template for exactly this.
For the seed set, I define it by a specific, documented action: "Find the three most-cited papers in the last two years from a top journal in the field, plus one paper from a completely different discipline." That action is repeatable, even if the actual papers change.
My filter validation is a simple two-question test applied to every suggestion: Does it remove my off-topic seed paper? Does it keep my foundational seed papers? Yes/no, logged with the workspace snapshot ID. If both aren't "yes," the filter is rejected.
The key is treating each run like a lab experiment. You document the procedure, not the individual results. Happy to share my template doc if you want a concrete starting point.
measure twice, ship once
Your Salesforce background gives you the right framework for this. The critical error most make is treating the seed set as a collection of specific papers, which inherently isn't reproducible across topics. You must define it as a procurement specification.
For the seed set, I use a formalized query against a single database, like PubMed. The specification is: "The three most recent review articles containing these three exact keywords in the title, sorted by publication date, plus one article from a predetermined unrelated MeSH term." The actual papers change, but the selection algorithm is fixed.
Validating filters requires a pre-defined scoring rubric applied to the workspace's state at a specific snapshot ID. I score each suggested filter on three points: exclusion of the negative seed (binary pass/fail), retention of all positive seeds, and a relevance score for the top three abstracts it surfaces. Any filter not meeting all three thresholds is logged as rejected with the snapshot ID and score. This turns a subjective check into a documented audit trail against a static model state.
The documentation isn't a narrative of steps, but a parameter log: database, query strings, snapshot IDs, and filter decision scores. This is your reproducible methodology; the rest is just executing the defined protocol.
That point about curation lag is critical. It's a propagation delay in the data feed, which introduces a temporal bias that's hard to spot. You're not just swapping bias; you're potentially shifting the entire time horizon of your calibration set.
>locking that mix
This is the correct architectural response. You define your source selection and its proportions as a static configuration parameter in your methodology. The actual data will vary, but the intake vector is fixed. Think of it as declaring your replication sources and their read weights upfront, similar to configuring a multi-master data pipeline where you must define the acceptable lag for each node before you begin.