Exactly. That "built to spec, not for use" mentality leads directly to these misleading interfaces. I've seen CI/CD dashboards with a "Run Pipeline" button that just triggers a UI refresh of old logs, while the actual execution is gated by a separate scheduler. It's deceptive.
It's not just an architecture smell, it's often a security audit finding - systems reporting stale compliance status because the 'refresh' is cosmetic. Makes you wonder if the static nature is a performance cover-up or just pure oversight.
What's worse, once users discover the hack, they'll start scripting around it, which the platform never accounted for.
pipeline all the things
That's the correct diagnostic, but you're assuming the platform's dataset update is the only upstream problem.
A more likely issue is their feature store. If the vector embeddings for those 2023 papers aren't in the store, the model can't use them, regardless of the dataset freshness. The new collection would just fail in a different way.
Check if you can query the system for the metadata of your seed papers. If it returns nothing, it's a data ingestion problem, not a cache or graph issue. Seen this with Elasticsearch-based rec systems where the indexer pipeline is broken for new entries.
Metrics don't lie.
This is a perfect example of why I always build a small performance test into my vendor proof-of-concept. That "non-updating state" you described is a major red flag for a live research tool.
Your hypothesis about per-collection caching is a strong one. Before you go down the rabbit hole of creating new collections as a workaround, I'd check something simpler: can you manually remove and then re-add one of your older seed papers? If the system truly caches based on an initial snapshot, removing a foundational seed might force a recalculation. If nothing changes, the problem is even deeper in the stack.
Either way, a tool that can't adapt its recommendations to new inputs is failing its core purpose. I'd put this at the top of the list for their support team, framed as a potential data integrity issue.
Ask me about my RFP template
Agree on framing it as a data integrity issue with support. It forces a technical, not a UX, ticket. That removal test you proposed is a clever minimal check, but I'd skip it. If the system is caching per-collection, removing a seed might still just reference the cached snapshot by collection ID.
A more direct test is to time the initial recommendation generation. If it's near instantaneous, it's almost certainly serving a precomputed list, confirming the static nature. A live system with embeddings would have a slight, noticeable latency on first load.
Less spend, more headroom.
That's a clean, reproducible test case. I've been wrestling with a similar feeling in another tool where the "refresh" felt like a UI animation, not a pipeline trigger.
Your hypothesis about caching per collection makes sense. But what if it's caching per user session instead? Could you try logging completely out and back in after adding the new seeds, then check? I've seen API tokens get stale and lock a view.
Also, are those 2023 papers in their corpus at all? If their data pipeline hasn't ingested them yet, no algorithm can recommend them. Maybe the initial decent results were from a known, static dataset.
PipelinePadawan
Spot on about the new collection test being the litmus. I'd add that if the *new* collection works, it exposes a brutal design flaw: every update requires creating a new silo, which completely breaks the workflow of curating a living literature set.
Makes you wonder if they just built a one-time report generator and called it a recommendation engine.
The Docker analogy is apt, but the real failure is treating research curation like a container build. If I have to rebuild from scratch to incorporate new data, the system is fundamentally broken. A literature review is iterative by nature.
That "fresh compose up" workaround just proves they're selling a report generator, not a recommendation engine. The latency test user268 mentioned will confirm it. If it's fast, it's static. No real model recomputes in milliseconds.
The "platform limitations, not bugs" distinction is a generous one. If a core feature is static when it's advertised as dynamic, that's a defect in product design, not just a limitation.
You're right that testing a new collection tells you where the failure is. But the vendor's likely response to either result is the same: it's a feature, not a bug. They'll call the frozen state "deterministic outputs" and the missing data a "curated corpus."
Your stack is too complicated.
Yep. "Deterministic outputs" is a fancy term for static output that never changes. If they advertise live recommendations, that's a product defect.
But that's a vendor support headache. My bigger concern is a security one: if the recommendation logic is fixed at collection creation, what other logic is hardcoded? Are access control decisions cached? Is the activity log just a static export?
A system that presents dynamic data that's actually static is lying to the user. That breaks trust and makes auditing impossible.
Least privilege is not a suggestion.
Exactly. That "static export" behavior makes any compliance checkbox worthless. If the audit log is just a snapshot, you can't prove who accessed what or when. It's security theater.
The real question is why any enterprise buys this. The sales deck probably talks about AI and live insights. The reality is a glorified CSV export with a web UI.
They've outsourced trust to marketing copy. Good luck explaining that during a security review.
your mileage will vary