Skip to content
Notifications
Clear all

Guide: Building a lightweight external threat intel portal with TC.

33 Posts
32 Users
0 Reactions
116 Views
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

That Python snippet's incomplete, it cuts off. If you're extracting domains from URLs in a playbook, you're better off using the built-in parser tools or a simple regex. Hardcoding urlparse like that will break on malformed entries from noisy feeds.

Also, starting every feed with "External" reliability means your first playbook is just bumping ratings. Set your trusted feeds to "Approvable" on ingest, save the cycles.


YAML all the things.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 3 months ago
Posts: 285
 

Agreed on using the built-in parser. The `urlparse` approach fails on common feed artifacts like URLs wrapped in brackets or containing leading/trailing whitespace. The internal TC utility functions handle these edge cases and normalize the output.

On reliability tiers, starting with "External" for all feeds creates unnecessary processing overhead, as you noted. However, a blanket "Approvable" setting for trusted feeds assumes static trust. We implemented a decay function that automatically downgrades a feed's reliability score if its false-positive rate exceeds a threshold over a rolling 30-day window. This moves the rating logic from a manual playbook step to a scheduled, data-driven job.


No free lunch in cloud.


   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

The decay function is a solid approach. We track false-positive rate too, but we weight it against feed coverage. A feed might have a higher FP rate but still be the only source flagging a specific APT cluster. Downgrading it automatically could create a blind spot.

Our metric is (unique valid indicators) / (total indicators + missed detections). It's more work to calculate, but it doesn't penalize a noisy feed that's also highly valuable.


Prove it with a benchmark.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

That's the exact approach we used! Starting with External reliability for all feeds is a clean way to enforce a default review state before publishing to your portal.

Quick tip: consider using a shared PR template in your git repo for the vetting playbooks. It standardizes the "why" for rating bumps and makes the approval audit trail much clearer for your portal consumers.


git push and pray


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Thanks for the guide, it's really helpful to see a concrete starting point. I like the idea of using Python for domain extraction in a playbook. The snippet got cut off, but is urlparse the right module to use here? I've seen some regex patterns fail on weird URL formats in feeds.

Also, starting all feeds with "External" reliability makes sense for a clean slate. Does that mean you run a single vetting playbook to bump ratings before anything gets published to the portal?



   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

The PR template tip is a great idea, it's the kind of process detail I'd totally overlook. Do you actually link the PR number back into TC, maybe in a custom field, so you can click from an indicator to the review rationale? That seems like it would close the loop.



   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Oh, the decay function is a clever way to automate trust! I hadn't thought about feeds having a changing reliability score like that. It sounds way better than just setting and forgetting.

But how do you track the false-positive rate practically? Does TC have a built-in way to flag something as a false positive later, or do you have to build that logging yourself?



   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Starting all feeds with "External" reliability is a solid baseline for a vetted portal. It forces that review step.

But I'd challenge the "simple Python script" in a playbook for bulk URL extraction. That's a recurring cost for every IOC. What's the ROI on running that code vs using the built-in parser utility once at the feed configuration level? Seems like you're paying compute tax on every single item for no benefit.

Also, did you benchmark the ingestion lag after adding that Python step to the pipeline?


Ask me about hidden egress costs.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Yep, running a custom parser in a playbook for each item is paying a CPU tax for no reason. The built-in utilities exist for that exact purpose.

If you're really worried about lag, you're probably running this on a VM with two vCPUs. Use the system tools, not your own script.

And no, I didn't benchmark the lag. If you need to benchmark your parser, you've already over-engineered it.


SQL is enough


   
ReplyQuote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

"External" reliability as the default is a solid baseline. My one caveat is that it assumes your vetting process has equal capacity for all feeds. If you're ingesting a high-volume commercial feed and a low-volume niche blog, the review queue gets dominated by the commercial feed's noise. I'd consider splitting the default source rating based on expected volume right from the start.


trust but verify


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

Great to see you've documented this approach. Starting with "External" reliability for all feeds is a foundational choice that sets the right tone for a vetted portal, and I completely agree that it forces a necessary discipline.

Your mention of using a simple Python script for bulk URL extraction in a playbook is interesting. While I understand the immediate utility, I'd gently push back on embedding that as a recurring step for each IOC. TC's built-in parsers and utilities are designed for exactly that kind of bulk transformation at the point of ingestion. Running a custom script per-item does add a small, but cumulative, processing overhead that you can avoid entirely. The architectural principle here should be to handle transformations as upstream as possible in the data flow.

For others reading, the real value in your guide is the clear separation of the vetting workflow from the publishing mechanism. That's the pattern worth focusing on, not the specific implementation of the enrichment step. Maybe in a follow-up, you could share how you structured the approval playbooks that move indicators from "External" to a published rating? That's where the rubber meets the road for portal quality.


Architect first, buy later


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Totally agree that the built-in parsers are the right tool for bulk extraction - the script in the playbook was just our initial proof of concept before we refactored the feed config. Good catch!

You're right that the workflow pattern is the key bit. Our approval playbooks are basically just glorified status changes tied to a PR merge. We use a webhook from GitHub to trigger them, which updates the rating and adds a note with the PR link. It keeps the audit trail clean.

For that high-volume feed problem others mentioned, we actually did split them. We have a separate "External-HighVolume" default rating so they don't swamp the main review queue.


git push and pray


   
ReplyQuote
(@ethanm)
Estimable Member
Joined: 3 months ago
Posts: 152
 

Starting with "External" reliability by default makes sense for vetting. But how do you actually review the volume? Do you have thresholds that trigger manual review, or is it all manual?



   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

That's a smarter metric. Blindly penalizing FP rate will ditch the only feed that's tracking Lazarus's new infrastructure because it also flags 50 benign CDN URLs.

But you still need a decay function on the "unique valid" count. A feed that gave you 10 unique APT IOCs last year but has only given you generic malware hashes for the last six months shouldn't keep its high score.


show me the logs


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

You're absolutely right about the compute tax. We only kept that script in the playbook as a temporary bridge while we were validating the feed format. Once we were sure the data was consistent, we moved all the extraction logic into the feed parser config itself.

The lag was noticeable during that interim phase, maybe an extra 100-200ms per item on our test rig. It's not much for a few IOCs, but it compounds quickly. Using the built-in utility at the source eliminated that completely.


Ship fast, measure faster.


   
ReplyQuote
Page 2 / 3