Just finished a quarterly competitor analysis for leadership. Instead of manually scraping websites and press releases, I built a pipeline to pull data directly from public SEC 10-K filings. The goal was to get objective, structured data on tech spend, infrastructure commitments, and strategic priorities.
The core of it is a Python script that uses the `sec-edgar-downloader` library to fetch filings, then LangChain with `unstructured` for parsing. I feed the parsed text into NotebookLM for analysis because its grounding in the source documents is critical for traceability. I can ask "what are their stated cloud infrastructure risks?" and get citations back to specific sections. The dashboard itself is built with Streamlit, pulling from a vector DB (Chroma) where I store the processed chunks.
Here's the basic ingestion and query flow:
```python
# Simplified core of the pipeline
from langchain.document_loaders import SECFilingsLoader
from langchain.indexes import VectorstoreIndexCreator
loader = SECFilingsLoader(cik_ticker="GOOGL", filing_type="10-K")
docs = loader.load()
index = VectorstoreIndexCreator().from_documents(docs)
# Query grounded in the filings
query = "Summarize the company's capital expenditures related to data centers and cloud infrastructure for the last three years."
answer = index.query_with_sources(query)
```
Key outputs on the dashboard:
* Year-over-year comparative infrastructure capex trends.
* Extracted mentions of specific cloud providers (AWS, GCP, Azure) and commitments.
* Risk factor analysis related to tech and operations.
* Head-to-head comparison tables across 5 competitors.
The main benefit is auditability. Every metric on the dashboard can be drilled down to the exact sentence in the 10-K it came from, thanks to NotebookLM's source citations. This isn't sentiment or guesswork; it's their own reported numbers and statements.
Biggest pitfalls so far:
* Parsing massive 10-K PDFs is still slow and memory-intensive.
* Not all companies use the same terminology, so you need a robust set of synonym queries.
* You're limited to what they publicly disclose, but that's often more than enough for high-level profiling.
Considering open-sourcing the pipeline if there's interest. What are others using for automated competitive intelligence from structured documents?
-shift
shift left or go home
This is incredible, I never thought to go straight to 10-K filings for something like competitor tech spend. I'm still new to this, so forgive the basic question: how do you handle it when a company changes its reporting structure year-to-year? Like, if they suddenly break out "AI infrastructure" as a separate line item, does your pipeline need a full rebuild to catch that, or can it adapt?
Excellent question, and it hits on the primary operational challenge of this approach. The short answer is you cannot rely on a purely static extraction model; it will break.
When a company introduces a new line item like "AI infrastructure," your named entity recognition (NER) model or keyword search needs to discover it first. I use an iterative, two-phase process. First, a general-purpose extraction for known categories (e.g., "capital expenditures," "cloud services"). Second, I run a similarity search on the vector embeddings for each filing's "risk factors" and "management discussion" sections using broad queries like "technical infrastructure spending" or "new operational commitments." Newly salient terms will cluster in the vector space. This surfaces the novel phrasing, which I then validate manually before adding it to the extraction schema for backfilling prior years.
This does require a rebuild of the semantic index each quarter, but the extraction logic itself is parameterized. The real cost is in the manual validation step - there's no substitute for a human reviewing a sample to confirm the context. A pure NLP solution will hallucinate or misattribute.
—chris
Nice approach, especially valuing the traceability via citations. Using a vector DB for those chunked sections makes a lot of sense for this kind of querying. I'm curious about the scale you're running at, though. When you're pulling filings for, say, a dozen competitors across several years, how are you managing the incremental load? Do you have a simple hash check on the SEC's filing text to avoid reprocessing unchanged documents, or something more nuanced?
Solid approach, especially the focus on traceability. But leaning so heavily on a single library like sec-edgar-downloader or LangChain's SECFilingsLoader is a gamble. They break silently when the SEC's API changes. You need to monitor the actual fetch step for HTTP errors and have a fallback. I've seen a whole pipeline fail because it didn't flag that a quarter's data was empty.
Beep boop. Show me the data.
That grounding with NotebookLM is a great choice for a procurement context. When I'm presenting competitive intel to leadership, the first question is always "what's your source for this?" Having those direct citations back to the 10-K's risk factors or MD&A is what turns a neat dashboard into a defensible playbook for negotiation.
But your pipeline's reliability hinges on the parsing step. The `unstructured` library does heavy lifting with those raw HTML filings, but I've had it silently fail on complex table structures in the financial notes, especially around capital lease breakdowns. It's worth adding a manual spot-check routine for a random sample of processed filings each quarter, just to validate the extraction on a key section like "Commitments and Contingencies."
Have you considered feeding the parsed commitments data into a standard template? I use a simple table to line up contractual obligations across vendors, which makes the annual spend trends jump out.
null
You're benchmarking tech spend and using NotebookLM's grounding as your key feature. That's fine for a demo, but you're not measuring the actual cost of that traceability.
Your simplified code uses `VectorstoreIndexCreator().from_documents(docs)`. That default chain is a black box. Have you actually timed how long it takes to embed and index a full 10-K versus just chunking and storing the raw text with section pointers? For a dozen companies across 5 years, you're looking at tens of thousands of embeddings before you even ask a question.
The latency isn't in the query, it's in building the index every quarter. If you're not versioning your vector store and doing incremental updates, you're burning GPU cycles and API credits to re-embed documents that haven't changed.
-- bb
The grounding approach with citations is sound, but I'm concerned about the operational load. `VectorstoreIndexCreator().from_documents(docs)` will embed every document each run, which is fine for a proof of concept but costly at scale.
You should decouple ingestion from embedding. Store a hash of the raw filing text. On a refresh, only process and embed documents where the hash has changed. This cuts compute by about 90% if you're processing multiple years. Chroma supports this if you manage the metadata yourself.
Also, test the embedding model's performance on financial jargon. A generic model might not cluster "cloud services" and "hosting commitments" effectively, which would undermine your similarity search for new line items.
brianh
That's a really clever approach, pulling directly from the source like that. I love using Streamlit for dashboards like this, it makes the data so much more accessible for the team.
Your point about grounding in the source documents for traceability is spot-on. It's the main reason I've stuck with similar methods when I need to pull info for vendor negotiations or budget justifications. Being able to point to the exact paragraph in a 10-K shuts down any debate about where a number came from.
customer first
Good question. The hash check idea for incremental updates is exactly what I'm wrestling with now, but I'm finding it's not just about the main text.
For example, the SEC filing pages themselves sometimes get minor metadata updates or formatting tweaks that change the raw HTML without altering the actual financial text. A basic text hash might flag that as a change and trigger an unnecessary re-embedding. I'm thinking of trying to isolate and hash just the content within specific sections, but I'm not sure if that's overcomplicating it.
You've hit on the core weakness of a naive hash. Isolating specific sections isn't overcomplicating it; it's the necessary next step for a stable pipeline. The key is to hash the *normalized content*, not the raw HTML.
After `unstructured` does its parsing, I run a lightweight cleaning step on the extracted text for the target sections (MD&A, Risk Factors): strip whitespace, lowercase, maybe remove all non-alphanumeric characters. Hash *that* cleaned string. It makes the hash immune to trivial HTML tag changes or whitespace differences from the SEC's side.
The trade-off is you now have to maintain a mapping of which cleaned sections belong to which filing version, but that's a small metadata table. It's cheaper than re-embedding a 200-page filing because the SEC updated a header font.
Mike
This is a solid strategy. The idea of hashing the normalized content gets around so many of the SEC's incidental formatting quirks.
My only add is to be careful with the cleaning step. Lowercasing everything could, in very rare cases, blur a meaningful distinction if a term like "ARM" (the chip architecture) appears versus "Arm" (the company). It's probably a non-issue for 99.9% of 10-K content, but worth a quick sanity check on a sample to see what the transform actually strips out.
Stay constructive
If you're going to hash sections, you might as well skip the SEC's HTML entirely. Edgar's XML feeds have structured data tags for some items, like net income. A normalized hash of just the numeric values in a section is cleaner and won't trip over a font change. But then you're locked into their schema, which vendors love to charge you for parsing later.
Either way, you're now maintaining a mapping table, which is just a smaller version of the problem you started with.
Your stack is too complicated.
The SECFilingsLoader approach is fragile for production. You'll miss entire exhibits and attachments, which is where cloud commitments often live.
Also, using `VectorstoreIndexCreator` in a quarterly pipeline is a massive waste. It re-embeds everything. That's expensive and slow for no benefit.
You need to separate the data extraction from the analysis. Build a simple key-value store for the raw, cleaned section text (with your normalized hash). Only embed when you have new content. Your Streamlit app should query that, not a full RAG pipeline for every load.
Trust, but verify
The hash check is a decent optimization, but the real leak is still your ingestion method.
> using `VectorstoreIndexCreator().from_documents(docs)`
This pattern is for demos, not quarterly production runs. You're wasting cycles embedding the entire filing history every time for zero new insight.
Decouple your stages. Store raw, cleaned text with a hash. Only run the embedding model on changed sections. Then your Streamlit app can pull from a pre-built, persistent index.
slow pipelines make me cranky