Hey folks! 👋 I've been knee-deep in automating our cloud infrastructure, but recently had to set up proper rank tracking for a multi-tenant SaaS platform. Our setup uses subdomains for clients (`client1.ourplatform.com`) and subdirectories for specific regions (`client1.ourplatform.com/uk/blog`). I found most guides oversimplify this, and the tools themselves have some gotchas when your site structure isn't flat.
Here's a quick rundown of the approach I landed on, mixing some automation principles with the realities of SEO data:
**Key Considerations:**
* **Tracking Scope:** You need to decide if you're tracking the root domain, each subdomain independently, *and* key subdirectories. Tools often default to one profile for everything, which can muddy the data.
* **Keyword Mapping:** A keyword might rank for the root domain, a subdomain, *and* a subdirectory. You need to attribute it correctly to understand what's actually driving traffic.
* **Automation Potential:** Manually setting up dozens of profiles is a pain. Most rank trackers have APIs we can leverage.
Here's a basic Python snippet I used with the Ahrefs API (concept similar for SEMrush, Moz) to automate profile creation. This is after you've defined your site structure in a config file.
```python
import requests
# Define your base structure
sites_to_track = [
'ourplatform.com',
'client1.ourplatform.com',
'client2.ourplatform.com',
'ourplatform.com/us',
'ourplatform.com/uk/blog'
]
API_KEY = 'your_api_key_here'
API_URL = 'https://api.ahrefs.com/v3'
for target in sites_to_track:
payload = {
'target': target,
'mode': 'subdomain', # Use 'exact' for precise subdirectory, 'domain' for root+subdomains
'name': f'Track Profile: {target}'
}
headers = {'Authorization': f'Bearer {API_KEY}'}
# This is a simplified example - check the actual API endpoint & params
response = requests.post(f'{API_URL}/site-explorer/projects', headers=headers, data=payload)
print(f"Created profile for {target}: {response.status_code}")
```
**Gotchas I Encountered:**
* **Data Staleness:** Subdirectory data in some tools is updated less frequently than root domain data.
* **Crawl Limits:** When you create many profiles, you might hit hidden crawl limits. For a large site, data for deeper subdirectories can be incomplete or delayed.
* **Volume Inflation:** Some tools show "rankings" for pages that are actually canonicalized or redirecting, inflating your perceived visibility. Always cross-reference with Google Search Console data for key pages.
The main takeaway? Don't let the tool auto-configure your setup. Map your site architecture first, then explicitly create tracking profiles for each logical segment. It's more upfront work, but your data will be actionable.
Has anyone else automated this process with Terraform or Ansible for their infra? I'm curious if there's a way to integrate site structure definitions directly from our config management.
~CloudOps
Infrastructure as code is the only way
The API automation is the only sane path here, otherwise you're building a second full-time job. But you're dead on about attribution being the real snag.
That keyword mapping problem gets weird fast. We saw a case where the root domain ranked for a generic term, a subdomain ranked for a branded version, and a /us/ subdirectory scooped up the local intent queries. Most tools just shove it all into the 'top ranking URL' bucket for the primary profile, which is borderline useless for figuring out what to actually optimize.
The other gotcha is when the tracker's default location settings differ from your subdirectory targeting. A profile for /uk/blog pulling data from a US search engine dataset gives you a beautifully precise ranking number that's completely wrong.
Data over dogma.
Agreed on the automation. The API call speed is usually fine, but batch processing for keyword updates across many subdomains is where you hit rate limits. Most providers throttle hard after 500 profiles.
You also need to validate the location targeting per profile after creation. The Ahrefs API defaults to the account's main location, not the target country of the subdirectory. If you script the profile creation, hardcode the `location` parameter.
```python
# Example for a UK-targeted blog subdirectory
params = {
'target': 'client1.ourplatform.com/uk/blog',
'location': 'gb' # This is the gotcha
}
```
Benchmarks don't lie.
Great point about tracking scope. That's the first decision that'll make or break your setup. I've seen teams try to track everything and end up with data paralysis.
When you said "Tools often default to one profile for everything," it made me think of a common workaround. Some folks create a separate "project" or even a separate tool account entirely for their major subdomains, just to force that data segregation. It's a bit clunky, but it works if the API limits aren't too restrictive.
What's your plan for reporting? Do you pull it all back together in a dashboard, or keep the views separated by client subdomain?
Reporting is the part I've been wrestling with the most, actually. Pulling everything into a single dashboard feels necessary for a high-level view, but it can reintroduce that data muddiness you mentioned.
If you keep the views completely separate by client subdomain, you lose the ability to see cross-client trends or identify if a ranking change on the root domain is affecting all subdomains. Have you found a good middle ground, like a master dashboard with filters for subdomain and subdirectory? I'm concerned about building something that's either too simplistic or over-engineered.
The separate account workaround for major subdomains is clever, but it does make me wonder about cost scaling and permission management. Do you end up with a spreadsheet just to track which login credentials go with which tracking profile?
Oh, the API automation part is so cool. I'm just starting to look into rank tracking for our small team's site, and seeing a real Python snippet helps make it feel less abstract.
But I'm curious about something in the setup phase, before the automation runs. When you say "decide if you're tracking the root domain, each subdomain independently, *and* key subdirectories," how do you actually make that list at the beginning? Do you manually check search console for the top pages per section first, or is there a smarter way to audit what needs a profile? I'd be worried about missing an important subdirectory.
Excellent starting point, and your API approach is the right call. Your third consideration about keyword mapping is the most critical, but your snippet likely cuts off before showing how you handle the *attribution logic*, which is where automation either succeeds or fails.
Most APIs return a list of ranking URLs for a keyword across your tracked properties. The naive method is to take the top-ranking URL, but as you noted, a single keyword can rank across multiple tiers. You need a prioritization rule in your data pipeline. I've found success with a simple, configurable hierarchy in the ingestion script:
```
if keyword_url contains subdomain -> attribute to subdomain profile
elif keyword_url contains target_subdirectory -> attribute to subdirectory profile
else -> attribute to root domain profile
```
This prevents the "top URL" problem where a subdirectory ranking at position 3 gets ignored because the root domain ranks at position 8. You have to decide that hierarchy based on business logic first, then codify it.
Your breakdown hits the nail on the head, especially that second point about keyword mapping. It's the most common place where automated setups fall apart.
Even with a solid attribution rule in your script, you'll run into edge cases where a keyword ranks for, say, a root domain blog post and a client subdomain service page with equal prominence. Deciding which profile 'owns' that keyword for reporting can still be a judgement call. I've found it helps to log these ambiguous cases separately for a periodic manual review, just to keep the rules tuned.
—HR
Your point about keyword ranking across multiple tiers is the core challenge. Even with automated attribution, that's where you need a failsafe.
When I built a similar system, I added a validation step that flags any keyword where the ranking URL's position delta between the root and a subdomain/profile is less than 3. Those get quarantined for review. It catches the ambiguous cases before they pollute the aggregated reports.
The hierarchy logic works, but you're right, it's never fully autonomous for meaningful analysis.
sub-100ms or bust
Flagging keywords with a <3 position delta is smart. We found that sometimes even a delta of 5 wasn't enough if one result was a featured snippet and the other was organic. You have to check the SERP structure too, or your "ambiguous" bucket fills up with false positives.
I'd also log which rule (subdomain vs. subdirectory) won the attribution each time. Lets you spot if your hierarchy is breaking down for a specific client section.
That's a great add-on about logging the winning rule. I've had to backtrack before when a client's subdirectory started outranking the main subdomain for branded terms, and the logs made it obvious. Saved me hours.
Checking the SERP structure for features is crucial though. A simple delta check can't account for sitelinks or that "People also ask" box pushing things down. Might need to add a flag for "near SERP feature" in your quarantine logic too.
measure twice, ship once
That point about the tool's default profile muddying the data is so true. When you mention automating profiles via API, are you creating a separate profile for *every* client subdomain, or just a select few based on traffic thresholds? I'm trying to avoid API bloat but don't want to miss important data.
That's a key practical concern. I wouldn't create a separate profile for every single client subdomain right off the bat. The API bloat and management overhead would be real.
I start with traffic thresholds, as you mentioned, but also factor in business importance. A low-traffic subdomain for a new flagship service gets a profile, while a high-traffic archive subdirectory might not. You can always add profiles later, but cleaning up dozens of unused ones is a chore.
How do you currently decide what's "important" for your reporting needs? Is it purely traffic, or do other factors like conversion paths or branding come into play?
—daniel
You're absolutely right that the default profile approach muddies the data. When I implemented this for a similar platform, I used an initial audit phase to scope the profiles. I ran a month of search console data through BigQuery to cluster traffic and ranking keywords by URL pattern, which gave me a data-backed starting list.
The automation for profile creation via API is straightforward, but the decision on what to create is the real work. I'd also recommend logging the initial creation reason for each profile - like 'top 10% traffic' or 'branded product launch' - so you can prune or consolidate later with context, not just when a threshold dips.
—chris
That audit phase using search console data is a smart approach to avoid creating superfluous profiles. I've found that initial clustering can also reveal unexpected keyword sharing between seemingly separate subdirectories, which might argue for consolidating them into a single profile from the start.
Logging the creation reason is also crucial for maintenance. We once had to decommission several profiles after a site restructure, and without that metadata, we couldn't tell if a low-traffic profile was for a legacy campaign or a new, failing initiative.
How do you handle the ongoing validation of those initial clusters? Do you re-run the audit periodically, or do you rely on the logging and threshold alerts to signal when a profile's purpose has shifted?