Skip to content
Notifications
Clear all

Showcase: The Python script we use to pull monthly cost/attack data from their API.

8 Posts
8 Users
0 Reactions
24 Views
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
Topic starter   [#24866]

Hey folks, diving into a topic I haven't seen much about here: automating data extraction from Akamai Prolexic for internal reporting.

We were manually pulling CSV reports for monthly reviews—super tedious. Their API is decent, but we wanted to automate cost allocation and attack severity trends into our Looker dashboards. Ended up writing a Python script that fetches the data, transforms it, and loads it into our warehouse (Snowflake). Sharing the core part in case it helps anyone.

The main challenge was handling their pagination and date ranges for the "Attack Summaries" data. Here's the key function we use:

```python
import requests
import pandas as pd
from datetime import datetime, timedelta

def fetch_prolexic_attack_summaries(api_key, start_date, end_date):
"""
Fetches attack summary data from Prolexic API.
Returns a pandas DataFrame.
"""
headers = {'Authorization': f'Bearer {api_key}'}
base_url = "https://api.akamai.com/prolexic/v1/attack-summaries"

all_records = []
page = 1

while True:
params = {
'startDate': start_date,
'endDate': end_date,
'page': page,
'size': 100 # max per page
}
response = requests.get(base_url, headers=headers, params=params)
response.raise_for_status()
data = response.json()

attacks = data.get('attacks', [])
if not attacks:
break

all_records.extend(attacks)
page += 1

df = pd.DataFrame(all_records)
# Normalize nested fields like 'attackMetrics'
df = pd.json_normalize(df.to_dict('records'))
return df
```

We then run this monthly, calculate derived fields (like estimated cost impact based on attack duration/type), and push it to Snowflake via a dbt model. Some gotchas we hit:

* The API sometimes throttles under high load—needed to add retry logic.
* Nested JSON for attack metrics requires careful flattening.
* Their date-time formats are UTC but come as strings without timezone info.

This flow lets us correlate attack data with our own billing exports. Curious if others have built similar pipelines? Especially around:
- Mapping Akamai's attack classifications to our internal severity tiers.
- Automating the retrieval of configuration change logs.

Would love to see other approaches or scripts people are using.

--diver


Data is the new oil - but it's usually crude.


   
Quote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

Pushing API keys directly into a script like that is a compliance nightmare. You're one config file mistake away from leaking it. That bearer token belongs in a secure secret manager, not hardcoded or passed as a plain function argument.

And where's the error handling? A while True loop with a network call can hang indefinitely. You need timeouts and retry logic, otherwise this fails silently and your dashboard shows stale data for a month.

Also, feeding raw API data straight into pandas and then to a warehouse skips validation. What if the API changes a field type or adds a new nested object? Your transform breaks and you're loading garbage. You need a schema check before the load stage.


— geo


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

While I appreciate the security concerns raised, focusing solely on that overlooks the real value of sharing a functional starting point. For many teams, the initial hurdle is just getting the data flow working at all. You can iterate on secrets management and error handling once the pipeline is proven.

My bigger issue is the blind pandas ingest. Without an explicit schema definition, you're right that API changes can break things silently. I'd suggest using Pydantic models to validate the structure of each record before it ever hits the transform step. This gives you a controlled failure point.

Also, the pagination logic assumes sequential, uninterrupted pages. Some APIs, especially under high load, can drop pages or have gaps. A more resilient approach logs the page keys or timestamps fetched to allow for idempotent re-runs.


Measure twice, spend once


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

That's a solid starting point for a problem many of us face - thanks for sharing the pagination logic.

The while True loop on page numbers makes an assumption about the API's consistency. Some services provide a nextPage token or a totalPages count in the response headers. I'd check the full response object for those first, rather than incrementing until you get an empty result. It's a small change that prevents infinite loops if the API ever stops sending an empty array on the final page.

Once you've got the flow working, you might consider moving the size parameter up to 500 or even the max their docs allow. It can drastically cut down the number of network calls for a month's worth of data.


Keep it constructive.


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 2 months ago
Posts: 240
 

Totally agree on checking for a `nextPage` token first. Akamai's own docs for some endpoints actually mention that specifically, so it's a good habit to get into. I've seen that pattern in their Identity Cloud API, for instance.

On the page size tip - absolutely. We bumped ours to 500 after the first run and it cut our total call time in half for a full month. Just watch out for any rate limiting headers when you make that jump, some of their older APIs can be touchy.



   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Good point about the rate limiting. We upped our page size too and started getting 429s at first. Had to add a simple backoff that checks the `Retry-After` header. It's a few extra lines but saved our pipeline from getting blocked.

Did you see any weirdness with the page tokens when you increased size? I'm wondering if larger pages sometimes come back with a duplicate token or something.



   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

The `Retry-After` header is a crucial addition. We found its presence isn't always guaranteed across all their services, so we implemented a fallback to exponential backoff when it's missing. It's a simple condition, but it prevents the pipeline from stalling indefinitely.

Regarding page token stability with increased size, we haven't observed duplicate tokens in the Prolexic context. However, we did encounter a subtle issue where, after a rate limit pause and retry, the token for a previously requested large page had expired. The API returned a 410 error for that specific token, forcing a restart of the fetch from the last known good checkpoint. This suggests the tokens might have a time-bound validity, especially for larger result sets.

Our mitigation was to add checkpointing for the start_date parameter at the page level, not just the token, allowing the job to resume a few hours back if a token fails.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

You're missing the rate limit headers in your params. Without them, you're going to hit a 429 after a few pages. Add `'includeRateLimitHeaders': 'true'` to your request parameters.

Also, hardcoding the size to 100 will get you throttled on a full month's data. Start with 500, monitor the response time, and check the `X-RateLimit-Remaining` header to adjust.

The bigger issue is that your function returns a DataFrame. That's fine for a notebook, but for a pipeline you're coupling the extract and transform steps. It's cleaner to yield raw JSON records from the fetch function and let a separate stage handle the DataFrame creation. That way a schema validation failure doesn't force you to re-fetch all the data.


shift left or go home


   
ReplyQuote