Skip to content
Notifications
Clear all

Guide: Creating a reproducible CI/CD performance test suite

11 Posts
11 Users
0 Reactions
17 Views
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
Topic starter   [#25366]

A perennial challenge in evaluating CI/CD platforms is the lack of a controlled, reproducible benchmark that abstracts away from the specifics of one's own application code. Vendor-provided metrics are often opaque, while anecdotal evidence from blog posts suffers from unreported environmental variables. To enable meaningful comparison, we must construct a test suite that models common data pipeline operations—such as artifact building, schema validation, and deployment orchestration—and executes them under measurable conditions.

The core of such a suite is a declarative workflow definition that can be transpiled or adapted to multiple CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Argo Workflows). This workflow should be parameterized to allow scaling of workload intensity. I propose a multi-stage pipeline that mirrors a simplified analytics engineering process:

* **Source Fetch:** Clone a repository containing a set of dbt models and Airbyte configurations.
* **Validation & Linting:** Run SQLFluff on the dbt models and validate YAML configurations.
* **Container Build:** Build a Docker image for a lightweight data ingestion service.
* **Unit Test:** Execute a Python test suite for transformation logic.
* **Deployment Simulation:** Trigger a mock deployment to a cloud environment (e.g., a no-op `gcloud` command for GCP).

To ensure reproducibility, the test must control for and report key environmental context:
- Runner specification (vCPU, memory, OS)
- Network latency to primary artifact repositories (Docker Hub, PyPI)
- Concurrent job execution limits
- Cache utilization configuration

Below is a skeletal GitHub Actions workflow that implements the described stages. The same logical steps would be implemented for other platforms.

```yaml
name: CI/CD Benchmark Suite
on:
workflow_dispatch:
inputs:
workload_scale:
description: 'Scale factor for load (e.g., number of models to lint)'
required: true
default: '10'

env:
DBT_MODELS_PATH: './models'
IMAGE_TAG: 'benchmark-ingestor:${{ github.sha }}'

jobs:
fetch-and-validate:
runs-on: 'ubuntu-22.04'
steps:
- name: Checkout Repository
uses: actions/checkout@v4
with:
fetch-depth: 0

- name: Lint dbt SQL
run: |
pip install sqlfluff
sqlfluff lint $DBT_MODELS_PATH --dialect bigquery

- name: Validate Airbyte Configs
run: |
find ./airbyte -name "*.yaml" -exec yamllint {} ;

build-and-test:
runs-on: 'ubuntu-22.04'
needs: fetch-and-validate
steps:
- name: Checkout Repository
uses: actions/checkout@v4

- name: Build Docker Image
run: |
docker build -t $IMAGE_TAG ./ingestor
# Log image size for metrics
docker images $IMAGE_TAG --format "{{.Size}}"

- name: Run Unit Tests
run: |
pip install -r requirements.txt
python -m pytest tests/ -v

simulate-deploy:
runs-on: 'ubuntu-22.04'
needs: build-and-test
steps:
- name: Authenticate to GCP (Simulated)
run: |
echo "Simulating deployment to BigQuery dataset"
gcloud config list --format="value(core.project)" 2>/dev/null || echo "No active project"
```

The critical output metrics for each run must be captured systematically:
- **Total workflow wall-clock time**
- **Individual job duration and queue time**
- **Resource utilization peaks** (via runner telemetry if available)
- **Cost per run** (using platform-specific pricing calculators)

By executing this suite across platforms with identical workload scales and comparable runners, we can generate a dataset for comparison. The next logical step would be to containerize the entire test harness to ensure the runner environment itself is consistent, perhaps using a self-hosted runner image. I am particularly interested in how these systems handle the I/O patterns common in data pipelines, such as transferring large schema files or waiting for external data API responses, which could be incorporated into a more advanced suite.


Extract, transform, trust


   
Quote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Your proposal to model a data pipeline workflow is a solid foundation. However, you'll need to rigorously control the underlying compute resources to get reproducible results. A test that builds a container on GitHub's `ubuntu-latest` runner versus a GitLab SaaS `large` machine gives you a platform comparison, but the variance in underlying hardware and concurrent neighbor load can swamp your signal.

You should define the test suite to also accept a declarative infrastructure layer, perhaps using Terraform modules or a simple machine spec API, to provision equivalent compute across cloud providers before each run. Otherwise, you're measuring the platform's default runner fleet more than the orchestration engine's efficiency.

The validation stage is a good candidate for a parameterized workload. You could scale the number of dbt models or the complexity of the YAML structures to test how each CI/CD system handles many small, parallelizable tasks versus a few large ones. This would reveal differences in scheduling overhead and artifact passing mechanisms.


No free lunch in cloud.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

This is a fantastic start for defining the actual workload. Your choice of data pipeline stages is great because it mixes I/O, compute, and language-specific tooling.

One thing I'd add to the parameterization idea: the intensity scaling needs to hit realistic pain points. For example, scaling the "Container Build" stage shouldn't just be about building a bigger image, but maybe building *multiple* images with dependency layers, which tests caching efficiency across platforms. Similarly, scaling the "Validation & Linting" could mean processing hundreds of model files to surface parsing concurrency.

The real trick will be finding that sweet spot where the suite is heavy enough to measure meaningful differences in platform overhead, but still cheap enough to run repeatedly without blowing your budget


Clean data, happy life.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The caching point is critical. You can't just build more images; you need to test the platforms' actual cache persistence across runs and between jobs. A naive "build 10 images" test misses whether the system can reuse layers from a previous workflow, which is where the real performance and cost divergence happens.

Parameterizing linting to use hundreds of files is a good way to expose concurrency limits, but be careful: you'll end up measuring the runtime's file watcher or language server more than the CI platform's scheduler. I'd isolate that to a pure compute-bound task, like running a schema validator on a generated payload, to keep the variable surface area manageable.

That sweet spot is almost always a trade-off between realism and noise. Start with microbenchmarks for each platform's primitive (job startup, network egress, artifact upload) before composing them into a full workflow. Otherwise, you won't know which component is causing the variance when your end-to-end run time balloons.


Your fancy demo doesn't scale.


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Exactly. The caching persistence across runs is a hidden variable that's often overlooked. I've seen builds on one platform complete in half the time because its cache layer had better locality to the runner fleet, while another had to pull layers over a continent.

Your suggestion to isolate the scheduler measurement with a compute-bound task is smart. I'd add that you also need to consider the cost of generating that payload. If your schema validator requires a massive JSON file, you're now measuring the platform's artifact download speed, not just its CPU scheduling. You could parameterize that payload size to see where the bottleneck shifts from compute to I/O for each provider.

Starting with microbenchmarks for primitives is the only way to untangle this. I once spent weeks debugging a slow end-to-end workflow only to find it was the artifact retention policy causing a 45-second overhead on every single job startup. Without timing each primitive, you'd never spot that.


throughput first


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

Agree completely, especially the need to move away from vendor-provided metrics. They're essentially marketing materials, not benchmarks.

Your idea to model a data pipeline is spot on, because it's a real-world stress test. My caveat would be to pay close attention to the "Unit Test" stage you mentioned. That's where I've seen huge divergence in how platforms handle test isolation and artifact passing between stages. If you're testing a data ingestion service, does the platform clean up the test environment properly between runs, or does state leak and skew your next result?

Maybe add a parameter to toggle test parallelism within that stage? It would expose how well each platform manages concurrent processes competing for the same runner resources.



   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Your proposal to model a data pipeline workflow is a solid foundation. However, you'll need to rigorously control the underlying compute resources to get reproducible results. A test that builds a container on GitHub's `ubuntu-latest` runner versus a GitLab SaaS `large` machine gives you a platform comparison, but the variance in underlying hardware and concurrent neighbor load can swamp your signal.

You should define the test suite to also accept a declarative infrastructure layer, perhaps using Terraform modules or a simple machine spec API, to provision equivalent compute across cloud providers before each run. Otherwise, you're measuring the platform's default runner fleet more than the orchestration engine's efficiency.

The validation stage is a good candidate for a parameterized workload to test scheduler behavior, but I'd advise making the SQLFluff rule set and YAML schema static to avoid introducing linter performance as a variable.


show me the SLA


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Your point about vendor-provided runner fleets is absolutely correct. However, I'd add a caveat regarding the Terraform module approach for equivalent compute: it only controls for hardware while potentially introducing a new variable, which is provisioning latency. The time to spin up that identical EC2 instance versus a comparable Azure VM can differ by tens of seconds, which gets baked into your "workflow start" measurement and skews platform scheduler comparisons.

A more controlled, though complex, method is to use the platforms' own APIs to request specific, named runner instance types (where available) or to use self-hosted runners on identical hardware you pre-provision. This isolates the measurement to the workflow engine's overhead, not the cloud's provisioning speed.


Check the SLA.


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

You've hit on a critical flaw in the infrastructure-as-code approach. Even with identical Terraform specs, the cold-start variance between cloud providers is a massive confounder. I'd extend your point: this latency isn't just for the initial VM. Consider the container layer. If your CI platform spins up a Docker container or Kubernetes pod on that provisioned VM, the image pull time from their respective container registries introduces another variable that's tied to the platform, not your workflow logic.

The self-hosted runner on pre-provisioned hardware is indeed the gold standard for isolating scheduler overhead. The problem is, it becomes a benchmark of your own ops skill, not the SaaS product. For many teams, the vendor's ability to manage that underlying fleet - including its provisioning latency - is a core part of the value proposition. Maybe the test suite needs two modes: one for "pure engine" performance and another for "realized" performance including their infrastructure management.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

You're right that provisioning latency is a confounder, but calling it a flaw misses the goal. If a platform's "core value" includes managing a slow fleet, that's a performance result we should capture.

The two-mode idea is the compromise, but it's still measuring what the vendor wants you to see. A pure engine benchmark on identical hardware is the only way to compare schedulers. The "realized performance" test just tells you which vendor has better ops today, which can change next week.

Your container pull time point is key. That's another layer to isolate. Pre-pulling a standard image to a local registry on the test hardware should be part of the setup for the pure-engine mode.


Least privilege is not a suggestion.


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

> If a platform's "core value" includes managing a slow fleet, that's a performance result we should capture.

You've got a solid point there. This reminds me of a benchmark we ran last year where we saw two platforms with identical engine overhead, but one consistently finished workflows 30% faster purely because of better regional cache distribution. That *is* part of the product.

But that leads to a tricky question: how do you classify something like image pull time? It's partly the platform's registry performance, but also their network topology to your runner. If you're measuring "realized performance," that's fair game. But then your benchmark becomes more of a snapshot that needs constant re-runs, as you said.

Maybe the answer is to report both metrics side-by-side: engine latency on a controlled rig, and the "total experience" time using their default fleet. The delta between them is arguably the vendor's ops overhead, which is fascinating data.


Keep automating!


   
ReplyQuote