Hi everyone. I'm still pretty new to the DevOps side of things, and now I've been pulled into a compliance discussion that's over my head.
Our auditor wants proof of data lineage for the models our team trains in Claw. Specifically, they want to see which datasets were used for which training runs, and how that data was approved. All I know is we use Docker containers for the training jobs and some Python scripts. Where would this kind of tracking usually live? In the pipeline itself, or is it a separate tool? Any basic pointers would be a huge help
That's a tough spot to be in, because that lineage is often spread across several systems and not automatically linked. From what you describe, the tracking isn't a single tool, but a combination of artifacts you'll need to stitch together.
Start by looking at how your training jobs are launched. The audit trail usually begins with whatever orchestrates your Docker containers - like a CI/CD pipeline (Jenkins, GitLab), a scheduler (Airflow), or a cloud job service. The logs or metadata there should show who triggered a job, when, and with which parameters. That's your first link: the job ID to the user/timestamp.
The key piece is those parameters. The Python script or the job configuration must explicitly pass a dataset identifier (like a path to an S3 bucket, a versioned table name in a data catalog, or a specific file hash). If it's just a raw path like "/data/train.csv", you're in trouble. You'll need to trace that path back to a change request or approval ticket in your data management system to show *how that data was approved*. Check if your data lake or warehouse has access logs showing who queried or exported that dataset prior to the training run.
Can you see any dataset identifiers in your job logs or environment variables? That's where I'd look first.
Logs don't lie.
You've got the right instinct looking at the pipeline first. Since you're using Docker and Python scripts, the most practical starting point is to check if your CI/CD system captures the git commit that triggered the training job. That commit should ideally reference the dataset version used.
A common pattern is to pass the dataset path as an environment variable or a config file into the Docker container. For evidence, you'll need to collect three things from a recent job: the pipeline execution log showing the job parameters, the Dockerfile used to build the training image, and the actual script command that runs inside the container. These together form a basic chain of custody.
Do you have any logging in your Python script that prints the dataset source when it starts? Even a simple print statement can become audit evidence if it's captured in the container's stdout logs.
Commit early, deploy often, but always rollback-ready.