Skip to content
Notifications
Clear all

Thoughts on the new Pulumi Automation API? Could be a game-changer for our CI/CD.

3 Posts
3 Users
0 Reactions
27 Views
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
Topic starter   [#17068]

Hey folks, been knee-deep in CI/CD pipeline refactoring for our multi-cluster K8s setup, and I've been experimenting with the Pulumi Automation API for the last few weeks. I have to say, this feels like it might fundamentally shift how we think about infrastructure deployment within our automation loops.

Traditionally, our pipelines have been a mix of `pulumi up` shell commands and a patchwork of scripts to manage state and handle dynamic config. It worked, but it was clunky, especially when we needed to make decisions based on previous stack outputs or handle rollbacks gracefully. The Automation API lets you treat Pulumi programs as a library, essentially embedding them directly into your Go, Python, or Node.js automation code.

Here's a tiny, concrete example from our world. Imagine you need to provision a namespace and then create a ConfigMap within it, but only if a certain condition from a previous deployment is met. With the Automation API, you can do this in a single, typed, and testable process:

```typescript
// Inside a Node.js CI runner script
import { LocalWorkspace } from "@pulumi/pulumi/automation";

async function deployConditionalConfigMap(clusterEnv: string) {
const stack = await LocalWorkspace.createOrSelectStack({
stackName: `k8s-${clusterEnv}`,
projectName: "app-config",
program: ourInfraProgram, // This is just a function
});

const currentOutputs = await stack.outputs();
const shouldCreateConfig = currentOutputs.namespaceStatus?.value === "ready";

if (shouldCreateConfig) {
await stack.up({ onOutput: console.log });
console.log("ConfigMap deployed based on namespace state.");
}
}
```

The real power, in my view, comes from these trade-offs:

* **Pro: Unprecedented Flexibility.** Your CI/CD code and your infra code are now in the same logical context. You can loop, branch, and react using your programming language's full power, not just what a CI YAML syntax allows.
* **Pro: Improved Observability.** You can capture events, outputs, and state transitions directly in your app's logs, feeding them directly into your observability stack (think OpenTelemetry) without parsing CLI output.
* **Con: Increased Complexity.** You're now writing more code vs. declarative YAML for the pipeline itself. This requires good software practices (testing, error handling) that some platform teams might not be ready for.
* **Con: Vendor Nuance.** You're committing more deeply to Pulumi's model. While the core concepts are portable, the Automation API code itself is not.

For us, migrating our GitOps workflows to use this API is a significant refactor. We're moving from a "Terraform-like" outer shell script wrapper to an "SDK-driven" model. The question I'm wrestling with is whether the increased control and debuggability is worth the lift. Has anyone else here started down this path, particularly in a Kubernetes context? I'm especially curious about handling state import/export and concurrent operations safely.

—Chris


Prod is the only environment that matters.


   
Quote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

It might shift your thinking, but probably toward more lock-in and complexity. Embedding Pulumi as a library just moves the clunky patchwork from shell scripts into your codebase. Now your application logic is permanently married to Pulumi's runtime and state management model.

That conditional ConfigMap example sounds neat, but you're describing a basic workflow that any decent pipeline orchestrator should handle. Introducing a whole new programming model to avoid writing a simple "if" statement in your existing CI system feels like a solution in search of a problem.

Have you run into the fun of debugging when your embedded program's state diverges from the Pulumi service? The promised "testability" tends to melt away when you realize you're still mocking out a cloud provider.


Show me the data


   
ReplyQuote
(@jasonc)
Estimable Member
Joined: 3 months ago
Posts: 60
 

You're hitting on the crucial distinction that makes the Automation API compelling: it's not just about replacing shell commands. The real shift is treating infrastructure as a true programmatic dependency within your orchestrator, enabling direct control flow and state interrogation.

That "clunky patchwork of scripts" you mentioned for handling dynamic config and rollbacks is precisely what this replaces. Instead of parsing JSON outputs from `pulumi stack output` or managing separate state files, your orchestration code can call `stack.outputs()` directly and use the results in a type-safe way to make decisions. The rollback logic becomes a try-catch block around `workspace.up()`, with the ability to programmatically revert or trigger a notification workflow.

The complexity argument is valid, but it exchanges one form of complexity for another. You're trading the complexity of coordinating disparate scripts and state-passing mechanisms for the complexity of a direct library dependency. The latter, in my experience, is far more amenable to unit testing and integration patterns we already use for application code. You can mock the `LocalWorkspace` class itself, for instance, without mocking an entire cloud provider.

Have you looked into using the inline programs feature? It lets you define the entire Pulumi program as a function within your automation script, which really blurs the line between infrastructure code and pipeline logic. It's a double-edged sword, but fascinating for self-contained deployments.


API whisperer


   
ReplyQuote