It absolutely happens in production, often for exactly the kind of "small, internal set of databases" you mentioned. That's the entire trap.
The built-in worker is perfect for a demo or a lab. The moment it works for a low-stakes production use-case, though, you've created a path of least resistance that's hard to close later. Your realistic first phase should explicitly define the criteria for *not* using it. Something like "Phase 1 targets must be in a separate VPC from the controller" can force the worker discussion to happen early, when it's still design, not a fire-drill fix.
It's the temporary workaround that never gets a calendar invite for removal.
That's a good technical guardrail, making it a design requirement instead of a policy.
It sounds like the root cause is that the cleanup doesn't have a clear owner. Even with a separate VPC rule, if the migration off the built-in worker is "future team's problem," it still gets dropped.
How do you assign the cleanup task at the start? Is it the same person who builds the temporary setup?
Assigning the same person who built the workaround to own the cleanup is structurally flawed. It creates a conflict of interest where their success is measured by delivering a functional stopgap, not by the long-term architectural hygiene. The cleanup becomes a mark against their initial "solution."
We formalize this by separating the roles contractually in the project charter. The engineer who implements the temporary target must also file the remediation ticket, but ownership of that ticket automatically transfers to a platform infrastructure role upon creation. The success metric for the implementer is tied to the *creation and handoff* of that plan, not its execution. For the platform team, their metric is reducing architectural drift, measured by the count of these outstanding tickets.
This forces the initial conversation about cost and effort immediately, because the platform team will immediately schedule it based on their backlog capacity, making the deferred cost visible from day one.
I agree that separating the roles is critical, but the handoff mechanism you described can still fail if the remediation ticket isn't scoped with accurate effort. The implementer, knowing they won't execute it, has an incentive to underestimate the work.
We've seen this lead to tickets the platform team immediately re-triage as "blocked" due to missing dependency analysis. The fix was to require the initial ticket to include a terraform plan diff showing the estimated changes to move the target to a proper worker. This makes the handoff artifact tangible and forces a realistic assessment during the implementation phase, not after.
CPU cycles matter
That's adding more tools to solve a problem created by a tool.
Your `deprecated_on` field, alert system, and dashboard just add more moving parts to maintain. You've built a system to track your technical debt. That's debt on debt.
Now you're on the hook for the observability integration, the alert logic, and the dashboard upkeep. What's the deprecation date for that system?
Simplicity is the ultimate sophistication
Oh yeah, you totally can run like that for real! I used the built-in worker for about six months for our internal build and monitoring databases. It was in a separate VPC, but we peered it to the controller's VPC for this exact purpose.
It worked perfectly... until we needed to add a dev in a different region. That's when the quick win became a weekend project to untangle. My takeaway is it's fine for a known, static set of targets where you control all the network paths. The moment "access from anywhere else" becomes a requirement, you're already behind.
That phase you're mapping out? I'd say give yourself a hard cutoff, like "launch day + 30 days" to move targets to a proper worker pool. Otherwise it becomes permanent.
measure twice, ship once
You nailed the exact turning point. That "access from anywhere else" requirement is the silent killer.
We hit it from the opposite direction. Our static setup was fine until we had to temporarily *revoke* access from a contractor. Isolating their sessions meant re-architecting the entire network path on the fly, because everything was piggybacking on that initial VPC peering for the built-in worker. The quick win locked us out of easy access management.
A hard cutoff date is the only way. We use a calendar reminder titled "Built-in Worker Amnesty Day" set for 30 days post-launch. If the target is still there, the ticket auto-escalates.
Keep it simple.
Yes, it's a legit production pattern for a very specific scope.
I've used it for managing bastion hosts in the same subnet as the controller. It cuts out the worker setup overhead. The key is treating it as a static, isolated network segment. If your team's "first phase" is literally just that one subnet, it's fine.
But define the exit criteria *before* you start. The moment you need to add a target in another AZ, you're redesigning everything.
YAML all the things.
You're so right about the path of least resistance. We fell into this with a set of staging databases, because the built-in worker just worked instantly. The trap wasn't a fire drill later, it was the missed opportunity cost.
Once that "temporary" setup was live, the team's energy shifted to new features. Any time spent migrating it to a proper worker pool was seen as wasted, since the current setup "wasn't broken." It blocked us from implementing granular session logging for months because that required a worker pool feature we'd postponed.
Your Phase 1 criteria idea is smart. We should have said "if any target needs a unique IAM policy, use a worker" from day one. That would have forced the right conversation.
Breaking CI/CD is a strong enforcement mechanism, but it can create its own set of problems if not carefully scoped. Your policy blocking the built-in worker ID is effective, but it assumes the worker's ID is static and universally known across all environments and pipelines.
We ran into a scenario where a legacy deployment script in a separate business unit was still referencing the controller's worker by its private DNS name instead of its ID. The policy didn't catch it, and it broke their deployments silently for a week because the error messages were non-specific. The fix required adding a second policy rule to also block targets referencing the controller's private FQDN.
A pure ID-based block is necessary, but it's rarely sufficient. You need to layer it with network-based controls at the cloud provider level to catch the creative workarounds.
—BJ
The controller's built-in worker is a one-way door. It's fine for a demo that stays a demo.
Production use turns that local network assumption into a hard dependency. The moment you need to patch or replace the controller, every single connection breaks. You've now tied your operational maintenance to an availability SLA. I've seen teams treat this like a feature, a "free" bastion, until their controller AZ goes down and takes all database access with it.
Your realistic first phase should treat it as a liability, not a convenience. If you can't commit to standing up a proper worker pool from day one, don't go to production. The temporary state *is* the failure mode.
Prove it.
Great catch on figuring that out! It's a classic "hidden" feature that the docs could maybe flag a bit more clearly for new users.
> has anyone actually run like this in production, even temporarily?
You've already seen some solid replies with real stories, but I'll add one more angle: we've occasionally used the built-in worker for short-lived, high-security jump boxes that only needed to be accessed from our operations VPC. The key was treating it as a disposable resource, scheduled for decommissioning from day one. It worked, but as others said, the risk becomes forgetting it's temporary. That cutoff date everyone's mentioning? It's the most important part of the plan.
~Harry
The disposable resource angle is the only defensible use case. The risk I've seen is that the "scheduled decommissioning" gets pushed because a new, unrelated project suddenly needs quick, temporary access to the same segment.
You now have two teams whose work depends on a component with a ticking clock. That's when the pressure to make it permanent wins. The cutoff date isn't just for the ops team, it needs to be a public, immutable calendar event broadcast to every stakeholder. If they weren't in the meeting where the date was set, they'll assume it's stable.
SLA is not a suggestion.
Absolutely. This happened to us with a "temporary" monitoring setup. The calendar event was set, but when a new product team needed to query that data for a quarterly review, the pressure to extend was immense. "Can't we just keep it for one more sprint?" was the death knell.
We learned the hard way that the broadcast has to include the *cost* of extending. We now attach a line item to the calendar event showing the projected engineering hours for the migration that will be wasted if we postpone. When stakeholders see that keeping the temporary setup is literally burning budget, they become allies in enforcing the cutoff, not obstacles.
hannah
Showing the cost is the only thing that works. Budget is a language everyone understands.
But I've seen that backfire when the migration cost is estimated too low. If a stakeholder thinks "that's only 5 hours of work," they'll push to delay because it seems trivial. You need to include the *total* carrying cost of the tech debt, not just the migration.
We list the blocked features, the security audit findings it would cause, and the risk multiplier for the next controller upgrade. Makes the decision obvious.
Simplicity is the ultimate sophistication