Hey everyone, my org just finished rolling out GitHub Copilot to all our devs (around 100 people). I'm coming from a Linux sysadmin background, so watching this rollout was fascinating from an infra and process angle.
We saw some immediate friction. Our internal, older APIs aren't well-documented in the codebase, so Copilot kept suggesting wrong or outdated methods. It also really struggled with our custom Helm chart structures. On the plus side, it stuck perfectly for boilerplate stuff—Dockerfile stages, basic Kubernetes manifests, and even some Terraform modules for AWS. My team loved it for writing Ansible playbooks.
For those who've been through this, what were the biggest pain points in your stack? And more importantly, how did you adapt your onboarding or documentation to make the AI suggestions more useful? Trying to help our platform team smooth this out.
Interesting to hear about your experience with undocumented APIs. We had a similar issue with our internal Python packages. Copilot kept suggesting deprecated functions.
What helped us was adding small, targeted docstrings to the most used internal functions. Just a one-line description of what it does now. The platform team also created a short "common pitfalls" guide for our stack.
Did you track how often devs accepted the suggestions for boilerplate versus custom code? I'd be curious if acceptance rates were a good signal for where documentation was needed.
Our rollout had the exact same friction with custom Helm charts. It kept hallucinating non-existent values.
We found two things that worked:
- Created snippet libraries for the problematic patterns (like our ingress template) and pointed Copilot to those files first.
- For the older APIs, we didn't have time for full docstrings. Instead, we added a simple // DEPRECATED: use package/v2/whatever comment above the old function declarations. That stopped most of the bad suggestions cold.
Your point about boilerplate is spot on. That's where it's actually reliable. We now explicitly tell new hires: use it for Docker, K8s, and Terraform boilerplate. Ignore it for our internal service layer until you know the ropes.
The snippet library approach is solid. We did that with our custom Jenkins pipelines. Copilot kept suggesting wrong agent labels until we seeded it with a few real examples from our repo.
But that DEPRECATED comment trick is dangerous. If your IDE's linter doesn't flag it, a human might still copy the old function by accident. Better to actually remove export or add a runtime warning.
Our rule is stricter: boilerplate only, full stop. No suggestions for internal business logic at all. Cuts down the noise.
That's super interesting about the custom Helm charts. We've been piloting it with a small support tools team and hit the same wall with our internal ticketing system's API client.
I've got a super basic question, maybe others too: how do you even track "acceptance rates" for suggestions? Is that a built-in Copilot metric or do you need separate tooling?
Also, we haven't tried the snippet library idea for common patterns yet. Did you find it worked better than just improving the main code docs?
Ask me in a year
Your point about Terraform modules sticking is interesting. We saw the same pattern, but it broke down when we got to more complex, state-dependent logic within our modules, like intricate depends_on chains or for_each with local maps.
For onboarding, we added a short section to our internal "Platform 101" doc specifically about when not to trust Copilot. It's essentially a decision tree: if it's a standard AWS resource block, green light. If it's referencing one of our internal modules or a custom variable structure, red light.
Our biggest adaptation was training the team to use the "ignore this suggestion" shortcut aggressively. The friction dropped significantly once developers learned to dismiss bad patterns immediately instead of letting them linger in the editor.
Your bill is too high.
That decision tree approach is brilliant, and it mirrors our internal guidance almost exactly. The Terraform state-dependent logic is a perfect example of where these tools fall apart. They're pattern matchers, not dependency solvers.
We found the same with `for_each` constructs referencing dynamic data sources. Copilot would confidently suggest a map structure that looked correct but would fail during `terraform plan` because it couldn't infer the data flow. It created a subtle false confidence.
Training developers on the "ignore suggestion" shortcut was our biggest win too, but we paired it with a quick team rule: if you have to ignore the same pattern three times in a session, add a comment to the relevant module or file. It turns the friction into a documentation signal.
The Ansible playbook success makes sense. It's great for templatized YAML. The friction with custom Helm charts is universal; it's trying to infer patterns from a mess of values files.
You need to feed it the right context. We forced a pattern: every service's Helm chart must include a `_helpers.tpl` with a standard comment block defining our custom labels and selectors. We then pointed Copilot's context to that file first. The bad suggestions for ingress and probes dropped by about 70%.
For the older APIs, we didn't have time for docstrings either. We wrote a small script that ran as a pre-commit hook, adding `// COGNITIVE_LOAD: High - see internal/pkg/v2` above any function in our deprecated directories. It's a brute-force signal for the tool to avoid that area.
shift left or go home
Pointing Copilot's context to a standard `_helpers.tpl` file first is a clever workaround. Did you have to guide developers on how to set that up in their IDE, or was it more of a team-wide config change?
I'm curious about the `// COGNITIVE_LOAD` comment. Does that actually work as a signal for Copilot, or is it mainly for the human reading the code? I thought it only read docstrings and code structure.
Oh wow, thanks for sharing this. I'm on a much smaller team just starting to talk about Copilot, and your breakdown of what stuck versus what broke is exactly what I need to bring to my lead.
The part about Ansible playbooks is really interesting - I wouldn't have guessed that. Does your team find it helpful for the more complex tasks, like handlers or weird conditional logic, or is it mostly just for the basic play structure?
Also, from an infra angle, did you run into any weird network or latency issues with 100 people hitting it at once? That's one of our team's big worries.
The parallels to our rollout are striking, especially your observation about boilerplate sticking. That pattern holds true across most compliance-focused code as well; we saw high acceptance rates for generating standard audit log structures or compliance check stubs.
Your question about adapting onboarding is key. We found the most effective method was to integrate specific Copilot guidance directly into our existing security and code review checklists. For instance, under the "Infrastructure as Code" section, we added a bullet: "For generated Terraform/AWS blocks, verify resource arguments against current provider documentation before commit." This frames the tool as a starting point that requires validation, aligning with our existing control mindset.
Regarding your friction points, undocumented internal APIs are a major risk vector. We instituted a lightweight process: any API method flagged as problematic during Copilot rollout was added to a quarterly tech debt review for proper documentation or deprecation. It turned the AI's confusion into a prioritization signal for the platform team. Did you experience any pushback when trying to formalize those gaps?
—at
The snippet library is a decent stopgap, but it's treating the symptom. The root problem is you're feeding it a custom, undocumented mess. That library will drift over time and then you've got two sources of truth.
Your "DEPRECATED" comment trick is worse. That's a style guide band-aid, not an actual deprecation. You're trusting a comment to do a linker's job. If the function is still exported and callable, it's not deprecated, it's just confusing. Either remove it or wrap it in a proper deprecation warning that throws at runtime.
— geo
> That's a style guide band-aid, not an actual deprecation.
You're right, but you're also describing every legacy codebase I've ever seen. Formal deprecation cycles are a luxury for teams with enough runway to pay down tech debt. Most of us are just trying to stop the tool from making the mess worse.
The comment trick isn't for the linker, it's for the pattern-matching machine. If adding "COGNITIVE_LOAD" cuts the bad suggestions by half, I'll take the band-aid.
CRM is a means, not an end.
That's a smart idea to pair the shortcut with a documentation rule. I like the concept of turning annoyance into a signal.
How do you manage that in practice, though? Is the comment just a simple "Copilot struggles here" note, or is there a more structured format your team uses to capture *why* the pattern fails, like mentioning the specific dynamic data source?
Also, have you found any difference in how this works between, say, Visual Studio Code and Neovim with the Copilot plugin? I'm curious if the suggestion mechanics are consistent enough for that three-ignore rule to be portable across editors.
Your point about Helm charts really resonates. We had the same struggle with our custom provisioning API, which has a lot of nested JSON config. Copilot would guess the structure wrong half the time.
We ended up creating a small, well-documented "snippet library" of example calls in a `/examples` directory and told everyone to keep it open as a separate tab. It gave Copilot the right patterns to copy, which cut down the bad suggestions a lot.
For the older APIs, we just slapped a `// DEPRECATED: use v2.client instead` above the old methods. It's not elegant, but it seemed to steer the suggestions toward the newer endpoints. Did you try anything similar, or was the lack of docstrings the main blocker?
Webhooks or bust.