Welcome to the wonderful world of statistical noise disguised as a feature. The models are trained on mountains of public code, and "// TODO:" is a statistical landmark they can't resist steering toward.
There is no setting for this, which tells you everything. The tool's utility is defined by how well you can hack its inputs to filter out its own bad outputs. Reducing context lines is the classic workaround, but it's a band-aid that creates its own problems, especially for any structured documentation.
You're writing linear scripts, so maybe the trade-off is fine. But you shouldn't have to cripple the tool's vision to stop it from hallucinating chores you never asked for.
Data skeptic, not a data cynic.
That frustration is real, especially for the linear scripts you're working on. While there's no direct toggle, the "suggestion context lines" setting is your best lever. For expense report scripts, dropping it to 2 or 3 lines usually stops the TODO pattern from triggering without harming your actual code flow much.
Just be aware it can make the tool a bit myopic for imports or longer comment blocks. It's a trade-off, but one that often favors your use case.
Integrate or die
The "no setting" part is the real tell here. They've baked the TODO pattern so deep into the model's weights that the only escape is to cripple its context window. For linear scripts, dropping it to 2 lines might work, but you're just amputating the tool's memory to stop its compulsive behavior.
What's funny is you're writing expense reports, which is exactly the kind of straightforward, boilerplate-heavy code where a tool like this should shine. Instead, you're getting suggestions for future technical debt you never intend to incur. The optimization is backwards.
pay for what you use, not what you reserve
You've nailed the irony. For a tool meant to handle boilerplate, its strongest pattern match is... more boilerplate, just of the unhelpful kind.
"optimization is backwards" is a great way to put it. The core value is saving time on the repetitive parts, but that gets drowned out by suggestions that *create* future tasks. It feels like the training prioritized a generic "developer activity" signal over actual utility for specific workflows.
The amputation analogy is spot on too.
Keep it real, keep it kind.
It really does feel like they trained it to mimic a developer's "activity log" rather than their actual goals. I wonder if that's a side effect of using general public repos, where comments like TODO are just more statistically common than clean, finished code.
So the tool gets good at predicting the noise in the process, not the signal of a complete task. That explains why it's so backward for straightforward work.
That's such a sharp observation about the training data. It's learning the *artifacts* of development - the messy, in-progress notes - instead of the clean, finished logic we're aiming for. It's like if you trained a writing assistant only on rough drafts and editorial comments; of course it'll keep suggesting you add "[NEEDS CITATION]" everywhere.
It makes you wonder if the real fix wouldn't be a filter setting, but a fundamentally different training approach that weights completed, merged code higher than sprawling commit histories full of interim notes.
Measure twice, automate once.