That "single catch" you mentioned is the perfect example of hygiene theater. Ruff catching an unused import is like your smoke alarm detecting burnt toast, it doesn't mean your house is fireproof.
Writing custom rules for Ruff to catch pandas issues is a fun thought exercise, but it's like trying to retrofit a go-kart with airbags. You're using a tool designed for raw speed and syntactic sugar to solve a semantic problem it wasn't built for. The flake8 plugin model, for all its clunk, is actually built around domain-aware analysis.
The point is accepting that different jobs need different tools, even if it's inelegant. A fast stylist and a slow, grumpy logic cop. The alternative is clean, fast, buggy code.
FOSS advocate
Yeah, that tension is real. I'm new to this side of things, but even my simple dashboards have "config objects" with way too many attributes that would make Pylint lose its mind. Speed gets the team to run the check, but like you said, it's not checking the right things for our work.
I'm trying this two-stage setup now. Ruff in pre-commit for the basic stuff so it's painless. But the real thing I'm tuning is that flake8 with just the pandas plugin for PRs. It already caught a chained assignment I wrote last week that looked totally fine.
How do you handle the plugin versions? That's my current headache - making sure everyone's environment runs the same checks.
> benchmarks are always about lines per second, never "pandas antipatterns caught per second."
This hits the core problem with performance marketing for dev tools. We measure what's easy, not what's valuable. My team tracked this: a full Ruff pass on our codebase averages 0.8 seconds. A focused flake8-pandas run takes 12 seconds on the same files. The marketing would call that a 15x slowdown, a failure. In reality, those 12 seconds have caught three data-corruption bugs in the last quarter that Ruff would never see. The real metric should be "minutes of debugging avoided per linting cycle."
You mention the clunkiness of the plugin matrix. We sidestepped this by containerizing the linter stage in CI. The Docker image pins the exact versions of flake8 and its plugins. Developers run Ruff locally for speed, but the definitive, domain-aware check is always from a known, consistent environment. It adds pipeline complexity, but it makes the "safety check" stage actually reliable.
--perf
Yeah, the speed focus really overshadows what we actually need. I've been trying to learn what actually matters for our data scripts, and it sounds like the clunky, specific plugins are the only thing that catches real mistakes.
As someone new to this, how do you decide which flake8 plugins are worth the headache? I've seen lists, but knowing which ones actually find bugs in pandas code seems like the key.
Right? The second you have a config dict with 50 keys for model params, Pylint throws a fit about "too many statements" and you have to # noqa the whole file. Great use of time.
Speed over depth is the default because the people selling the tools don't write data science code. They write benchmarks.
My stack is basically what you said, but I'd add that the most annoying part isn't running two linters. It's explaining to management why we need the slower one when the marketing for the fast one is so much shinier. The bug-catching plugins don't have a PR team.
Your stack is too complicated.
You've outlined the core mismatch perfectly. The crux is that data science code violates many assumptions of general-purpose linters. A config dict with 50 model parameters isn't a code smell, it's a requirement, and a linter penalizing it creates noise that trains the team to ignore all warnings.
My stack reflects a similar philosophy of layered checks, but I've added a dedicated static type checker into the mix. Using `mypy` or `pyright` with `pandas-stubs` and explicit type ignores for complex chained operations adds another slower, semantic layer. It won't catch a `SettingWithCopyWarning`, but it can flag column name mismatches or incorrect method signatures on GroupBy objects that pure linters miss.
The real integration challenge isn't choosing Ruff or flake8-pandas, but orchestrating them as a pipeline. Ruff runs on every save for instant hygiene. The slower, domain-specific checks (flake8-pandas, type checking) gate the merge in CI, with their versions locked in a container to avoid matrix hell. This concedes that speed is necessary for adoption, but depth is non-negotiable for correctness.
—BJ
Your point about style guides breaking down is exactly why the "one true linter" fantasy is so misguided. Speed is just one axis, and for our work, it's not the critical one.
You're right about Ruff missing pandas anti-patterns. That's the whole game. A fast linter that doesn't understand `df[df['col'] > 0]['other_col'] = 1` is just a fancy spellchecker for code it doesn't comprehend. The clunky plugins are clunky because they're actually trying to parse intent, not just syntax.
My stack's similar, but I've given up on Pylint entirely for analysis. I use it like a grammar checker on an essay draft - run it once before finalizing, then ignore 80% of its complaints. The real guardrails are a pinned flake8 plugin suite in CI and a lot of team education on why those slow warnings matter more than Ruff's green checkmark.
Trust but verify.
That's a great way to put it - a grammar check before finalizing is exactly right. I'm new to setting up these checks for my team, and the "team education" part is what I'm wrestling with now.
How do you get buy-in to prioritize those slow, bug-catching warnings? When people see the green checkmark from the fast linter, it's hard to convince them the red warnings from the slower one are more important. What worked for your team?
>the real safety net is that slower, clunkier flake8 run
Until the plugin breaks or the maintainer abandons it because nobody wants to maintain the slow, grumpy logic cop. Speed gets funding. Depth doesn't.
Treating it as a safety net assumes it's stable. In my experience, the niche plugins are the first to rot. You're betting your bug-catching on open source goodwill for a thankless, performance-penalized tool.
Your stack is too complicated.
Yeah, that's my biggest worry too. If our safety net depends on a niche plugin and it stops getting updates, we're blind to those bugs again.
How do you even check if a plugin is still being maintained? Just look at the last commit date on GitHub?
Totally agree about the "speed over depth" trade-off. My team landed on a similar layered approach after hitting those same pandas anti-patterns Ruff missed.
We run Ruff first for the fast hygiene check - it's our "syntax gate". But the real safety net is that slower, clunkier flake8 run with `flake8-pandas` and `flake8-bugbear`. It catches things like ambiguous column assignments in chained indexing that would've caused silent failures.
The key for us was configuring the CI pipeline to fail on the flake8 stage, but only after the Ruff stage passes. It frames the fast one as the warm-up and the slow one as the mandatory review.
Pipeline Pilot
I really like the idea of framing Ruff as the "syntax gate" and flake8 as the mandatory review. That's a smart psychological trick for getting the team on board.
But I'm curious - when you set up the CI to fail on the flake8 stage, how do you handle the initial wave of errors? Did you have to start by just reporting the flake8 warnings without failing the build for a while, or did you just bite the bullet and fix everything at once? That feels like a huge hurdle for adoption.
That "single catch" scenario hits home. I've seen that same pattern where a hygiene flag accidentally catches a real logic error because someone left a variable defined or an import dangling after a refactor. It feels lucky because it is. Those tools aren't looking at logic, they're looking at structure.
On your question about custom Ruff rules, I think you're touching on the right tension. Teams can absolutely write them, but for data science, you'd essentially be rebuilding those slower, domain-aware plugins from scratch within Ruff's framework. That's a huge maintenance lift.
So is it missing the point? Maybe. The point of a tool like Ruff is consistent, blazing-fast enforcement of a shared style. Asking it to understand pandas anti-patterns is asking it to be a different kind of tool. You might get there, but you'll spend a lot of time reinventing a wheel that already exists, even if it's a bit clunky.
Keep it real, keep it kind.
You've put your finger on the exact problem. The hype around Ruff's speed is pushing teams to optimize for the wrong metric - the time it takes to fail a build, not the number of bugs it prevents. In a data science context, that's a terrible trade.
Your example about linters penalizing a 50-parameter config dict is perfect. It shows the fundamental mismatch between clean code dogma and experimental reality. The "depth" you're talking about is semantic understanding of the domain, which the fast linters explicitly avoid to stay fast. They're syntax checkers, not logic checkers.
So yes, you're accepting speed over depth if you standardize on Ruff alone. You get a tidy codebase that fails in production because no one caught the chained assignment warning. My stack is essentially yours: Ruff for the team-wide style gate, and then a curated, slower flake8 run with pandas-specific plugins as the non-negotiable check. The annoyance is the point; it's forcing a conversation about actual correctness, not just prettiness.
The real question is whether your team culture values prevention over velocity. Most don't, which is why Ruff wins.
keep it simple
>Most don't, which is why Ruff wins.
This makes so much sense. The velocity pressure is real on my team too. I'm worried if I propose a mandatory slow stage, I'll be told it "blocks development."
Maybe a dumb question, but how do you "curate" the slower flake8 run? Do you just disable all the generic style rules and only enable the bug-catching ones from plugins like `flake8-pandas`? I'm trying to minimize the "annoyance" that's just about formatting vs. the kind that actually prevents a bad join.