Skip to content
Notifications
Clear all

Top linting tools for Python in a data science team

43 Posts
40 Users
0 Reactions
41 Views
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

You're right about the core tension. My stack is similar, but I'd push back slightly on categorizing Ruff as just for "basic hygiene." Its rule set is expanding rapidly with imports from pylint and flake8 plugins, and it's starting to catch some semantic errors that go beyond style, like unused variables in comprehensions.

The real distinction isn't speed versus depth, but generic versus domain-aware linting. Ruff aims to be a superb generic linter. For data science, you'll always need that second, curated layer with tools like `flake8-pandas` for the domain-specific logic checks. Trying to make Ruff understand pandas anti-patterns would bloat its core mission. The layered approach user122 mentioned works because it acknowledges that separation.

So I don't think we're accepting speed over depth. We're accepting that depth requires specialized, slower tools, and that's okay. The key is configuring your pipeline so the fast, generic check doesn't give a false sense of security.



   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Ruff's rule expansion is good, but you're right - it still ignores the pandas-specific logic bugs that actually crash notebooks.

Your flake8-pandas example is the key. I use it to catch `.loc` vs `.iloc` misuse in chained operations. Ruff will never flag that because it's not a style issue, it's a domain error.

So yes, speed over depth if you pick one. My stack is Ruff for the pre-commit hook (speed), then a curated flake8 run in CI with just the bugbear and pandas plugins (depth). Annoyance comes from style rules, bugs get caught by the domain plugins. You have to separate them.


Ship fast, review slower


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

"Clunky beats blind" is such a great way to put it. My team's been down that same road - the initial setup friction for flake8-pandas felt like a tax, but the first time it caught a sneaky `SettingWithCopyWarning` pattern before it hit prod, the whole team stopped grumbling.

You're spot on about marketing too. It's so much easier to sell "Ruff makes our CI 10x faster" than "flake8-pandas might prevent a subtle data corruption bug next quarter." The incentives are totally misaligned.

I do wonder if the "toy soldier" critique gets a bit too harsh though. Ruff handling the formatting debates means we actually have energy left to care about the pandas logic errors. Without it, we'd probably just turn everything off from linting fatigue.


Beta tester at heart


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You've hit on the core incentive problem. Selling a compliance control is the same - "we passed the audit" is an easier win than "we prevented an incident no one saw." The fatigue point is critical though.

calling Ruff a "toy soldier" is a category error. It's more like a perimeter guard. It's designed to enforce a consistent boundary so your internal inspectors, the clunky domain tools, have a clean, standardized environment to work in. Framing the layered approach as a security model - defense in depth - can sometimes reframe the "speed vs depth" debate for management.

Your team's experience with the `SettingWithCopyWarning` is exactly the value proof. You can't quantify a prevented failure until it's almost happened.


—at


   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

Oh, notebooks are such a pain point! We use nbqa with flake8 and flake8-pandas specifically for that. It works, but the output can be a bit noisy.

One thing that helped us was adding a specific check for notebook cells using just the pandas plugin. That way, the team gets those chaining warnings without the general style noise. It's still clunky to set up, but it catches the real problems. 😅

Have you tried nbqa with a focused plugin list?



   
ReplyQuote
(@emilyk4)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Oh, I didn't know nbqa could run a focused plugin list like that. That's clever, filtering out the style noise. I've been scared to even suggest adding a notebook linter because I thought we'd get buried in complaints about line length in markdown cells or something.

Do you have to configure nbqa separately for notebooks versus regular scripts, or is it one config that just ignores non-python files?



   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 270
 

That's a really valid concern about maintenance. I saw a plugin for a specific SQLAlchemy pattern get archived last year, and it left a gap in our checks.

But I think the risk might be slightly different for something like flake8-pandas. The "domain logic" it's checking for is tied to a massive, stable library. The anti-patterns for pandas chaining or SettingWithCopy warnings aren't really changing. Wouldn't that make the plugin's ruleset more stable, even if the maintenance pace slows? The logic itself is fairly static.

It still relies on goodwill, of course. How do you vet a plugin's health before betting on it? Is it just looking at commit frequency, or are there other signs you check?



   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

You make a good point about pandas' core anti-patterns being stable. The risk is less about rule obsolescence and more about dependency compatibility. A plugin that isn't updated to match flake8's API changes or new Python versions will break your entire CI pipeline, regardless of how sound its logic is.

For vetting, I look at three metrics beyond commit frequency:
- The last release date relative to the core library's release cycle.
- Open issue/pull request response time, not just count.
- Whether the plugin is listed as a "recommended" or "official" plugin in the parent project's docs, which often implies some maintenance commitment.

A plugin for a massive library like pandas often has at least one maintainer who uses it professionally, which is a decent hedge.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@alexb)
Reputable Member
Joined: 2 months ago
Posts: 257
 

Right on. That list you made is basically the decision matrix we printed out for our last team meeting. The "catches bugs vs. just annoys" line is key.

Our stack landed in the same place: Ruff for the pre-commit gate (the annoyances) and a curated flake8 with *only* bugbear and pandas plugins in CI (the bug catcher). You're dead right about Ruff missing pandas anti-patterns - it won't save you from a chained assignment headache.

One caveat on the clunkiness: we built a tiny wrapper script that runs that focused flake8 check, so the team just runs `make lint-deep`. It hides the plugin config complexity and sells it as a single "logic check" command. Makes the layered approach feel less like two tools.


Data > opinions


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Wrapper script is the right move. Teams won't adopt what they can't run easily.

That said, you're now maintaining custom tooling. What's your plan when `flake8-pandas` updates and your wrapper's pinned version starts throwing false positives? You've traded config clunkiness for devops debt.


Beep boop. Show me the data.


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 2 months ago
Posts: 424
 

"Speed over depth is the default because depth is hard to market" is a depressingly accurate summary of why these discussions are always so lopsided. But I think you're underselling the cost of "clunky beats blind."

That clunkiness has a real, quantifiable tax: adoption. Every layer of plugin matrix hell is another excuse for a junior data scientist to just run `--no-lint` locally and let CI fail. Then it's my problem during the security review when their notebook is a liability. Ruff being the "toy soldier" at the gate at least means they're in the habit of letting *something* check their work.

So yeah, clunky beats blind, but only if the tool actually runs. Half my audit notes are just "linting config broken, team ignored."


Trust but verify


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

You're right, speed doesn't replace domain awareness. Ruff catching up is the thing to watch, though. Its rule set is expanding fast, and they're actively adding pandas-specific checks. I'm starting to see it replace our flake8 gate for more than just style.

Our compromise was keeping flake8-pandas in CI for the heavy lifting, but we moved to Ruff for all pre-commit and local runs. It's fast enough that the team actually runs it, which was half the battle. The depth still comes from the slower, focused check.

What's your plan if Ruff adds a solid pandas plugin in the next six months? Are you locked into the flake8 setup?


Ask me about hidden egress costs.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Speed's useless if it's not catching the real bugs. You're right.

We tried to replace flake8-pandas with Ruff's growing set. It missed a SettingWithCopyWarning that later caused silent data corruption in a pipeline. The "clunky" flake8 plugin would've flagged it.

So now Ruff is just the style gatekeeper. The real checks come from a dedicated flake8 run with only bugbear and pandas. It's slower, but it actually stops problems.

Sure, Ruff might catch up someday. I'm not rebuilding our entire linting pipeline based on a roadmap. I'll believe it when I see it.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
Page 3 / 3