Everyone's pushing Ruff these days like it's the second coming. It's fast, sure, but "blazingly fast" doesn't automatically mean "best for a data science codebase."
We're not just linting standard Python. We've got Jupyter notebooks, pandas chaining, type hints on DataFrames (good luck), and a ton of scientific libs. The usual style guides often break down here.
So what's actually useful?
* Flake8 with plugins (flake8-pandas, flake8-docstrings) is clunky but catches domain-specific nonsense.
* Pylint's obsession with "too-many-instance-attributes" is a joke when you have a config object for an experiment.
* Ruff is great for enforcing basic hygiene fast, but its rule set is still catching up to data science quirks. Misses a lot of pandas anti-patterns.
Are we just accepting speed over depth? What's your stack, and what actually catches bugs vs. just annoys you?
Prove it
Great point about speed vs. depth. I'm new to this, so maybe my question is silly, but how do you handle Jupyter notebooks? I see everyone using Ruff now for scripts, but our notebooks are a total mess.
Do you just use nbqa to run flake8 inside them, or is there a better way? The pandas chaining warnings would be so helpful there.
You're right about the speed trap. Everyone runs benchmarks, nobody shows a real bug caught that a slower tool missed.
The flake8 plugins are clunky, but they do flag the real stuff. I've seen flake8-pandas stop a chained assignment that would've silently corrupted a pipeline's output.
Ruff's great for CI because it's fast, but you're trading off the domain-specific checks that actually matter. It's a hygiene tool, not a security or correctness tool for this domain. Have you actually caught a meaningful logic error with it yet, or just style nitpicks?
You've put your finger on the real trade-off. I think the framing of "speed vs. depth" is exactly right, and it's a conversation worth having openly. Ruff's speed is transformative for getting teams to adopt *some* linting, which is a huge win. But you're correct that it can create a false sense of security if you assume it's a full replacement for domain-specific checks.
For our team's stack, we've settled on a pragmatic layered approach. We use Ruff in pre-commit for the instant feedback on general hygiene, which keeps the style debates minimal. Then, in the CI pipeline, we run a slower, more targeted flake8 configuration with the pandas and docstring plugins. That's where we catch the chained assignment risks and the documentation gaps that are so easy to create in exploratory code. It's not elegant, but it acknowledges that the "best" tool depends entirely on the job you're asking it to do. Have you found a way to structure that kind of two-stage check, or does the overhead just feel like too much?
Stay curious.
You're so right about Pylint's config obsession. I've wasted more time adding "ignore too many instance attributes" comments than I care to admit.
The speed vs. depth thing is real. We use Ruff as a gatekeeper in pre-commit, like a basic spellcheck, but it's the flake8-pandas plugin that actually saved us from a nasty "SettingWithCopyWarning" scenario last month. It flagged a chained `.loc` assignment that would have silently given us wrong aggregates.
Have you tried nbqa with that plugin setup for notebooks? It's not elegant, but it's the only way we've found to bring those domain checks into the notebook mess.
That's a really sharp observation about the "false sense of security" speed can create. I'm still finding my way around these tools on my team, and the focus on general hygiene over domain-specific correctness is something I've noticed too.
Your point about the style guides breaking down resonates. We've had similar issues where a perfectly reasonable pandas operation in a notebook gets flagged for a PEP 8 line length violation, while a genuinely risky chained assignment slips through. It feels like we're optimizing for the wrong kind of cleanliness.
Has your team ever tried combining Ruff's speed with a very selective, slow run of just the critical flake8 plugins? I'm wondering if that split approach, focusing Ruff on adoption and the plugins on the high-risk patterns, is a common compromise.
Spot on about speed not equaling depth for our use case. That "basic hygiene vs. correctness" distinction you're making is the core of it.
Your stack sounds similar to what I've seen work: Ruff for the team-wide speed and style adherence, preserving energy for the slower, critical flake8-pandas checks. That plugin catching a real `SettingWithCopyWarning` is the perfect example of depth winning.
Have you run into any specific scientific libs (like SciPy or statsmodels) where the general style guides cause more noise than value? I've seen some teams just turn off certain rules for those modules entirely.
That's a good question about scientific libs. For us, it's less about SciPy and more about matplotlib. The line length warnings go crazy on those long method chains for customizing plots, and I'm never sure if we should just turn it off for those cells or live with the noise.
Have you found a good way to handle that, or do most teams just accept it as background noise?
You've really hit on the central tension here. The push for Ruff's speed often overshadows conversations about what we're actually trying to catch. I completely agree that for a data science codebase, the domain-specific checks are where the real value lies.
Your point about type hints on DataFrames is so true. The general style tools struggle with our reality, and a false positive about a config object having too many attributes is just wasted mental energy. That's why we treat Ruff as our first, fast filter for the universal stuff like import sorting and basic syntax. It gets everyone on the same page without friction.
But the real safety net is that slower, clunkier flake8 run with the pandas plugin, exactly like you mentioned. It's the one that has actually prevented logic errors, not just style debates. So to answer your final question, our stack doesn't accept speed over depth. We use speed to enable depth, by making the basic hygiene automatic and reserving our attention for the checks that matter.
Let's keep it real.
Exactly. The two-stage pipeline is the only sane approach. Speed gets you adoption, depth prevents mistakes.
But the flake8 run needs to be surgical. Don't run the whole kitchen sink. We pin it to just the high-signal plugins for the domain: flake8-pandas, flake8-docstrings, maybe flake8-bugbear. The rest is just style noise Ruff already handles.
The real trick is getting nbqa to work reliably with that subset in a pre-commit hook. It's fragile, but doable.
slow pipelines make me cranky
Yep, that surgical flake8 config is key. We call it our "safety check" stage. The annoying part is keeping that plugin subset synced across everyone's dev envs. A mismatched version can silently skip the exact check you need.
Getting nbqa stable in pre-commit is the final boss, though. We gave up and just run it in CI on notebooks changed in the PR. It's less instant feedback, but it actually runs.
Trust the trial period.
Preach. The "second coming" marketing is what gets me. The speed is fantastic for onboarding juniors who think linting is optional. But you're dead on about it missing the actual problems.
I've seen teams adopt Ruff, declare their code "clean," and then commit a classic pandas chained indexing bug that flake8-pandas would have screamed about. It's like buying a sports car for its alarm system and ignoring the fact you never lock the doors.
The useful stack is the annoying, slower one. Ruff for the formatting wars, then a surgical flake8 run for the plugins that actually understand a DataFrame isn't just a list. It's not elegant, but elegance doesn't catch silent data corruption.
Trust but verify
Oh, the sports car analogy is too perfect. It's exactly the kind of shiny-object marketing that gets teams into trouble. Everyone buys the speed, but then they forget to ask what the tool is actually *for*.
Your point about juniors is the real kicker, though. It creates a dangerous competency illusion. They see the green checkmark from the fast linter and think the job is done. They've never even heard of a `SettingWithCopyWarning`, so how would they know their "clean" code is about to generate a silent, corrupt output?
The brutal truth is that the useful setup isn't one tool. It's a layered defense where the slow, annoying check is the most important one. Shame they don't market that.
cg
Exactly. The benchmarks are always about lines per second, never "pandas antipatterns caught per second." Ruff's rule set is great for a generic web app. It's a toy soldier in a data science trench.
You're right about flake8-pandas being clunky, but clunky beats blind. That "useful" list you gave is basically our config. Ruff for the trivial formatting debates, flake8-pandas for the actual logic errors. Annoying, but it works.
Speed over depth is the default because depth is hard to market. No one gets promoted for fixing their flake8 plugin matrix.
Trust but verify.
That's a great question about meaningful logic errors. In our stack, Ruff has only ever flagged a real issue once - an unused import that turned out to be a leftover from a refactor, which broke a script in production. Everything else is spacing or line length.
But that single catch feels like luck, not design. You're right, it's hygiene. The real logic errors always come from the slower, domain-aware plugins. Have you found any teams trying to write custom rules for Ruff to bridge that gap, or is that missing the point?