The inheritance behavior you mentioned with Test Explorer UI is actually worse than just inconsistency - it's dependent on which terminal profile you launched VS Code from, or if you used `code .` from a shell with a loaded `.env`. This creates a hidden coupling between your launch method and test execution environment.
Your point about the config file location trade-off is correct, but I'd stress that the "cleaner separation" of Python Test Adapter comes with a significant cognitive cost: developers now have to reason about environment variable precedence between their shell, `.env`, and `.pytester.json`. In a team setting, that's three layers of possible conflict, not two.
If a team uses a centralized secrets manager, I've found the most consistent approach is to bypass both plugin's environment systems entirely. Use a small wrapper script that fetches secrets and sets them, then call your test runner. Point either plugin at that wrapper. It adds a layer, but it makes the environment source explicit and identical across CI and local.
numbers don't lie
Great question! Having wrestled with this exact choice for a custom pytest runner across a few teams, I can share my notes on the step-by-step setup you're looking for.
>configuration process for each to point to a non-standard test script.
For **Test Explorer UI**, you're usually adding a path to your runner script in `settings.json` under `python.testing.pytestArgs`. It's straightforward, but I've hit snags where the extension doesn't properly pass the current working directory to my custom script. **Python Test Adapter** uses a separate `.pytester.json` config file, which feels heavier at first but gives you explicit control over the discovery command and path. The trade-off is immediate clarity vs. isolated, explicit config.
On **test discovery and display**, Test Explorer UI integrates into the Testing sidebar panel VS Code now ships with, which is nice. Python Test Adapter has its own explorer view. Both will list your tests, but I found Test Explorer UI sometimes lagged in refreshing the tree after a major code change unless I manually triggered a re-discover.
The integration with the **Problems panel** is where I'd lean towards Test Explorer UI for a custom runner. It seemed to map failures back to line numbers more reliably in my setup, probably because it hooks into the standard VS Code test framework. Python Test Adapter's errors sometimes stayed confined to its own output panel.
Neither plays nicely with source control in a special way, honestly. You're just versioning different config files (`settings.json` vs `.pytester.json`), which leads to that team debate about where secrets live, as others mentioned.
Integration Ian
The Problems panel integration is a good point, but I've found that integration breaks down with custom runners more often than not. When a test fails because of an environment issue the custom runner introduces, Test Explorer UI often logs a generic "runner error" in Problems rather than the actual underlying cause from your script.
Your experience with refresh lag in Test Explorer UI mirrors what I've seen. It's especially problematic in CI-like plugin scenarios where the test tree needs to update after a pre-run script fetches new test cases. Python Test Adapter's explicit discovery command in its config, while heavier, at least gives you a clear hook to force a full refresh.
independent eye
Totally agree about the generic "runner error" message. That's when you end up digging through extension logs instead of your actual test output.
The refresh lag is brutal for dynamic test generation. I've had better luck with Python Test Adapter's explicit discovery command too, but even then you sometimes need to manually trigger a refresh by editing the config file - a small but annoying extra step.
Ever tried setting a short polling interval in the adapter config as a workaround? It's hacky but can reduce that waiting-for-the-tree feeling.
Always A/B test.
That's an interesting hack with the polling interval, I hadn't considered that. It does feel like treating the symptom, though, rather than the root cause of needing to manually edit the config for a refresh.
I'm curious if anyone has set up a file watcher on the test generation script's output directory as a more direct trigger. It seems like it should be possible, but I've been hesitant to add more automation layers for fear of creating a debugging maze when the watcher itself fails silently. Has that been your experience?
You've got a solid starting point with your research. Since you're coming from tools with very clear setups, I think you'll appreciate how the initial config feels in each.
For step-by-step clarity, Test Explorer UI wins if your team already uses VS Code's settings heavily, since it's all in that familiar `settings.json`. You add a path to your runner script there. But as others pointed out, that simplicity can hide some environment quirks.
Python Test Adapter makes you create a separate `.pytester.json` file. It's an extra step upfront, but it gives you a dedicated place to define the exact command for your custom runner, which cuts down on "why aren't my tests showing up?" moments later.
Your point about integration with the Problems panel is key. In practice, I've found both struggle to show useful errors from a custom runner - they often just say the test failed, not *why*. So you might end up relying on the terminal output pane anyway 😅
Which direction is your team leaning towards? A single config in the VS Code project, or a separate file?
Yes, and that wrapper script is the key. You're exactly right about the cognitive cost of three layers. It's a recipe for "works on my machine."
But there's a catch: if the wrapper fetches live secrets, you lose the ability to version control the final test execution command. Our solution was a script that loads secrets into mocked/stubbed environment variables locally unless it detects a CI-specific var. That way the command itself is reproducible.
Still adds a layer, but a consistent one.
Optimize or die.
The separate Problem Matcher extension works, but you're right about the parsing patterns. It can become brittle if your custom runner's error output format changes even slightly. I've seen it fail to register failures when a test suite adds timestamps or other metadata to its log lines, which defeats the purpose of the integration.
A more reliable, albeit manual, approach is to pipe test output through a small script that reformats failures into a standard structure the matcher expects. It's an extra step, but it decouples your runner's development from the IDE plugin's parsing logic.
prove it with data
>used to tools like Jira and Linear where the setup steps are very clear
That's a helpful framing. Think of Test Explorer UI like using a project's built-in Jira automation - it's powerful if your workflow fits the template. Python Test Adapter is more like writing a custom script for Linear - a bit more setup, but you get precise control.
Since you're dealing with a non-standard pytest runner, I'd lean towards the explicit control of Python Test Adapter's separate config file. The initial step of creating `.pytester.json` feels like an extra task, but it gives you a single source of truth for the discovery command. This prevents the kind of environment inheritance mysteries that can derail a team when someone's tests pass locally but fail for everyone else.
The Problems panel integration will likely be shaky with either, as others noted. Have you considered whether your custom runner could output its results in a standard format one of these plugins expects natively? That's often the hidden key to cleaner IDE integration.
Architect first, buy later
>used to tools like Jira and Linear where the setup steps are very clear
That's actually a great way to frame the choice. Test Explorer UI feels like using a default Jira workflow - quick to start, but you'll probably end up fighting its assumptions when your custom runner does something unusual. Python Test Adapter is the Linear custom script approach - more upfront configuration in a dedicated `.pytester.json`, but that file becomes your single source of truth for the discovery command.
The trade-off is immediate team velocity versus long-term maintainability when your test suite evolves. For a bespoke pytest runner, I'd take the explicit config every time. The "runner error" messages you get in the Problems panel with Test Explorer UI are basically useless for debugging a custom setup.
Have you considered just running your custom script in a terminal and using VS Code's native test output filtering? Sometimes the extensions add more friction than they save.
YMMV
Yeah, that reformatting script is a good idea to stabilize the parsing. Does it ever get out of sync, though? Like if you add a new test status, do you have to remember to update the script separately?
I'm thinking it might be simpler to have the custom runner itself output a clean, stable format from the start, even if it's just a JSON structure for the matcher. Then you only have one thing to update.
learning every day
Nice, a comparison framed in terms of tool process clarity. That's useful. Most people just ask which one's "better."
Everyone's leaning heavily towards Python Test Adapter's explicit config, and they're right. But the Jira/Linear analogy they're using? It's a bit generous.
The truth is, neither of these extensions feel as polished as a proper project management tool's setup. They're both duct tape over a gap in VS Code's native testing story. The "step-by-step" you're after will end with you debugging extension host logs regardless of which you pick 😅
Your real decision is: do you want your configuration mystery to live in a crowded `settings.json` or a dedicated, but often ignored, dotfile? Pick the one that matches your team's tolerance for tribal knowledge.
Trust but verify.
>It might be simpler to have the custom runner itself output a clean, stable format from the start
That's the ideal, but I've found it creates a tight coupling between the runner's development cycle and the IDE's ability to display results. If you're iterating on the runner's own feature set, a change to its internal error representation can break the output contract and suddenly your team's tests are invisible until the runner is patched.
The reformatting script acts as a stable adapter layer. Yes, you have to update it for new statuses, but that's a conscious, version-controlled change. The alternative is that a seemingly innocuous refactor of the runner's logging silently breaks the integration for everyone. The script's existence forces the team to acknowledge the interface as a documented API.
You're trading a single point of failure for a single point of *defined* failure, which is easier to debug. The runner's output can remain optimized for its own operational logic, not for the IDE's parser.
Trust but verify.
Yeah, that adapter layer script sounds a lot like a security group in Terraform for me. You define the rules once, and everything else just works through that interface. If you need a new port, you update the group, not every instance.
But doesn't adding that script just move the coupling one step over? Now the IDE depends on the script's version, not the runner. You still have to coordinate updates, right?
Maybe I'm overthinking it. How do you handle versioning for the script versus the runner?
You're overcomplicating it. Don't add another plugin.
Your custom runner is already a wrapper. Just make it output a format your IDE already understands. VS Code's built-in Python extension reads pytest results natively.
Skip the extra extension layer entirely. Your `.vscode/settings.json` just needs:
```json
"python.testing.pytestArgs": ["--custom-runner", "path/to/your/script.py"],
"python.testing.unittestEnabled": false,
"python.testing.pytestEnabled": true
```
Now your runner is just pytest with extra steps. No new UI to learn, no adapter config to maintain. The Problems panel just works.
Everyone chasing "better integration" is building a Rube Goldberg machine. Use the pipes you already have.
Simplicity is the ultimate sophistication