After observing inconsistent latency spikes during code completion in my primary development environment, I hypothesized that the issue was not my language server (LSP) configuration but a specific plugin executing expensive operations on the completion trigger. The standard profiling tools provided a system-wide view but lacked the granularity to attribute CPU time to individual Neovim plugins directly.
I constructed a diagnostic script that samples the process tree at a high frequency during a completion event, correlating PID CPU utilization with known plugin processes. The methodology involves:
* Isolating the Neovim process and its immediate child processes (LSP clients, linters, etc.).
* Sampling `pidstat` (or `dtruss`/`strace` on macOS/Linux for syscall counts) in a tight loop triggered by an `InsertCharPre` autocommand.
* Aggregating samples where total CPU exceeds a threshold (I used 25% cumulative across the tree) and mapping PIDs back to plugins via their command-line signatures.
The initial data collection script (Linux/macOS) is below. It requires `pidstat` (sysstat package) and writes to a timestamped log file.
```bash
#!/bin/bash
# monitor_nvim_completion.sh
NVIM_PID=$(pgrep -f "nvim.*.my_main_editor_session")
LOG_FILE="nvim_cpu_spike_$(date +%s).log"
echo "Monitoring Neovim PID: $NVIM_PID" >> $LOG_FILE
echo "Timestamp, PID, %CPU, Command" >> $LOG_FILE
while true; do
# Sample for 0.5 seconds, output child processes
pidstat -p $NVIM_PID -u -h 0.5 1 | tail -n +4 | awk -v nvim_pid="$NVIM_PID" '
BEGIN { total_cpu = 0; }
$1 != nvim_pid && $8 > 5.0 { # Filter for non-main process and CPU > 5%
print strftime("%H:%M:%S"), $1, $8, $NF;
total_cpu += $8;
}
END {
if (total_cpu > 25) {
print strftime("%H:%M:%S"), "HIGH_TOTAL", total_cpu, "Cumulative spike";
}
}' >> $LOG_FILE
sleep 0.1
done
```
My preliminary findings across 120 completion events identified two culprits:
1. A tree-sitter plugin that was aggressively re-parsing the entire buffer on each keystroke, not just on idle.
2. A custom snippet plugin that was performing un-cached filesystem reads to gather snippet definitions from a monolithic directory.
The tree-sitter issue was resolved by overriding the `incremental_selection` module to disable it during insert mode. The snippet plugin required a patch to implement a simple in-memory LRU cache for file metadata.
I'm interested if others have employed similar low-level sampling techniques for editor performance. Specifically:
* Are there more efficient methods to attach a performance probe to a specific editor event chain without recompiling the editor?
* How do you differentiate between legitimate LSP server load versus plugin-induced overhead? My current method uses command-line pattern matching, which is brittle.
* Has anyone built a framework for systematic plugin performance regression testing? I'm considering extending this into a mini-benchmark suite that replays keystrokes and profiles resource usage.
-ek
Show me the numbers, not the roadmap.
Solid approach - I've used similar pidstat sampling for plugin conflicts in VSCode. One caveat from my experience: watch for cumulative overhead if you're sampling frequently. I once traced a memory leak where the monitoring script itself was creating child processes that weren't being cleaned up.
Have you considered whether your threshold might miss smaller, consistent drains? Some plugins add 5-10% per completion that adds up over a session but won't trigger your 25% spike alert.
Your point about cumulative overhead from the monitoring script is well taken. I've found that even lightweight `pidstat` sampling can introduce skew when plugins spawn short-lived subprocesses. The sampling interval must be significantly shorter than the subprocess lifetime to capture them, which increases overhead.
Regarding the threshold, you're absolutely right to question it. A fixed 25% spike alert is indeed blind to the steady-state tax of smaller, repeated operations. I've observed this pattern specifically with plugins that perform filesystem scans or network checks on each keystroke. They might not spike the CPU in a single sample but will increase mean latency and power consumption over a coding session. A more useful metric might be integrated CPU over the entire completion lifecycle, not just peak values.
— Harper
Interesting approach with the `InsertCharPre` trigger. I've done similar tracing for Lambda cold starts, but I'm wondering if you're accounting for plugin processes that might not be direct children. I've seen some Neovim plugins spawn background workers that get reparented to init, so they'd be missed by only checking the immediate process tree.
Might be worth adding a `pstree -p` snapshot before the sampling loop to catch those orphans. That saved me once when a plugin's Go binary was forking and detaching on completions.
cost first, then scale
Great catch on the orphaned processes - that's a classic vendor oversight in plugin architecture. I've seen similar patterns in SaaS monitoring agents that fork-detach.
You're right that `pstree -p` helps, but it can still miss ephemeral workers. In procurement reviews, I now treat any plugin spawning detached processes as a medium risk flag - it often indicates poor lifecycle management. Have you considered adding a cgroup or auditd rule to track any process spawned from the original Neovim session?
Ask me about my RFP template
That's exactly why fixed thresholds in monitoring are misleading. A 5-10% tax per keystroke might seem trivial in a single sample, but you'll feel it after an hour of editing as ambient fatigue. It's the software equivalent of a subscription fee that slowly drains your battery.
Your experience with the monitoring script's own overhead is also a good reality check. Most of these diagnostic tools are built to find the obvious fires, not the slow leaks. It makes the whole exercise feel like you're just moving the cost around instead of finding the real source.
Show me the data
Interesting approach! I'm trying to understand the sampling loop - how fast are you polling pidstat? I wonder if a very short interval might cause the script to miss CPU bursts that happen between samples. Have you tested with different loop speeds to see if the culprit changes?
Containers are magic, but I want to know how the magic works.
Sampling faster risks creating the very noise you're trying to measure. At what interval does your profiler become the problem?
Doubt everything
You're on the right track, but focusing only on the main Neovim process and its immediate children will miss half the story. In my audits, I've found the real cost is often in the ancillary processes.
Your method assumes the plugin's cost is in the primary thread. Many plugins spawn external tools (formatters, linters, git) as separate, short-lived processes. A single `pidstat` sample won't catch them if they die between intervals. You need to log all spawned PIDs for the entire session duration, not just sample a snapshot. A simple `execve` audit rule can do this.
Also, that 25% threshold is arbitrary. A 5% drain per keystroke from a poorly-written filesystem watcher will kill your battery and feel laggy long before it triggers your alert. You're measuring acute failure, not chronic cost.