I've been evaluating various AI coding assistants for internal knowledge base querying, specifically for our DevOps and platform engineering teams. Our internal support documentation spans over 1000 markdown files covering CI/CD pipeline configurations, cloud cost anomaly procedures, and infrastructure as code templates.
I used Perplexity's "Pro Search" with a custom-built ingestion pipeline to process the documents. The goal was to assess its accuracy in retrieving and synthesizing specific, technical instructions. Here's a breakdown of the methodology and results.
**Methodology:**
* **Dataset:** 1000+ `.md` files from our internal wiki (Confluence export).
* **Query Set:** 50 pre-defined questions based on real support tickets (e.g., "How do I remediate AWS RDS cost spikes reported in CloudWatch?", "What is the approval workflow for deploying a new Terraform module to production?").
* **Accuracy Metric:** A correct answer required:
* Retrieving the correct document snippet.
* Synthesizing steps from multiple docs when necessary.
* Providing exact code block or configuration parameter when asked.
**Key Findings:**
* **Overall Accuracy:** 78% on first-answer retrieval (39/50 queries fully correct).
* **Strengths:**
* Excellent at finding specific error codes and matching them to documented runbooks.
* Code block extraction from documents was nearly flawless. When a doc contained a `terraform` or `github-actions.yaml` example, it was cited verbatim.
* Performed well on cost optimization queries, likely due to the structured nature of the data (e.g., "steps to resize Azure Kubernetes Service node pools").
* **Weaknesses:**
* Struggled with queries requiring synthesis of *contradictory* procedures (e.g., old vs. new deployment methods). It would often cite both without clarifying which was current.
* Hallucinated specific file paths or repository names that were similar but incorrect (e.g., `infra/terraform/modules/networking` vs. `infra/terraform/networking/modules`).
* For vague ticket descriptions ("The pipeline is broken"), it failed to ask sufficient clarifying questions and would retrieve a generic troubleshooting guide.
**Sample Query & Output:**
**Query:** "What is the command to purge the build cache for the legacy Jenkins agents, and which team manages the playbook?"
```bash
# Perplexity's cited answer (correct):
ssh jenkins-legacy-master-01
sudo /opt/scripts/clean_jenkins_cache.sh --all --force
# Note: The playbook is maintained by the Platform Reliability team. See docs/infra/playbooks/jenkins-maintenance.md
```
The citation directly matched the source document. In incorrect cases, the command would be slightly off (e.g., a wrong flag or hostname).
**Conclusion:** For highly structured internal technical documentation, Perplexity performs reliably for fact retrieval. Its propensity to synthesize can be a liability when documentation is outdated or contradictory. For our use case, it's suitable as a first-line support assistant, but requires human verification for any procedural changes or architecture decisions. I would not recommend it for autonomous execution of retrieved commands without a review step.
Numbers don't lie
Interesting approach! I always wondered how these AI tools would handle internal docs, especially when they get really specific like "approval workflow for deploying a new Terraform module." That's the kind of detail I'd need.
78% on the first try seems promising, but I'm curious how you handled the other 22%. Were the errors mostly about missing the latest doc updates, or did Perplexity start hallucinating steps for your internal processes?