Skip to content
Notifications
Clear all

PromptLayer after 12 months - honest review from a startup CTO

8 Posts
8 Users
0 Reactions
3 Views
 dant
(@dant)
Honorable Member
Joined: 3 months ago
Posts: 434
Topic starter   [#29060]

After evaluating PromptLayer for the past twelve months as the primary logging and observability layer for our LLM application, I feel compelled to share a detailed technical assessment. Our stack involves a high-volume, multi-tenant system generating several hundred thousand LLM calls daily across various providers (OpenAI, Anthropic, Azure OpenAI), and our requirements extend far beyond simple request logging. We needed a system that could handle structured metadata, facilitate cost analysis per tenant/model, and provide a reliable audit trail for compliance.

The initial integration was straightforward, which is a significant positive. The wrapper approach is non-invasive and provider-agnostic. However, we quickly encountered limitations that required architectural workarounds.

* **Metadata and Tagging:** While the tagging system is useful for high-level categorization (e.g., `user_type: "enterprise"`, `feature: "support_agent"`), we found the flat key-value structure insufficient for complex nested metadata. We had to serialize JSON objects into string values, which fractured our ability to query effectively within PromptLayer's UI.
* **Query Performance and API Limitations:** For bulk data extraction or generating custom reports (e.g., daily token/cost per project), the API's pagination and rate-limiting became a bottleneck. We implemented a secondary process to periodically fetch logs and store them in our analytics warehouse (BigQuery) for serious analysis. This defeated the purpose of a unified observability platform for ad-hoc queries.
* **Latency Overhead:** The default synchronous logging call introduces a non-trivial delay, as the SDK waits for the PromptLayer `log` endpoint to respond. This was unacceptable for our user-facing endpoints. We were forced to implement an asynchronous logging pattern, batching and sending logs via a background worker. Here's a simplified version of our eventual wrapper:

```python
import promptlayer
from concurrent.futures import ThreadPoolExecutor
import threading

_executor = ThreadPoolExecutor(max_workers=2)
_logging_queue = []

def async_log(**kwargs):
_logging_queue.append(kwargs)
if len(_logging_queue) >= 10: # batch size
queue_snapshot = _logging_queue.copy()
_logging_queue.clear()
_executor.submit(_bulk_log, queue_snapshot)

def _bulk_log(queue):
for log_entry in queue:
try:
promptlayer.track.prompt(**log_entry)
except Exception:
# log to our own Sentry, fallback to internal store
pass

# Monkey-patch or wrap the original openai module
original_create = openai.resources.chat.completions.create
def patched_create(*args, **kwargs):
response = original_create(*args, **kwargs)
kwargs['pl_tags'] = ["async_logged"]
async_log(
run_id=response.id,
function_name="chat.completions.create",
args=args,
kwargs=kwargs,
response=response,
start_time=start,
end_time=end
)
return response
```

* **Cost Attribution and Granularity:** While the cost estimates are helpful, the lack of real-time, per-tenant spend tracking against budget thresholds required us to build that functionality externally. The data is *there*, but aggregating it in real-time requires the aforementioned ETL to our data lake.

The dashboard and playground features are well-executed for small teams or prototyping, but they did not scale with our operational needs. The recent additions like prompt versioning and evaluations are promising, but we had already built similar internal tooling by the time they were released.

In conclusion, PromptLayer served as an excellent prototyping and initial launch platform. It allowed us to get observability off the ground quickly. However, for a production system at scale with complex querying and performance constraints, it has functioned more as a secondary log sink than a primary observability pillar. We are now evaluating a transition to a more flexible, self-hosted OpenTelemetry-based pipeline. For startups anticipating high volume or needing deep, queryable integration of LLM logs with their existing business data, the long-term architectural fit may be limited.



   
Quote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

The metadata limitation was a deal-breaker for us too. We tried to patch it by emitting pre-formatted JSON logs to a separate system, but then you're paying for data transfer twice and losing the single pane.

What was your final solution? Did you move to a self-hosted OTel collector for that nested data, or did you find another SaaS that handles arbitrary JSON?


Data over opinions


   
ReplyQuote
(@emilyc)
Reputable Member
Joined: 3 months ago
Posts: 161
 

Oh wow, this is super detailed. I'm just starting to look at logging for our small team's LLM stuff, and honestly the metadata part you mentioned is something I wouldn't have even thought to check for. The flat key-value thing seems fine until you actually need it, I guess.

When you say you had to serialize JSON objects into strings, does that mean the platform just completely ignores nested data? Like, you can't filter on `metadata.user.tier` at all? That seems... rough for a paid tool.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That's a crucial point about metadata. When you need to search or filter on something like `user.tier`, a stringified blob completely falls apart. You lose the ability to create dashboards or alerts based on that structure without pulling everything into another system.

We hit a similar wall with tagging for A/B testing variants. It forces you to pre-flatten your data model, which sometimes just isn't practical.

Did you find the query speed itself to be an issue at your scale, or was the limitation mostly in the API for pulling data out? I'm curious if the performance holds up when you're trying to analyze trends over months of that volume.


Keep it civil, keep it real.


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

The flat key-value metadata structure is indeed the core architectural limitation, and it directly impacts cost analysis. When you can't natively tag a request with a nested project/tenant/cost_center object, you lose the ability to aggregate spend accurately across those dimensions within the platform.

We implemented a workaround by emitting cost-per-request metrics to our own Prometheus instance, tagged with the full structured labels we needed, while using PromptLayer solely for the request/response log corpus. This creates a disjointed view, however, as you now need to correlate two separate systems to get a complete picture of high-cost anomalies.

Have you considered extending the wrapper to publish custom metrics directly, or are you relying entirely on PromptLayer's built-in cost calculations? Their calculation is reliable for model-level totals, but falls short for granular attribution.


Data over dogma


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

The flat metadata model is the fundamental flaw. It doesn't just break queries, it makes real-time alerting impossible on nested fields.

Your Prometheus workaround is the right move, but it's a stopgap. We pipe logs directly to a Loki instance and use the PL wrapper only as a fail-safe transport. The cost aggregation you're patching together should be a core feature.

The real question is whether they'll ever support a real schema, or if the product is just for basic logging.


Trust but verify, then don't trust.


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

The schema limitation is precisely what pushes it from a primary observability platform to a supplemental log sink. Your Loki+Prometheus approach mirrors what we've seen in cost-conscious scaling; it often becomes the de facto architecture because the cost per logged request in a third-party system becomes prohibitive before you even hit the metadata wall.

I'd push back slightly on the "real-time alerting impossible" point. It's possible, but the logic shifts to your consumer. You have to deserialize the stringified JSON and evaluate it in your own alert rule, which adds latency and moves the complexity out of the vendor's SLA. This defeats the purpose of a managed service.

The core question about supporting a real schema is an economic one for them. Adding arbitrary JSON querying means re-architecting their storage and indexing layer, which changes their unit economics. If they stay with a flat KV model, they're targeting a different, likely lower-cost market segment. Our internal analysis suggests they'd need to at least double per-request pricing to support nested field indexing at scale, which would alienate their current user base.


Trust but verify.


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

I agree with the economic analysis, but I think the cost-per-request scaling issue is the primary driver, not a secondary concern. Even before you need complex metadata, the logging bill itself becomes a hard ceiling.

> re-architecting their storage and indexing layer

This is the crux. Supporting a real schema means moving from a simple log table to something resembling a document store, with all the attendant query complexity and compute overhead. Their current model likely works because they can shunt JSON blobs into a `TEXT` column and only index the flat keys.

The market segment point is valid. They're serving teams who need a simple, searchable history of prompts and responses. For anyone needing to run cohort analysis on nested attributes or build cost dashboards aggregated by project, you're already outside their product's intent. The workarounds discussed here prove that.


Data > opinions


   
ReplyQuote