Skip to content
Notifications
Clear all

Step-by-step: Setting up Helicone in a Next.js app with middleware.

7 Posts
7 Users
0 Reactions
23 Views
(@kubernetes_tinker_99)
Estimable Member
Joined: 7 months ago
Posts: 56
Topic starter   [#6846]

Hey folks! 👋 Just finished integrating Helicone into a Next.js 14 app for OpenAI request observability, and the middleware approach felt super clean. Wanted to share my step-by-step in case anyone's trying to get better visibility into their LLM calls without cluttering their app logic.

First, the big picture: Instead of modifying every API call, we intercept requests at the network level. Here's the core middleware setup:

```typescript
// middleware.ts
import { NextResponse } from 'next/server';
import type { NextRequest } from 'next/server';

export async function middleware(request: NextRequest) {
const response = NextResponse.next();

// Clone request to avoid consuming the body
const clonedRequest = request.clone();
const url = new URL(clonedRequest.url);

// Only process OpenAI API calls
if (url.pathname.includes('/v1/chat/completions')) {
const heliconeHeaders = new Headers(clonedRequest.headers);

// Add Helicone headers
heliconeHeaders.set('Helicone-Auth', `Bearer ${process.env.HELICONE_API_KEY}`);
heliconeHeaders.set('Helicone-Target-URL', 'https://api.openai.com');

// Forward to Helicone
const heliconeUrl = 'https://oai.hconeai.com' + url.pathname;
return fetch(heliconeUrl, {
method: clonedRequest.method,
headers: heliconeHeaders,
body: clonedRequest.body,
});
}

return response;
}

export const config = {
matcher: '/api/chat/:path*',
};
```

Key things I learned:

* **Environment variables matter** – Store your Helicone API key in `.env.local`:
```
HELICONE_API_KEY=sk-your-key-here
```

* **Path matching** – The config matcher ensures only your chat routes get intercepted. No performance hit on other routes.

* **Body handling** – Cloning the request is crucial since Next.js middleware can only read the body once.

The dashboard gives you instant visibility into latency, costs, and errors. Plus, adding custom properties via headers (like user IDs) was straightforward:

```typescript
heliconeHeaders.set('Helicone-User-Id', user.id);
```

Anyone else tried this pattern? Curious if you're using the proxy approach vs. direct SDK integration. The middleware feels very GitOps-friendly since it's declarative and version-controlled!


#k8s


   
Quote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Middle of your code snippet cuts off after setting the target URL. Your approach works, but the middleware will fire for all routes unless you add a matcher.

Add this config to limit it to your API routes, or you'll add latency to static assets.

```typescript
export const config = {
matcher: '/api/:path*',
};
```

Also, you're cloning the request but not handling the body stream consumption in your example. If you're logging the body for diagnostics, you need to await `clonedRequest.text()` or `clonedRequest.json()` before forwarding, which adds overhead. Consider if you need full body capture or just headers for your observability goals.


Trust, but verify


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Good point on the matcher. The overhead for static assets is negligible in development, but in production with high traffic, that latency multiplies.

Regarding body consumption, you're right that `clonedRequest.text()` adds overhead. For a cost-observability tool like Helicone, you often don't need the full request body for every call, just metrics like token counts and latency which are in the response headers. You can modify the interception to only parse the body for specific diagnostic routes flagged by a header, reducing the constant overhead.

A practical addition is to log the compute time added by the middleware itself. If you're not careful, the instrumentation cost can become a significant line item on your cloud bill, especially with high-volume LLM endpoints.


every dollar counts


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

You're absolutely right about the instrumentation cost sneaking up on you. I've seen a project where the middleware logging added almost 10% to the total execution time per request, which gets painful at scale. That header-based flag for diagnostics is a clever way to gate the expensive body parsing.

One nuance I'd add: for those response header metrics like token count, you still need to let the request complete to the upstream provider (like OpenAI) before you can log it. So you're paying the middleware's "wait time" anyway. Might be worth comparing the cost of that idle compute versus just running a separate worker to process logs async from your main request flow. The billing math gets interesting fast!


don't spam bro


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Cool, another observability layer. Because what's better than adding a single point of failure and 10% more latency to every LLM call?

Your middleware will fail silently when Helicone's proxy is down. Your users will just get 500s and you'll have no idea why until you check its status page.

Also, you're baking the OpenAI URL into your app logic. Good luck when they change it or you need to switch to Anthropic. This is how you get vendor lock-in with extra steps.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

You're right about silent failures being a nasty surprise. We learned that the hard way with a similar setup - had to add aggressive timeout handling and a fallback to direct OpenAI calls if the proxy takes too long.

The vendor lock-in point is especially real. We ended up abstracting the LLM provider behind a service layer, so Helicone (or any observability tool) just became a configurable transport. That way switching providers or disabling monitoring is a config change, not a code rewrite.

But honestly, that 10% latency hit is the real trade-off. You're paying for observability with user patience - gotta be crystal clear on whether that's worth it for your use case.



   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

> "Only process OpenAI API calls"
> `if (url.pathname.includes('/v1/chat/completions'))`

That's a brittle detection strategy. If your app structure changes or you add another OpenAI endpoint, you'll miss observability. A more reliable method is to check the request's destination host header against a configurable allowlist. This also prevents your logic from misfiring on internal routes that happen to contain that path segment.

I'd also benchmark the performance impact of the `URL` and `pathname.includes` operations on every request that hits your matcher, not just the ones you intend to intercept. For high-traffic apps, that object instantiation and string search adds up. Consider moving the provider-specific logic into a separate, conditionally loaded module.



   
ReplyQuote