I've been running Pika through its paces for the last three weeks, integrating its video generation API into a prototype real-time analytics dashboard. The core value proposition hinges on predictable operational cost, which is currently being completely undermined by what appears to be a significant discrepancy between reported credit consumption and actual generated assets.
My monitoring pipeline logs every single API call. According to the Pika dashboard's usage report for the last 7 days, I've been charged for 1,247 credits. However, my own immutable audit log, which correlates request IDs to successful generation outputs stored in our data lake, only accounts for 842 completed generations. That's a discrepancy of 405 missing generations, or roughly 32.5% of the billed credits. This isn't a rounding error or a minor glitch; it's a substantial financial leak.
I've ruled out the obvious:
* **Failed generations:** The API's response status codes and bodies are logged. We have zero `5xx` errors in the period, and all `4xx` errors (bad prompts, etc.) are accounted for separately and are not included in the 842 successful outputs.
* **Async job polling:** We use a simple, idempotent polling mechanism. The logic is straightforward: we call `POST /generate`, receive a `generation_id`, then poll `GET /generation/{id}` until status is `completed` or `failed`. Only upon `completed` do we log a success. The code is below.
```python
# Simplified audit log snippet
def log_generation_attempt(request_id, prompt, status_code, response):
audit_log.append({
"request_id": request_id,
"timestamp": utc_now(),
"prompt_hash": hash(prompt),
"http_status": status_code,
"response_body": response,
"billed": None # Populated later via webhook
})
def handle_webhook(event):
# Pika sends this on 'completed'
if event['status'] == 'completed':
match = find_audit_entry(event['generation_id'])
if match:
match['billed'] = event.get('credit_cost', 0)
match['final_url'] = event.get('url')
```
* **Webhook duplication:** The webhook handler includes a check for existing `generation_id` to prevent double-counting. Our data shows no duplicates.
The only correlation I can find is that the credit overage seems to spike when using specific motion control parameters (`--motion intensity`). However, the API returns a success response even if the motion effect is subtle or, to my observation, non-existent. Is the system billing for a more complex model inference attempt regardless of the visible output quality?
This is a critical issue for anyone building a production workflow. Cost predictability is a cornerstone of data pipeline design. I need to know:
1. Is there a hidden, non-idempotent retry mechanism on Pika's side that consumes credits but doesn't always surface to the client?
2. Are credits consumed for *attempted* processing that fails internally after returning a `200 OK` to the initial request?
3. Where is the detailed, line-item credit billing log? The current dashboard aggregation is useless for audit purposes.
Without transparent, itemized billing that we can reconcile against our own logs, it's impossible to trust the platform for any serious data workflow. Has anyone else done a systematic audit and found similar gaps? Share your methodology and findings.
—davidr
—davidr