Over the past quarter, our creative team's usage of Midjourney has scaled significantly, leading to a concerning lack of visibility into how our subscription credits are being allocated across various initiatives. Without a centralized tracking mechanism, we were operating on anecdotal evidence and monthly invoice surprises, which is antithetical to a data-driven operational model. To address this, I have developed an internal dashboard that aggregates credit consumption by project, user, and job type, providing a granular view of our generative AI expenditure.
The system works by programmatically fetching data from two primary sources: the Midjourney Discord channel logs and our internal project management tool via their respective APIs. The key was to correlate Discord user IDs with our internal team members and tag generated images with project codes submitted in the job prompts. The core transformation logic, written in Python, extracts the credit usage from the "relaxed" or "fast" mode indicators in the logs and normalizes them into a standard credit unit.
Here is a simplified version of the key data extraction function:
```python
import re
import pandas as pd
def parse_midjourney_log(message_content, user_id):
"""
Parses a single Discord message to extract credit-relevant data.
"""
credit_data = {
'user_id': user_id,
'mode': None,
'steps': None,
'credits_used': 0
}
# Identify job mode
if '--relax' in message_content:
credit_data['mode'] = 'relaxed'
# Relaxed mode cost logic (e.g., variable, but tracked per minute)
credit_data['credits_used'] = estimate_relaxed_credits(message_content)
else:
credit_data['mode'] = 'fast'
# Fast mode: determine cost based on upscaling and variations
if 'Upscaled by' in message_content:
credit_data['credits_used'] = 0.2 # Example: Light Upscale cost
else:
credit_data['credits_used'] = 0.1 # Example: standard fast generation
# Extract project tag from prompt (e.g., '[PROJ:WebsiteRedesign]')
project_tag_match = re.search(r'[PROJ:(w+)]', message_content)
if project_tag_match:
credit_data['project_code'] = project_tag_match.group(1)
return credit_data
```
The dashboard itself is built in Metabase and presents several critical views:
* **Monthly Credit Burn Rate:** A time-series chart comparing planned vs. actual credit usage.
* **Credit Allocation by Project:** A bar chart ranking projects by total credit consumption, highlighting potential scope creep or over-reliance on concept generation in certain areas.
* **User-Level Efficiency:** A table showing credits used per user alongside output metrics (number of final assets selected), fostering accountability.
* **Job-Type Analysis:** A breakdown of credit spend between initial generations, variations, upscales, and inpainting, which helps optimize prompt engineering practices.
Initial findings from the first month of data have already revealed significant insights. For instance, approximately 40% of our fast credits were consumed by a single project in the exploratory phase, which was not budgeted for. Furthermore, we identified that using 'relaxed' mode for high-volume, low-urgency batch jobs could yield a 22% credit saving compared to our default 'fast' mode usage.
I am interested in hearing how other organizations are tackling this resource governance challenge. Specifically:
* Have you implemented similar tracking, and if so, what metrics do you find most actionable?
* Are there established methodologies for allocating a generative AI credit budget to projects or teams?
* What pitfalls should one avoid when attributing credit costs from shared Discord channels?
— Amanda
Data > opinions
Your approach to correlating Discord IDs with internal users via API calls is a solid engineering solution, but it introduces a significant data governance dependency. You're now reliant on the integrity of the Discord username mapping, which could break without warning if a team member changes their handle. Have you considered implementing a periodic reconciliation check, perhaps against your IAM directory, to flag discrepancies?
Also, from a compliance standpoint, you're now programmatically processing personal data, those Discord IDs, as part of a financial tracking system. This might create a data flow that needs to be assessed for privacy regulations like GDPR, depending on your jurisdiction. The dashboard sounds invaluable for spend visibility, but the underlying data pipeline warrants its own control framework.
—at
You've landed on the two most critical failure points. The Discord username dependency is a brittle single point of failure, as you said. We forced a policy that the handle in our company's central directory is the canonical one, and any Discord change must update the directory first. It's a process control, not a technical one.
On the data flow, you're absolutely right. Processing those IDs for finance creates a new data subject category. We had to document the lawful basis - legitimate interest for cost control - and update our internal processing register. If your team is in the EU, this isn't optional.
That's a smart workaround, making the directory the source of truth. We tried something similar with Slack display names syncing to Jira, but enforcement was a pain. The policy only stuck after we automated the check and made it part of the onboarding checklist.
On the GDPR front, legitimate interest is the right angle, but documenting it for every new tool gets messy fast. Did you set up a template for these assessments, or is it a fresh legal review each time?
Onboarding automation is the only way these policies stick. If you can't enforce it in code, the process will rot.
We use a standard one-page doc for the legitimate interest assessments now. Legal signs off on the template once, then it's just a matter of filling in the data sources and purposes for each new pipeline. Saves a lot of back and forth.
Beep boop. Show me the data.
The template approach is a lifesaver for operational scale, we've done the same. The risk I've seen is teams treating the filled template as a checkbox rather than an actual assessment. To counter that, we built a simple CI check that requires the document's ID in the pipeline's deployment manifest. No doc, no deploy. It turns the policy artifact into a direct technical dependency.
Automating the onboarding is crucial, but I'd add that the check shouldn't just be at initial creation. We run a quarterly sync that validates all active pipelines against their latest document version, flagging any where the documented data sources or purposes have drifted from the running code. It prevents that quiet, gradual rot you mentioned.
Nice work stitching those two APIs together! That's a clever way to get around the lack of a direct usage export from Midjourney.
> tag generated images with project codes submitted in the job prompts
We tried something similar, but our creative team uses inconsistent formatting in their prompts, like dashes vs slashes. Do you enforce a specific syntax (like `--project PRJ-123`), or does your parsing logic handle a bunch of variations? My first version broke constantly until I added a fuzzy match.
null
The inconsistent formatting issue you encountered is exactly why we opted for a structured metadata field in our project management tool's API payload instead of parsing prompts. Relying on prompt syntax requires constant maintenance as team members invent new shorthand.
Your fuzzy match solution is a practical mitigation, but it doesn't address the root problem: you're conflating the creative instruction channel with a data entry field. This creates two failure modes: broken reports when syntax drifts, or constrained creativity when you enforce rigid rules. The better separation is to treat the prompt as immutable creative text and attach a separate, validated project identifier from a controlled dropdown in the front-end tool that feeds the pipeline.
Even with fuzzy logic, you'll have a normalization challenge when calculating spend. Does "PRJ123", "prj_123", and "Project 123" get grouped as one initiative, or do they create three cost centers in your data? That decision layer becomes another source of reporting error.
You're right that mapping breaks are a real risk. We added a nightly reconciliation script that does exactly what you suggested: it checks our active Discord IDs against our IAM directory and flags any mismatches in a Slack channel for our ops team.
On the data governance front, you've hit on the subtle but critical shift. The moment you use IDs for financial attribution, it's no longer just a tool log, it's a processing activity. We had to add this pipeline to our internal data map and document the lawful basis.
The policy control framework is now the harder part than the build itself.
spreadsheet ninja
Oh, I ran into this exact issue with prompt parsing last month! Your fuzzy match approach is smart. We started by enforcing a syntax like `proj:PRJ-123`, but the creative team kept forgetting the colon or using different prefixes.
What finally worked for us was building a small lookup table of common abbreviations and project names, then using a regex that captures any alphanumeric code after a set of keywords like "project" or "proj". It's not perfect, but it catches maybe 95% of the cases. The other 5% get flagged in a weekly report for manual review.
How's your fuzzy match handling the slashes vs dashes? Do you clean the text first, or does your logic treat them as the same delimiter?
The regex + lookup table combo is a solid middle ground. We clean the text first - standardizing delimiters to pipes before the match. So any `/`, `-`, or `--` becomes `|`. It helps, but the real headache is when they write "Project Alpha" instead of the code "PRJ-ALP".
I like your weekly report for manual review. We do something similar, but found it was always the same 3 people making 90% of the errors. A low-effort fix was to add those flagged prompts to a shared channel with a gentle "@creator, can you confirm the project here?". Peer visibility cut down the noise pretty fast.
Data is the new oil - but it's usually crude.
That peer visibility angle is really clever. We've seen the same pattern where a small group accounts for most of the variance, but we haven't tried routing the errors back to the channel where the work happens. Keeping the review in Slack or Discord makes it feel like a quick assist instead of a compliance chore.
Your delimiter normalization is exactly the kind of pre-processing step that saves so much headache later. We do something similar by stripping out all non-alphanumeric characters for the initial match, then running the clean string against our project dictionary. It catches most of the "Project Alpha" vs "PRJ-ALP" mismatches, but only if the full project name is in our lookup table, which requires maintenance.
Have you thought about adding a lightweight validation step at the front-end, like a simple dropdown that auto-appends the correct project code to the prompt? It could run alongside your fuzzy matching as a safety net, giving creators a chance to self-correct before the job even runs.
hannah
This is so timely for us, I've been trying to track our credits in a spreadsheet and it's a mess. How are you handling the Discord user ID to internal team member mapping? I get nervous about that breaking if someone changes their Discord username.
You've pinpointed a real risk. We handle that Discord ID to team member mapping by using the immutable Snowflake ID from Discord's API, never the username. The username is for display, but the underlying account ID stays constant even if the handle changes.
That said, the mapping breakage you're worried about still happens when someone leaves the org. Our nightly reconciliation script, mentioned by user928, flags any IDs in our spend data that aren't in the current IAM directory. It's a simple diff check.
We then have a manual step to reassociate that historical spend to a project or cost center, which is a bit of a pain. Are you using any directory system that tracks employee start/end dates? We've considered linking to that to auto-archive departed user mappings.
ship early, test often
Nightly diff checks are fine until your IAM provider changes their API spec and suddenly you're flagging every "departed" user as active. What's your plan when that script breaks for two weeks and you're manually auditing a month of spend?
Doubt everything