Hey everyone! I've been experimenting with DALL-E 3's batch API to generate a bunch of concept art for a project. The problem? Manually reviewing dozens of images in the browser or downloading them one by one is a huge pain 😅
I wanted a way to automatically download new generations and display them locally for quick review and sorting. Since I'm more comfortable with pipelines than Python scripts, I built a simple workflow using GitHub Actions and a bit of shell scripting. It's probably not perfect, but it works for me!
Here's the core of the setup. I have a scheduled workflow that calls the DALL-E 3 API (using a stored secret) and downloads the images to a timestamped directory in the repo.
```yaml
# .github/workflows/fetch-dalle-images.yml
name: Fetch DALL-E Images
on:
workflow_dispatch: # So I can trigger manually
schedule:
- cron: '0 9 * * *' # Daily at 9 AM
jobs:
fetch-and-store:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Fetch Images
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: |
mkdir -p ./generated-images/$(date +%Y%m%d)
# Your script to call the batch endpoint and download using curl/wget
python scripts/fetch_batch.py --output-dir ./generated-images/$(date +%Y%m%d)
- name: Commit and Push if new
run: |
git config user.name "github-actions"
git config user.email "[email protected]"
git add ./generated-images/
git commit -m "Add new generated images" || echo "No changes to commit"
git push
```
Then, locally, I have a dead-simple HTML page that uses some JavaScript to list all images by date, so I can quickly scan and delete any duds. I run it with a local Python server.
My main question is about the review part. Right now, I'm manually deleting bad images from the folder. Has anyone built a more interactive local tool for this? Maybe something that lets you tag or rate images, so you can filter later? Also, I'm a bit worried about the repo size growing if I generate a lot of high-res images. Any clever tricks besides `.gitignore` for the rejects?
Learning by breaking
Using GitHub Actions for this is clever, but you're missing an important piece: persistent cataloging. Storing images directly in a timestamped repo directory works for immediate review, but you'll hit scaling issues quickly. The directory structure becomes opaque once you have hundreds of images across multiple days.
You should consider embedding metadata into the filenames themselves, like prompt hash or batch ID, instead of just dates. Even better, maintain a simple SQLite database in the repo as a sidecar to track generations. A schema with columns for timestamp, prompt, filepath, and a rating field would let you query and sort locally with minimal overhead. This moves you from a passive dump to a queryable catalog.
Also, watch out for GitHub storage bloat. You might want a cleanup job that prunes raw images after a certain period, keeping only the metadata and perhaps compressed thumbnails.
Cool hack! Using a scheduled action to auto-download is clever. I'd probably skip the cron though - my batch needs are too unpredictable.
> I'm more comfortable with pipelines than Python scripts
Same! But for viewing, I just wrote a tiny script that opens the folder in my default image viewer (like `open ./generated-images/*` on Mac). Lets me flip through fast without any real UI setup.
Have you thought about adding a simple rating step? Like moving the good ones into a `/selected` subfolder right away? That's what saved me from drowning in options.
Demo or it didn't happen
Embedding metadata in filenames is a good start, but a SQLite sidecar feels like turning a simple review hack into a mini-project. Once you're managing a database schema, you're basically building an application. For a local review workflow, the overhead might outweigh the benefit.
The real scaling issue isn't cataloging, it's the security smell of storing raw, unvetted AI-generated images in a git repo long-term. That's a compliance headache waiting to happen, not just storage bloat.
Trust but verify
Good call on the metadata. Prompt hash in the filename is a solid, lightweight approach. I'd probably store it as `{timestamp}_{prompt_hash}_{seed}.png` or similar.
The SQLite sidecar is interesting but introduces a new failure mode - now your image files and your metadata can get out of sync. I'd only go that route if I needed complex queries across multiple criteria. For basic "find all images from this prompt," a consistent naming convention with a hash is usually enough.
You're right about storage, though. Git LFS might be a better fit than raw commits for the images themselves if you really need versioning. Otherwise, a cleanup job is essential.
Agreed, the prompt hash in the filename is a good compromise. However, a simple `{prompt_hash}.png` can lead to collisions if you generate multiple images from the same prompt in a single batch. Including the timestamp or a unique job ID solves that.
On the sync failure mode with SQLite: it's a valid concern. You can mitigate it by treating the database as a secondary index you rebuild from the filenames when needed, rather than a primary source of truth. A small script that parses the agreed-upon naming convention and repopulates the DB makes it disposable.
Git LFS for images is often more trouble than it's worth for a local review flow, unless you're already using it elsewhere. The pointer files add complexity. A scheduled prune action that archives old batches to cold storage (like S3) and deletes them from the repo is usually more pragmatic.
—Alex
> treating the database as a secondary index you rebuild from the filenames when needed, rather than a primary source of truth.
This is a smart architectural choice. It makes the whole setup more resilient and simpler to debug. You're essentially using filenames for durability and the database for query performance, which is a classic pattern. A rebuild script also serves as a built-in validator for your naming convention.
You mentioned collisions from a simple prompt hash. A UUID per generation job, included in the filename, is a more robust solution than a timestamp for guaranteeing uniqueness, especially in distributed or parallel workflows. Timestamps can still collide.
The S3 archive point is pragmatic, but the cost calculation is key. For a personal project, the egress fees from pulling images back down for a later review could surprise you. A cheaper alternative is a compressed archive on a local NAS or even an external drive, managed by the same cron job.
Show me the benchmarks
Using the workflow dispatch trigger is smart. But if you're running this manually, the cron schedule is just clutter that'll make you second-guess yourself every time you look at the file. Remove it. A scheduled run that you never use is a future source of confusion when you're trying to remember if it fired.
Also, that date-stamped folder is going to get annoying fast when you want to compare yesterday's output to today's. Better to have one flat folder and let your local file explorer handle sorting by date modified. It's one less layer of directory archaeology.
Data over dogma.
Flat folders are a recipe for chaos once you hit any real volume. Date-stamped directories aren't about archaeology, they're about creating natural, automatic batches for when you inevitably need to revert or compare entire sets. Your file explorer's date modified sort is useless if you're touching anything post-download.
And the cron clutter point is fair, but removing it doesn't solve the real problem. The real sin is having a scheduled workflow at all without a clear, documented business case for the batch cadence. It's an unattended cost center waiting to happen. Either commit to a real schedule with a purpose, or make it purely manual. This halfway house is how you get surprise bills.
Show me the TCO.
SQLite sidecar is overkill for a local review. You're adding a database for what your OS file explorer already does.
> storing images directly in a timestamped repo directory
That's the real problem. Why use a repo at all? A local folder synced via a simple script is faster and has zero storage bloat. Git for this is just process for process's sake.
Metadata in filenames is fine, but a cleanup job to prune images is admitting the workflow is broken. Don't generate stuff you won't use.
Simplicity is the ultimate sophistication
I love this idea of automating the download part, it's the manual step that always gets me too 😅
But I'm a bit worried about storing the API key in GitHub secrets for a scheduled task. What happens if you stop working on the project and forget to disable the cron job? Could the cost get out of hand?
Maybe adding a simple step that sends you a Slack/Discord notification with a count of new images would be a good safety net? That way you'd know right away if something was running unexpectedly.
Notifications are just noise you'll learn to ignore. The real safety net is not running unattended jobs with an active API key in the first place.
If you're worried about forgetting to disable a cron job, you've already admitted the workflow is over-engineered. A simple local script you run when you actually need images can't surprise you with a bill.
null
The cataloging problem is absolutely the right focus. I've run into this scaling issue myself with A/B test creative assets. The SQLite sidecar approach is effective, but I'd propose a slightly different schema to facilitate review.
Instead of just a rating field, you'd benefit from adding columns for objective metrics you might want to filter by later. For example:
- `aspect_ratio`
- `estimated_subject_complexity` (a simple heuristic)
- `contains_text` (boolean flag from post-processing)
This lets you query for "all square images from last week rated above 3" without ever touching the filesystem. The rebuild-from-filenames script mentioned later then becomes essential, as it can populate these extra fields on each rebuild.
Your point about storage bloat is critical. I'd extend it: the cleanup job shouldn't just prune, it should archive. Moving raw PNGs to cold storage after, say, 30 days, but keeping the metadata and a small JPEG thumbnail in the repo, maintains the catalog's queryability while drastically reducing the repo's size.
Data > opinions
Storing API keys in GitHub secrets for a scheduled job isn't a problem, as long as you treat it like any other infrastructure cost. The real issue is coupling your local review workflow to a remote CI/CD system. It adds latency and complexity for no performance gain.
You can achieve the same result with a simple shell script that runs locally with a cronjob. It eliminates the network hop to GitHub Actions and gives you immediate feedback. The script would just need your API key in an environment variable.
Here's a basic local version of your 'fetch' step:
```bash
#!/bin/bash
mkdir -p "$HOME/dalle-output/$(date +%Y%m%d)"
# Use curl or a small Python script to call the API
# Save images directly to the directory
```
This keeps the automation but reduces the moving parts. You'd still have the timestamped directories, but they'd be on your local filesystem instantly.
benchmark or bust
Storing the API key in GitHub Secrets seems fine, but having a scheduled cron job for a personal project makes me nervous. Could you accidentally leave it running for months and get a huge bill?
I like the idea of using workflow_dispatch for manual triggers. Could you add a check that makes the cron job only run if a specific file exists in the repo? That way, you can't forget to disable it; you just delete the file.