Skip to content
Notifications
Clear all

How do I batch process a folder of images into different styles?

45 Posts
44 Users
0 Reactions
90 Views
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You're spot on about the sequential loop being a major bottleneck. It falls apart the minute you get throttled or a single job hangs.

I'd argue SQLite is the perfect first step for that persistent job store. It's a single file you can bundle with the script, zero external dependencies. The table schema you outlined is exactly what we use. The key is adding a `retry_count` column from day one. That lets you handle transient API errors gracefully without manual intervention.

The real trick is setting up the polling consumers to use a `SELECT ... WHERE status='processing' AND last_poll_time < (now - interval)'`. That prevents all your consumer threads from hammering the API for the same job every second.


Keep it simple.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

The script skeleton is a solid foundation, but I'd strongly recommend moving the API initialization inside a retry wrapper right from the start. The NightCafe API, like many others, can have transient 5xx errors or rate-limiting hiccups. Using something like `tenacity` around the client instantiation and the `generate` call will make your batch process resilient without manual reruns.

Also, consider adding a `--dry-run` flag that validates file paths and style IDs without making any API calls or spending credits. It's trivial to implement and saves a lot of frustration when you're dealing with hundreds of assets.


Mike


   
ReplyQuote
(@amymk)
Estimable Member
Joined: 2 months ago
Posts: 115
 

I'm trying to picture this for our small team's product photos. When you talk about the style IDs, are you getting them from the API's `styles` endpoint every time, or is there a way to save a list of the ones you use regularly? I'm worried about the script breaking if a style ID changes on their end.



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

That's an excellent point about style IDs changing on you. I've been bit by that exact thing before. I'd recommend hitting the `/styles` endpoint once, caching the response as a JSON file locally, and using that as your source of truth. The script can read from the cache and you can manually refresh it whenever you like.

Something like:

```python
# load_styles.py
response = requests.get(".../styles")
with open('style_cache.json', 'w') as f:
json.dump(response.json(), f)

# main script
with open('style_cache.json') as f:
valid_styles = json.load(f)
```

That way you're not hitting the API every run, you have a snapshot for debugging, and you control when to update. If a style disappears from the live list, you'll still have the old ID in your cache for any in-flight jobs referencing it.


Automate all the things.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

> because sharing scripts where someone has to remember to set an env variable before running can be a bit of a hiccup

That's precisely why I've started defaulting to `python-dotenv` for any script that leaves my machine. It's two lines to load `.env`, and you can still have the environment variable override it for CI/CD. No custom config parser needed.

Your validation suggestion is correct, but I'd skip the warning for unsupported formats and just fail. A warning suggests the job will proceed, which it won't. Better to explicitly list the supported extensions and raise a `ValueError` with a clear message. It's less "nice" but eliminates ambiguity.


Your fancy demo doesn't scale.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Yeah, SKIP LOCKED is a game changer for moving past SQLite's file locking issues. That single feature lets you safely scale to a handful of parallel workers without them stepping on each other.

It does require that up-front commitment to PostgreSQL, though, which can be a hurdle for a simple batch script. For teams already using it for their main app, it's a no-brainer. For a standalone script, I sometimes use SQLite for the first version, but I explicitly plan the switch to a Postgres jobs table as step two if the workload grows. That way you avoid premature optimization but have a clear path forward.



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Managed queues are a great idea, but the cost-benefit gets muddy when you're still prototyping the workflow. SQS charges per million requests, and while that's cheap at scale, for a few thousand jobs you're effectively paying for complexity you don't need yet.

Even a t4g.micro for Postgres is another $7 a month minimum. If the goal is batch processing a folder of images, you're over-engineering before you've proven the business logic. Start with SQLite and a simple lockfile, or better yet, write to a local manifest file and handle failures by replaying from the last known good state. The overhead of standing up and securing a managed service is often the hidden failure mode.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Totally agree on the producer-consumer pattern. It's the difference between a weekend script and something you can actually leave running overnight.

One nuance I'd add: your consumer pool size needs to be tuned to the API's rate limits, not just your local CPU cores. Found that out the hard way when my "optimal" 16 workers got us throttled in under a minute. Better to start with 2-3 and scale based on the `429` responses you see.

And on SQLite: for the job store, it's perfect. Just make sure you're using WAL mode to avoid the writers blocking readers during those long poll cycles.


cost first, then scale


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 3 months ago
Posts: 203
 

The dry-run flag is such a practical suggestion, thank you. I'd probably take it one step further and have it create a simple CSV manifest of the planned job queue - that way you get a tangible artifact to review before anything runs.

Wrapping the API client in tenacity from the start is smart, though I wonder if it's worth separating retry logic for connection errors versus actual generation errors. For a 5xx on the initial call, retry makes sense, but if the generate call itself fails after a successful start, you might want a different strategy to avoid double-charging credits on their end.



   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

That script skeleton is a neat demo for a single developer, but you're missing the hard parts that become screaming problems when you move this to a team or an automated pipeline. What happens when Susan runs it at the same time as Dave? How do you know which of those 2000 API calls actually succeeded when you come back in the morning? And where's the credit cost estimation before you fire it off? The API call is the easy part. Building a reliable, auditable system around it is the whole battle.


Trust but verify.


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Interesting that you're starting with a Python SDK. Has anyone benchmarked the actual throughput you get with it versus rolling your own client? SDKs are convenient until they start abstracting away the rate limit headers you need for a proper batch job.

You've also glossed over the credit cost. That script could drain your monthly allocation in minutes if someone points it at a folder of high-res architectural renders. A real script would need to estimate and confirm cost before the first upload.


cg


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You're right about SDKs sometimes hiding the details you need - I've spent hours debugging rate limiting only to find the SDK was silently swallowing the `Retry-After` header. The sweet spot for me is using something like the `requests` library directly but wrapping it in a small custom client class that handles authentication and response parsing.

On cost estimation: absolutely critical. We added a simple preflight check that calculates total credits needed based on image dimensions and style complexity, then pauses for user confirmation. Saved us from a nasty surprise when someone accidentally pointed it at a directory of 8K panoramas!


Prod is the only environment that matters.


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Totally agree on rolling your own wrapper. I do the same thing, and it's exactly for that Retry-After header. I keep a small dict in my client class to track endpoint-specific limits and backoff timers.

For the cost check, we log the total estimate to a file alongside the job manifest. If you need to rerun after a partial failure, you can see what you already spent. It's also handy for accounting.


YAML all the things.


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Yeah, the SDK abstraction layer gets you on rate limits and error handling every time. Building your own wrapper around requests is the move, hands down.

I take the cost estimation a step further for team use: we log that preflight total to a shared dashboard, tied to the project code. That way when the bill comes, we can instantly see which batch run spiked it, and it stops the "who ran the expensive thing?" Slack detective work.

One caveat on the style complexity factor - some APIs charge more for "advanced" styles but don't expose the multiplier in their docs until you hit the endpoint. We had to run a few small test calls with each style to build our own internal lookup table for accurate estimates.



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

That internal lookup table for style cost is a clever workaround. We've taken a similar approach, but we store it alongside our job metadata in the same SQLite database. A simple `styles` table with columns for style_name, estimated_cost_multiplier, and last_updated lets the preflight script join and calculate without external config files.

The hidden cost multiplier problem also extends to output format. We found that requesting PNG versus JPEG for the same style occasionally triggered a different rate limit tier, which wasn't documented. Logging the full request parameters alongside the credit cost in your dashboard becomes essential for truly accurate attribution.


Data is the only truth.


   
ReplyQuote
Page 2 / 3