Skip to content
Notifications
Clear all

How do I schedule a workflow to run on the last weekday of the month?

14 Posts
14 Users
0 Reactions
14 Views
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
Topic starter   [#25629]

I’ve been evaluating Flux for orchestrating our analytics pipelines, and I’ve hit a scheduling requirement that seems deceptively simple but is proving tricky to implement: I need a workflow to execute on the **last weekday of every month**. This is critical for our month-end financial reporting, where we must run our dbt models and refresh the corresponding Looker dashboards after the final business day’s data is ingested, but before the first calendar day of the new month.

The standard cron syntax (e.g., `0 0 28-31 * *` with additional weekday checks) feels clunky and error-prone, as it would trigger on multiple days and require complex conditional logic within the job itself to determine if it’s truly the last weekday. I’m seeking a more declarative, Flux-native solution that defines the schedule precisely at the orchestration layer.

My ideal solution would involve Flux’s scheduler understanding calendar logic. I’ve explored the documentation on schedule objects and intervals, but most examples revolve around fixed dates, times, or simple intervals (e.g., every Monday). Has anyone successfully implemented this pattern? I’m particularly interested in:

* Whether this can be achieved purely through the schedule definition in the Flux YAML, or if it necessitates a “guardrail” job that runs daily and uses an API call or a query to determine if the current day meets the criteria.
* The robustness of the solution across months with varying numbers of days and weekend placements (e.g., March 31st on a Saturday would mean the last weekday is Friday, March 30th).
* Any potential pitfalls with timezones in this context, as our data warehouse operates in UTC but our business days are defined in EST.

A conceptual outline of the logic I’m considering is below, but I’m unsure how to translate this into a concrete Flux schedule:

```yaml
# Pseudo-logic for the desired schedule
schedule:
type: "complex_calendar"
rule: "Last weekday of month"
time: "00:00"
timezone: "America/New_York"
```

Failing a native schedule type, I’d appreciate insights on the most reliable programmatic approach. For instance, would the best practice be to create a lightweight daily workflow that calculates `CURRENT_DATE = LAST_WEEKDAY_OF_MONTH` and then triggers the main reporting pipeline as a downstream dependency only when the condition is true? This seems to add overhead and move scheduling logic into the pipeline code, which I’d prefer to keep separate.

I’m eager to review any community workflows, YAML configurations, or even SQL date logic snippets used to identify the last weekday dynamically. Data quality hinges on predictable, accurate execution windows, so getting this pattern right is a priority.

- dan


Garbage in, garbage out.


   
Quote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

The calendar logic part is tricky. I'm in a similar spot trying to schedule month-end tasks.

Flux's `interval` schedule with a custom cron string might be your only option for now. You'd still need to handle the "last weekday" check inside the workflow logic, which isn't ideal.

Have you looked at whether you could use a simple `0 0 28-31 * 1-5` cron and then add a filter step at the start of your pipeline? It would run a few times but only proceed on the correct day.



   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

That cron expression (`0 0 28-31 * 1-5`) doesn't work. It triggers on any weekday *within* the 28-31 date range, not the *last* weekday of the month. You'd miss months where the 28th is a Saturday.

You're stuck with the conditional logic. Build a small script that checks if tomorrow is month N+1. Run it daily with a simple cron (`0 0 * * *`), and only proceed if the check passes. No wasted triggers.



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Yeah, that declarative scheduler ideal is the dream, isn't it? I've fought this exact battle for month-end commission calculations. Even if Flux's scheduler doesn't have native calendar logic, you can at least encapsulate the check cleanly at the very start of your workflow.

Instead of embedding messy `if` logic in your main pipeline, create a first step that purely answers "Is today the last weekday of the month?" using a simple script or a library. Have it exit with a failure code if not. Your workflow triggers daily, but the core logic only ever runs on the correct day. It keeps your main dbt and Looker refresh steps clean.

The real headache is testing the edge cases - February in a leap year, months ending on a Sunday, etc. You'll spend more time validating that calendar logic than you think!


Pipeline is king.


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

That daily trigger with an early fail step is just shifting the complexity. Now you've got a job that fails most days, clogging your monitoring with noise. Your alert fatigue goes up because you're constantly seeing expected failures.

The edge case testing is the real killer. You think you've covered Feb 29th, then a holiday Monday at month-end breaks your logic. Now your financial reporting is a day late.

Just use a proper job scheduler that understands business calendars. Cron isn't the only tool.


Don't panic, have a rollback plan.


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Exactly. Shifting to a "fail-fast" wrapper just turns a scheduling problem into an ops noise problem. You're trading one headache for another.

The real kicker is user49's point about holidays. If your business logic actually depends on a business day, a simple weekday check is already wrong. You need a calendar that knows about New Year's Eve falling on a Sunday, or bank holidays.

Cron and its cron-like descendants are built for periodic time, not calendar logic. Sometimes the right answer is to use a scheduler that understands dates, not just to keep hammering the same nail.


Show me the TCO.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

>I'm seeking a more declarative, Flux-native solution

Ah, the vendor promise. Good luck with that.

You're hitting the classic wall where "simple" scheduling logic isn't simple at all. The declarative ideal assumes the scheduler has a calendar brain, which Flux's schedule objects clearly don't. Your ask isn't even about the *last weekday*; it's about the *last business day*. That's a whole other rat's nest of holidays and regional calendars that no orchestrator's native syntax will solve for you.

You're stuck with a wrapper script or a different tool. Anyone telling you otherwise is selling magic beans wrapped in YAML.


Trust but verify.


   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

Wait, so you're saying the last *business* day is different from the last *weekday*? I never even thought about holidays. 😅

So if my company closes for a bank holiday on the 31st, a weekday check would still run, but the data wouldn't be ready. That's a much bigger problem.

Is there a common pattern for handling holiday calendars in pipelines, or is that always a custom script?


CloudNewbie


   
ReplyQuote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

You've hit on the core tension: the desire for declarative scheduling versus the reality of calendar complexity. While I share your wish for a Flux-native solution, my experience dictates that embedding this logic directly into the scheduler's definition is often the wrong layer to solve it.

The declarative ideal breaks down because "last weekday" is business logic, not just a time pattern. As others have noted, it often morphs into "last business day," which requires a holiday calendar. That's a data problem, not a scheduler problem. The most maintainable pattern I've used is to treat the calendar as a source of truth - a small table in your warehouse with valid business dates - and have the very first step of your workflow query it. The workflow can trigger on a safe, wide window (like the 28th through 31st at midnight), but it immediately checks that table and exits cleanly if today isn't the target day.

This keeps the scheduling simple and the complex date logic where it belongs: in code, version-controlled and testable alongside your dbt models. You can even generate that calendar table with dbt, incorporating company holidays. Trying to force this into any orchestrator's native schedule syntax, Flux or otherwise, makes the logic opaque and hard to audit.


Your data is only as good as your pipeline.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Yeah, that declarative dream is exactly what pulled me to Flux too. But reading this thread, I'm realizing it might be solving the wrong half of the problem.

You said you need this for month-end reporting *after the final business day's data is ingested*. Doesn't that mean you actually need to schedule based on when your data pipeline finishes, not just a calendar date? What if your source system's ETL is delayed by a holiday?

Seems like you'd still need a check for data readiness, even if Flux could magically schedule the last weekday.



   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Oh, that's a really good point about data readiness. I hadn't even considered that.

So even if Flux could magically pick the last weekday, you'd still need to check if the upstream data is actually there, right? Especially with holidays causing delays.

Does that mean the calendar date check is just the *first* gate, and you'd need another step after to verify the source data has landed?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Yes, that's precisely the logical extension. The calendar check is just the initial prerequisite. You then need to verify the existence or completeness of your upstream dataset before proceeding.

The cleanest pattern I've seen is to have your workflow's first step answer "Is today the last business day?" and, if so, then immediately execute a second step that polls for a known artifact from the ingestion pipeline. That artifact could be a success file in object storage, a specific partition in a table, or a flag in a metadata service.

You're scheduling the *opportunity* for the workflow to run, but the workflow itself manages the actual execution logic based on data state. This separates the concerns of timing from data readiness.


null


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

That separation of concerns is the only sane way to run this. We've used the "success file" pattern for years. The critical detail everyone misses is making the artifact check idempotent and stateful.

Your workflow shouldn't just poll for a file; it should write a lock or a processed marker once it proceeds. Otherwise, if your workflow pod gets rescheduled, you risk double execution on the same dataset. The artifact check step must be a transactional "check and claim" operation, not just a stat() call.

Also, your "Is today the last business day?" step better be a call to a maintained API or a database table you actually update for holidays. If it's a hardcoded script in your repo, you've just recreated the cron problem inside a container.



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Exactly right about the transactional check, that's a crucial implementation detail. A simple stat call is a race condition waiting to happen.

I'd add that the source of truth for "last business day" itself needs the same idempotent consideration if multiple workflows or teams are querying it. If it's an API, it should handle concurrent requests cleanly; if it's a table, you need to think about read consistency.

The pattern really becomes a two-part state machine: confirm the date, then claim the data. Miss one, and you're just building a more complex cron that fails in new ways.


—HR


   
ReplyQuote