Orchestration on Modal¶
How Stardag runs on Modal: what a deployed app contains, the two ways a build can be driven, and — for the reactive mode — how a build with no process of its own learns that something changed. For setup and recipes see Integrate with Modal; this page is the model behind it, and assumes Build & Execution.
What a deployed app is¶
StardagApp.finalize() registers a handful of Modal functions:
| function | runs |
|---|---|
worker_<name> (one per worker_settings entry) |
one task per input, self-reporting its lifecycle |
build |
a resident build: the ordinary scheduling loop, in a Modal container |
bootstrap, tick, tick_watchdog |
the reactive scheduler (below) |
All of them make registry calls, so all of them carry the registry secret.
A task's worker and resources are chosen by the app's worker_selector,
its concurrency-limit keys by its limit_key_selector; both are
deployed-app configuration, identical for every build of the app.
Detached execution and self-reporting workers¶
The Modal executor is detached: it spawns the worker and records the
Modal function-call id in the registry with the task's TASK_STARTED
event, rather than holding a blocking call open. The worker then reports
its own lifecycle — started (with its own call id), completed (plus
artifacts), suspended (dynamic dependencies), failed, interrupted — from
inside the container.
Three things follow, and every mode below rests on them:
- Re-attach instead of re-execute. A resumed build, or another build
wanting the same task, finds it
RUNNINGwith a live reference and attaches. Restarting an orchestrator does not restart your long tasks. - Cancellation is cooperative, and stopping is yours. Nothing reaches
into a container: a worker asks at its own checkpoints whether it is
still wanted and exits cleanly when it is not, writing no output and
reporting no completion. Cancelling, failing or completing a build
releases the claims held by every one of its plans — superseded ones
included — so its tasks are immediately available to the next build,
and the released tasks go to
CANCELLED(actionable), never a result. To end the containers themselves, rather than leaving them to notice on their own, is a build's executions (see Builds stop over executions) — a human-driven action, and deliberately not something anything automatic does. - Liveness is the backend's answer. A scheduler asks Modal whether a recorded call is running, finished or gone — no heartbeats.
A task's registry state is therefore accurate independent of any orchestrator's lifetime, which is what makes the orchestrator optional.
Two ways to drive a build¶
Resident — build_trigger(tasks) (or build_spawn): the build
function runs sd.build in one container for the whole build. Simple, and
the fastest for a DAG of many short tasks: the scheduling loop is a tight
in-memory loop. Its cost is that the container lives as long as the
longest task, and a three-day build with two long tasks pays for a
container that mostly waits.
Reactive — build_trigger(tasks, reactive=True): no resident process.
The build is driven by short-lived, idempotent scheduler ticks, spawned
when something changes. While a DAG churns, one tick lingers and behaves
like a resident loop; when only long tasks remain in flight, nothing runs
but your tasks. Reactive scheduling is also what makes cross-build
coordination, retries of infrastructure failures and checkpoint/resume of
preempted tasks work unattended.
Both modes share the registry, the claims and the limits, so they mix: a resident build on a laptop with Modal workers (a hybrid run) and a reactive build in the cloud coordinate through the same rows.
Reactive scheduling¶
A build's life¶
trigger (cheap, no target I/O):
mint or resume the build; register the roots; spawn bootstrap
bootstrap (one container, once per trigger):
fix the build's scope (this deployment + its settings) and apply the
settings; discover the DAG next to the target root; register it; check
every incomplete task can be rebuilt from registry data; arm the build;
spawn the first tick
tick (short-lived, single-flighted per build):
acquire the build's scheduler lease (held → exit)
loop:
clear the build's wake-up flag; read the frontier
act: spawn ready tasks detached (claim first), leave live claims
alone (a lapsed one is runnable again), heal completions,
retry a failed spawn within its own budget
terminal? → complete / fail the build (releasing its claims)
acted? → re-read immediately; else linger on the wake-up flag
on the way out: re-read the flag before and after releasing the lease
The bootstrap writes the reactive marker last, after the whole DAG is registered, so no tick ever sees a half-registered build. Discovery runs inside Modal because it is target I/O — a mounted volume there, a rate-limited API from a laptop — so triggering needs registry credentials only.
A tick rebuilds every task it schedules from the registry's stored
instance body, and from nothing else. That is what makes a running build
safe to re-plan under new code — the stored body is what that scope
constructed, so nothing carries over from the process that registered it.
It is also a
hard requirement on the app: rebuilding resolves a class through stardag's
polymorphic registry, which is populated by importing the defining
module, so the app must declare
task_modules.
The default infers it from the app's own package; an app where inference is
impossible (__main__, a loose script) is resident-only and its reactive
trigger is refused. The bootstrap dry-runs the reconstruction over the
whole discovered DAG and refuses a build it could not drive.
A tick's fan-out is bounded (max_concurrent_actions, default 50) and the
work one pass commits to is capped by a duration budget derived from the
tick function's own timeout, so a container never starts more than it
can live to finish. Truncation re-reads immediately; it is never a stall.
Retries and interruptions¶
A tick retries only the one failure no backend can retry for you — a spawn
that fails before any container starts — up to TickConfig.max_attempts
(default 2) within that single claim; nothing about the budget persists
across ticks. Three other failure shapes bypass it entirely, distinct from
each other and from it:
- A worker that dies with no restart coming (OOM, a crash, a timeout
nothing caught) simply lets its claim lapse, and the next claiming
start takes the execution over as a fresh attempt — up to
max_executions(default 20) executions of the task in the build, after which the tick fails it with the count instead, and the build'sfail_modeapplies. - A preemption is not the same. Modal restarts the execution itself, on the same call id, and the worker keeps its claim across the restart — see below.
- An exception inside your task. The worker reports it
FAILEDitself, and the tick never retries aFAILEDtask automatically — the build'sfail_modedecides, and onlystardag tasks retryor a re-trigger (not Modal's ownretries=, which never touches registry state) moves it back toPENDING.
A task past its timeout that caught the interruption, checkpointed and
raised ResumableInterruption is recorded INTERRUPTED and resumed, up to
max_interruptions (default 20). A task preempted the same way is not:
Modal restarts that input itself, on the same call id and in seconds, which
is better than a reschedule on every count — so the worker keeps its claim
and gets out of the way, recording only that a restart is now due. An
interruption the task did not catch is an ordinary failure: it had no plan
for one. Recipe and knobs: Preemption and
timeouts.
Which of the two it is comes from the exception the task caught, not from
how long the execution had been running — Modal raises KeyboardInterrupt
for a preemption and InputCancellation for a timeout, and only the first
restarts.
And it comes from the worker, which is the only thing that knows. A tick never probes a running task for what ended it: a live claim is left alone, whoever holds it, and a lapsed claim is simply runnable again — no probe, no report-grace knob, no reaper. The claim's own TTL (the executor's timeout plus a fixed grace) is what already gives a worker time to report before the registry would call its execution gone.
Wake-ups: how a build with no process learns something changed¶
A reactive build progresses only while a tick runs for it, and a tick runs only because something spawned one. The registry sees every write that can change a build's frontier but has no executor and never spawns — so a wake-up is two halves, done by two parties:
- The registry flags. Every change to a task's status — by any
worker, any tick, a resident build, an operator in the UI or CLI —
flags every other live reactive build whose active plan holds that
task. A
transition out of
RUNNINGalso flags the builds queued on the concurrency-limit keys the task held. Cancelling a build flags the build itself. The flag isneeds_tick_aton the build; setting it is part of the write's own transaction. - The scheduler spawns. A finishing worker sets its own build's flag and spawns a tick unless the registry says a scheduler already holds the build's lease (that tick will see the flag on its next poll). Every tick, at the end of each pass that acted and on every exit, asks the registry for the wake candidates: flagged builds with no live lease that were not handed out in the last ~2 minutes (or whose handed-out tick has since run and released its lease). It spawns one tick per candidate, on that build's own app. A resident build with Modal workers does the same after each result it processes.
The registry hands each build out once per window and records the hand-out, so twenty ticks asking at once produce one tick per flagged build, not twenty. There is no "which builds share a task with me" in the question — the flag already encodes relevance — so every tick is a bounded mini-watchdog for its environment, for free.
Skipping the spawn when a scheduler is live is safe only because a tick cannot exit past a flag it has not seen: it re-reads the flag once before releasing the lease (set → keep the lease, act again) and once after (set → spawn a successor). A wake-up that lands during the release either finds the lease gone and spawns, or finds it held and is picked up by that post-release read.
What this guarantees. Any recorded status change, and any freed concurrency slot, reaches every build it concerns, carried by the next scheduler pass anywhere on the deployment — seconds, while anything is running. What it does not: a change made while nothing on the deployment is ticking by something without a Modal client (the UI, a laptop-only build), and events that nobody writes at all — a worker that died without reporting, whose claim expires with nothing to notice. Those two are what the watchdog is for.
The watchdog¶
tick_watchdog is deployed on every app. It lists the running reactive
builds the app owns, spawns one tick for each, and returns — so a
sweep takes seconds whether the app is running one build or fifty, and
each build gets a container of its own with its full timeout, rather than a
share of the sweep's.
Those ticks do one pass and exit: a sweep is a safety net, not a
wake-up. A wake-up's tick lingers because something just happened and more
is likely to; a sweep looks at builds where nothing is known to have
happened. Lingering there would keep the tick function warm for the linger's
duration every period — a couple of stale RUNNING builds would be enough,
however few they are — for builds least likely to have anything to do. A
spawn for a build that already has a tick still starts a container, but that
tick finds the scheduler lease held and exits without acting.
With watchdog_period_minutes set it runs on that period; without it, it
runs when you invoke it — from the Modal UI or modal run — which is the
one-click recovery for a stalled build.
The default is off, and that is usually right: a standing sweep polls the registry whether or not anything is building, enough to keep a scale-to-zero database awake. Turn it on when leaving a build stalled for even a few minutes is unacceptable, and pick the period from how long that is — it is the recovery time for the two cases above, nothing else.
Turn it on if you run long detached tasks. Everything else in this chapter is triggered by a write: a status changes, the registry flags the builds it concerns, the next scheduler pass carries the wake-up. A claim expiring is not a write. Nobody records it, so nothing is flagged, and the recovery only happens when something looks — which the watchdog is the only thing that reliably does.
There is no way to make it self-limiting by scheduling a single wake-up at
the moment a claim is due to lapse. Modal has no delayed invocation: a
schedule is Cron or Period, both recurring and both fixed at deploy
time, and Function.spawn() takes no start time. "Wake at T" is therefore
only expressible as a container that stays alive until T — which for a
worker timeout measured in hours is the cost the reactive mode exists to
avoid — or as a recurring sweep, which is this.
So the period is a genuine trade: it is the longest a silently-dead worker can hold a claim before anything notices. A minute of sweeping against a day of a wedged task is usually the easy side of that.
The sweep itself is cheap in proportion to the environment: one pass per running build, per period, each in a container that exits as soon as it has looked.
App ownership¶
Each reactive build is owned by the app that triggered it (recorded in the registry). Only the owner's ticks drive it — a tick that reaches another app forwards the wake-up to the owner rather than running the build with the wrong code — and each app's watchdog sweeps only its own builds. A tick of the owning app that meets a build planned by an earlier deployment of the same app re-plans it under its own code (see below).
Deployments and code versions¶
In practice: Evolve a DAG Safely.
A deployment is a registry row for one code version of one app,
created by stardag modal deploy in two steps:
- Before the deploy, the CLI mints a
deployment_id(a uuid7) and records it —POST /deployments— which is when the registry assigns itsgeneration, monotonic per app. The id is baked into the Modal image as theSTARDAG_DEPLOYMENT_IDsecret every function reads. - After the deploy succeeds, the CLI marks the row live —
POST /deployments/{id}/activate— recording the Modal app id and image id the finished deploy learned. A failed activation exits non-zero: until the row is activated, no tick of the new code can plan (it exitssuperseded), and running reactive builds stay on the previous deployment.
"Current" for an app is the activated deployment with the highest
generation — order is fixed when the deploy starts, not when its record
lands, so a record that arrives late can never roll a build back to older
code. stardag modal deployments lists them, newest first, and marks each
app's current row. Nothing is kept alive beside the current deployment and
nothing needs collecting.
A local build (no stardag modal deploy involved) gets its deployment
row looked up or created at sd.build() start, keyed on
(environment, kind="local", code_id), where code_id is
STARDAG_CODE_ID if you set it, else the clean git HEAD SHA, else a
fresh, one-off id (warned) for a dirty tree. STARDAG_CODE_ID is
therefore the explicit pin when there is no git checkout to read a SHA
from — a CI image, a container built from an archive. A local deployment
row is never current and never superseded: it is authoritative for its
own plans from the moment it is created, has no separate activation step,
and the currency checks below apply to Modal deployments only — a local
driver at a new commit plans under a new scope, but the old commit's plans
stay usable rather than becoming unsealable.
A build's dependency edges belong to the instance, which belongs to a
scope — (deployment, settings), see Build &
Execution — and a
running build follows the live deployment. A reactive build progresses
by new spawns, so after a redeploy its next tick runs on the new code.
That tick finds the build's active plan names another deployment_id and
re-plans: it rehydrates the plan's root instances under its own code,
compares their task ids against the build's recorded roots, then runs the
static phase and seals under its own scope
((this deployment, the plan's settings)) — reusing a plan for that scope
if one already exists and is sealed. You see rolled_over in that tick's
summary. Discovery stops at completed tasks, so a redeploy costs one walk
of the incomplete part of each running DAG.
Three things follow from that:
- A tick still lingering on the old deployment sees the plan superseded
and exits (
superseded); the scheduler lease already guarantees one driver per build. - Workers are code-agnostic — the task id promises the output whatever code produces it — but each registers the dynamic dependencies it yields against its own plan and deployment. An old container's late yield is accepted into the superseded plan (a true fact about that scope, useful to any scope-mate), the rolled-over build's own instance has no dynamic edges yet, and its next tick restarts the parent under the new code: the old children keep running as accepted duplicate work.
- Executions the new plan no longer contains finish on their own; their targets are content-addressed, so they harm nothing.
A rollover only moves forward: the tick asks the registry which
deployment is current for its app and re-plans (or seals) only if that is
its own — checked again at /seal, under the same per-app lock deploy
activation uses, so two ticks racing under two new deployments cannot both
win. A tick of an older deployment that wins the scheduler lease late
exits superseded instead of moving the build back. That is why
stardag modal deploy records every deploy and exits non-zero if it
cannot. What makes the re-plan safe is that a task is rebuilt from the
registry's stored instance body under the new deployment's own code, so it
carries nothing from the code that planned the build. What cannot roll
over fails the build — a root whose task id changed under the new code, or
a task the new deployment cannot rebuild (its class gone, or outside its
task_modules) — and the remedy is a new build.
One process, one build's settings, for the lifetime of the process. Every process of a build — the bootstrap, each tick, each worker, a resident driver — applies the build's settings as environment variables for as long as it is serving that build; see Integrate with Modal for the one-input-per-container consequence this has for deployed ticks and workers.
Time-based recovery is the watchdog's job, and only the watchdog's. Everything else in this chapter — a wake-up, a rollover — is triggered by a write landing in the registry. A claim simply expiring is not a write: nobody records it, so nothing is flagged, and the only thing that notices is a periodic sweep. See the watchdog above for why that is a deliberate trade rather than a gap to be closed with a smarter wake-up.
A branch deployment is simply another app: give it its own name and it has its own single live version.
Choosing¶
| you have | use |
|---|---|
| many short tasks, one machine or one container | resident (sd.build, or the build function) |
| long or preemptible tasks, unattended runs, cross-build limits | reactive |
| a laptop driver with GPU tasks in Modal | resident with RoutedTaskExecutor — a hybrid run; it wakes reactive neighbours like a tick does |
Everything on this page beyond the first section requires a registry.