Exceptions¶
Stardag exceptions for error handling.
Exception Hierarchy¶
StardagError
├── APIError
│ ├── AuthenticationError
│ ├── AuthorizationError
│ ├── NotFoundError
│ └── TokenExpiredError
├── ResumableInterruption
├── ExecutionCancelled
└── ...
Base Exception¶
StardagError¶
Base exception for all Stardag errors.
API Exceptions¶
APIError¶
Base exception for API-related errors.
AuthenticationError¶
Raised when authentication fails:
- Invalid credentials
- Missing API key
- OAuth flow failure
Handling:
AuthorizationError¶
Raised when authenticated but not authorized:
- Insufficient permissions
- Wrong workspace/environment
- Resource access denied
NotFoundError¶
Raised on a 404: a build, plan, task or deployment that does not exist in
the environment (code names which, e.g. unknown_plan), or a route the
registry does not serve.
The last case is what an SDK and a registry from different release lines
look like. There is no version check in either direction: the SDK and the
registry are upgraded together, and a v2 SDK against a v1 registry (or the
reverse) fails on its first call with a NotFoundError whose detail is
FastAPI's "Not Found". stardag.exceptions.is_missing_route_error(e)
tells that apart from a missing resource. Upgrade the other side.
TokenExpiredError¶
Raised when authentication token has expired:
Handling:
try:
sd.build(task, registry=registry)
except TokenExpiredError:
# Refresh token and retry
os.system("stardag auth refresh")
ResumableInterruption¶
The one exception you raise rather than catch. It says: I saved my progress, run me again.
import stardag as sd
from stardag.integration.modal import MODAL_INTERRUPTIONS
class TrainModel(sd.TargetTask[sd.DirectoryTarget]):
def target(self) -> sd.DirectoryTarget:
return sd.get_directory_target(sd.get_default_relpath(self))
def run(self):
directory = self.target()
checkpoint = directory / "checkpoint.json"
try:
train(resume_from=checkpoint)
except MODAL_INTERRUPTIONS: # preemption OR the function timeout
save_checkpoint(checkpoint)
raise sd.ResumableInterruption("checkpointed") from None
directory.mark_done()
An interruption you do not catch is a failure, deliberately. Letting one propagate means the task had no plan for it — it hung, or the worker's timeout is too small — and both want the same answer: fail, under the scheduler's ordinary attempt budget. So there is no setting anywhere deciding whether a timeout was "expected"; the task answers by raising this, or by not raising it.
Catch the interruption types, never BaseException
A NameError is a BaseException too. A blanket catch sweeps up
ordinary bugs, and re-raising ResumableInterruption for one turns a
deterministic failure into a task that resumes until its budget runs
out. Use MODAL_INTERRUPTIONS (exactly KeyboardInterrupt and
modal.exception.InputCancellation).
except KeyboardInterrupt: is wrong the other way: InputCancellation
is not a KeyboardInterrupt, so it misses timeouts entirely.
ResumableInterruption is an ordinary Exception, not a BaseException:
you raise it from inside your own error handling, where a BaseException
subclass would be one more thing slipping past your control flow.
What happens next depends on whether a restart is still possible, which the
Modal runner reads off the interruption you caught — it is still on the
exception you raise, and raise ... from None keeps it there (that form
hides the "During handling…" preamble; it does not discard the original).
Caught a preemption, the runner re-raises an interrupt in its place so the backend sees a crashed container and restarts the input on the same call id, and records the preemption so a restart that never arrives is visible. Caught a function timeout or a cancel — when no restart is coming — it records an interruption for a scheduler tick to act on instead.
So raise it from inside the except block. raise
sd.ResumableInterruption(...) from None keeps the link, so does the same
statement with no from clause, and so does an explicit from err on a
saved exception. (A bare raise is not one of these: it re-raises the
platform exception, so the runner never sees a resumption request and
records nothing — on a preemption the backend restarts the input anyway,
but on a timeout or a cancel the execution dies and a later tick records a
retryable failure.) What loses the link is raising where the interruption is no
longer reachable — outside the block with no explicit cause, or on a
condition of your own. Then stardag falls back to comparing elapsed time
against the worker's declared timeout, which is a guess on a clock that
starts after the container does.
Resumption is bounded by TickConfig.max_interruptions (default 20), a
budget separate from max_attempts — see
Preemption and timeouts.
ExecutionCancelled¶
The other exception you raise rather than catch. It says: this execution is no longer wanted; stop without producing anything.
Stardag raises it for you at the two automatic cooperative-cancellation checkpoints — the start of each attempt, and each dynamic-dependency yield. You raise it where only you know a stop is safe:
import stardag as sd
for chunk in chunks:
if sd.cancellation_requested():
raise sd.ExecutionCancelled()
process(chunk)
The worker treats it as a clean exit, not a task failure: no output is written, no completion is reported, and no end-of-attempt event is recorded. What it does not do is return normally, and that is deliberate — a backend call that succeeds with no output is read by a scheduler as "the worker wrote it, eventual consistency", which would record a completion for a target that does not exist.
Ask sd.cancellation_requested() first rather than raising
speculatively: it is throttled, and it answers True only when the
registry positively said this execution has been superseded, its task
cancelled, or its build stopped running. See
Cancelling work.
Like ResumableInterruption, it is an ordinary Exception rather than a
BaseException, so it does not slip past your own error handling. A task
that catches it should re-raise.
Common Error Scenarios¶
Target Root Not Configured¶
# Error: No target root configured for 'default'
# Solution:
export STARDAG_TARGET_ROOTS__DEFAULT=/path/to/outputs
Task Not Complete¶
# A dependency failed to build
try:
sd.build(task)
except Exception as e:
# Check task completion status
print(task.complete()) # False
Serialization Error¶
# Output type cannot be serialized
# Ensure return type is JSON-serializable or use pickle
@sd.task
def my_task() -> dict: # JSON-serializable
return {"key": "value"}
Best Practices¶
- Catch specific exceptions - Handle
AuthenticationErrordifferently fromAPIError - Log error details - Exceptions contain useful debugging info
- Graceful degradation - Fall back to local builds if API unavailable