Skip to content

Exceptions

Stardag exceptions for error handling.

Exception Hierarchy

StardagError
├── APIError
│   ├── AuthenticationError
│   ├── AuthorizationError
│   ├── NotFoundError
│   └── TokenExpiredError
├── ResumableInterruption
├── ExecutionCancelled
└── ...

Base Exception

StardagError

from stardag import StardagError

Base exception for all Stardag errors.

try:
    sd.build(task)
except StardagError as e:
    print(f"Stardag error: {e}")

API Exceptions

APIError

from stardag import APIError

Base exception for API-related errors.

AuthenticationError

from stardag import AuthenticationError

Raised when authentication fails:

  • Invalid credentials
  • Missing API key
  • OAuth flow failure

Handling:

try:
    registry = APIRegistry()
except AuthenticationError:
    print("Please login: stardag auth login")

AuthorizationError

from stardag import AuthorizationError

Raised when authenticated but not authorized:

  • Insufficient permissions
  • Wrong workspace/environment
  • Resource access denied

NotFoundError

from stardag.exceptions import NotFoundError

Raised on a 404: a build, plan, task or deployment that does not exist in the environment (code names which, e.g. unknown_plan), or a route the registry does not serve.

The last case is what an SDK and a registry from different release lines look like. There is no version check in either direction: the SDK and the registry are upgraded together, and a v2 SDK against a v1 registry (or the reverse) fails on its first call with a NotFoundError whose detail is FastAPI's "Not Found". stardag.exceptions.is_missing_route_error(e) tells that apart from a missing resource. Upgrade the other side.

TokenExpiredError

from stardag import TokenExpiredError

Raised when authentication token has expired:

Handling:

try:
    sd.build(task, registry=registry)
except TokenExpiredError:
    # Refresh token and retry
    os.system("stardag auth refresh")

ResumableInterruption

from stardag import ResumableInterruption

The one exception you raise rather than catch. It says: I saved my progress, run me again.

import stardag as sd
from stardag.integration.modal import MODAL_INTERRUPTIONS


class TrainModel(sd.TargetTask[sd.DirectoryTarget]):
    def target(self) -> sd.DirectoryTarget:
        return sd.get_directory_target(sd.get_default_relpath(self))

    def run(self):
        directory = self.target()
        checkpoint = directory / "checkpoint.json"
        try:
            train(resume_from=checkpoint)
        except MODAL_INTERRUPTIONS:        # preemption OR the function timeout
            save_checkpoint(checkpoint)
            raise sd.ResumableInterruption("checkpointed") from None
        directory.mark_done()

An interruption you do not catch is a failure, deliberately. Letting one propagate means the task had no plan for it — it hung, or the worker's timeout is too small — and both want the same answer: fail, under the scheduler's ordinary attempt budget. So there is no setting anywhere deciding whether a timeout was "expected"; the task answers by raising this, or by not raising it.

Catch the interruption types, never BaseException

A NameError is a BaseException too. A blanket catch sweeps up ordinary bugs, and re-raising ResumableInterruption for one turns a deterministic failure into a task that resumes until its budget runs out. Use MODAL_INTERRUPTIONS (exactly KeyboardInterrupt and modal.exception.InputCancellation).

except KeyboardInterrupt: is wrong the other way: InputCancellation is not a KeyboardInterrupt, so it misses timeouts entirely.

ResumableInterruption is an ordinary Exception, not a BaseException: you raise it from inside your own error handling, where a BaseException subclass would be one more thing slipping past your control flow.

What happens next depends on whether a restart is still possible, which the Modal runner reads off the interruption you caught — it is still on the exception you raise, and raise ... from None keeps it there (that form hides the "During handling…" preamble; it does not discard the original).

Caught a preemption, the runner re-raises an interrupt in its place so the backend sees a crashed container and restarts the input on the same call id, and records the preemption so a restart that never arrives is visible. Caught a function timeout or a cancel — when no restart is coming — it records an interruption for a scheduler tick to act on instead.

So raise it from inside the except block. raise sd.ResumableInterruption(...) from None keeps the link, so does the same statement with no from clause, and so does an explicit from err on a saved exception. (A bare raise is not one of these: it re-raises the platform exception, so the runner never sees a resumption request and records nothing — on a preemption the backend restarts the input anyway, but on a timeout or a cancel the execution dies and a later tick records a retryable failure.) What loses the link is raising where the interruption is no longer reachable — outside the block with no explicit cause, or on a condition of your own. Then stardag falls back to comparing elapsed time against the worker's declared timeout, which is a guess on a clock that starts after the container does.

Resumption is bounded by TickConfig.max_interruptions (default 20), a budget separate from max_attempts — see Preemption and timeouts.

ExecutionCancelled

from stardag import ExecutionCancelled

The other exception you raise rather than catch. It says: this execution is no longer wanted; stop without producing anything.

Stardag raises it for you at the two automatic cooperative-cancellation checkpoints — the start of each attempt, and each dynamic-dependency yield. You raise it where only you know a stop is safe:

import stardag as sd

for chunk in chunks:
    if sd.cancellation_requested():
        raise sd.ExecutionCancelled()
    process(chunk)

The worker treats it as a clean exit, not a task failure: no output is written, no completion is reported, and no end-of-attempt event is recorded. What it does not do is return normally, and that is deliberate — a backend call that succeeds with no output is read by a scheduler as "the worker wrote it, eventual consistency", which would record a completion for a target that does not exist.

Ask sd.cancellation_requested() first rather than raising speculatively: it is throttled, and it answers True only when the registry positively said this execution has been superseded, its task cancelled, or its build stopped running. See Cancelling work.

Like ResumableInterruption, it is an ordinary Exception rather than a BaseException, so it does not slip past your own error handling. A task that catches it should re-raise.

Common Error Scenarios

Target Root Not Configured

# Error: No target root configured for 'default'
# Solution:
export STARDAG_TARGET_ROOTS__DEFAULT=/path/to/outputs

Task Not Complete

# A dependency failed to build
try:
    sd.build(task)
except Exception as e:
    # Check task completion status
    print(task.complete())  # False

Serialization Error

# Output type cannot be serialized
# Ensure return type is JSON-serializable or use pickle
@sd.task
def my_task() -> dict:  # JSON-serializable
    return {"key": "value"}

Best Practices

  1. Catch specific exceptions - Handle AuthenticationError differently from APIError
  2. Log error details - Exceptions contain useful debugging info
  3. Graceful degradation - Fall back to local builds if API unavailable
from stardag import APIError, AuthenticationError

try:
    sd.build(task, registry=registry)
except AuthenticationError:
    print("Auth failed - running locally")
    sd.build(task)
except APIError as e:
    print(f"API error: {e} - running locally")
    sd.build(task)