All case studies
Netflix

Data platform · Case study

Netflix: keep scheduled jobs on track

Take each failed or late job from first alert to a rerun, a fix by its owner or a recorded decision to leave it.

By OneShot · Updated

Read the solution ↓

What the workflow produces

Job incident record
Failure
Run, error, last good run and what changed between them.
Action taken
Reruns requested, their budget and their results.
Owner decision
Acknowledging team, fix or decision to pause the job.
Outcome
Next successful run and downstream jobs released.

The documented problem

Netflix describes Maestro as a general-purpose workflow orchestrator offered as a managed service to its data platform users. The project’s documentation says it serves thousands of users, among them data scientists, data engineers, machine learning engineers, content producers and business analysts.

According to Netflix, Maestro schedules hundreds of thousands of workflows and millions of jobs every day, and operates with a strict service-level objective even during traffic spikes. The source describes the scheduler, not the follow-through on an individual failed run. The OneShot workflow below is a proposed design for that follow-through; it is not part of Maestro, and Netflix is not described here as using it.

The OneShot solution

Keeping scheduled jobs healthy at scale

Keep every failed or late job moving to a rerun, an owner’s fix or an explicit decision, so no broken run sits unnoticed.

Work enters the queue
A failed run, a job past its expected finish or a rerun that failed again.
Decision owner
On-call engineer for the owning team
  1. 01

    Open the incident record

    Read the failed run, its logs and the last successful run of the same job. Record what changed between them and which downstream jobs are waiting.

  2. 02

    Rerun within limits

    The OneShot agent requests a rerun when the failure matches a cause the owning team has marked safe to retry, within the job’s approved compute budget. Anything else waits for the owner.

  3. 03

    Bring in the owner

    Send the owning team the failure evidence, what was already tried and which downstream work is blocked. Chase an unacknowledged incident before the next scheduled run.

  4. 04

    Record the resolution

    Confirm that the next run succeeds. Store the cause, the fix and the owner’s decision against the job; reopen the record if the same failure returns.

What counts as complete

The job has a successful run after the fix, or its owner has recorded a decision to pause or retire it.

Decisions and exceptions

  • A rerun that would exceed the approved budget waits for the owner’s approval.
  • Failures with an unknown cause are never retried automatically.
  • Changes to pipeline code stay with the owning team.

What to measure

  • Time from failure to acknowledgement
  • Reruns that succeed first time
  • Repeat failures per job
  • Downstream delay

Establish the baseline and review period before launch. Compare completed cases, unresolved work and reviewer corrections against the same scope.

The agent’s tool basket

OneShot gives its agents the tools to analyze records, communicate, research and act on authorized business data. The platform carries the workflow from the first action through to the recorded result.

  • Workflow schedules and job records

    Read job runs, logs and ownership; request reruns the owning team has marked safe to retry.

Workflow pricing

Scope the full workflow around your volume, business systems and required outcome.

Discuss workflow pricing →

Bring a workflow
like this one.