Data platform · Case study
Netflix: keep scheduled jobs on track
Take each failed or late job from first alert to a rerun, a fix by its owner or a recorded decision to leave it.
By OneShot · Updated
Read the solution ↓What the workflow produces
Job incident record- Failure
- Run, error, last good run and what changed between them.
- Action taken
- Reruns requested, their budget and their results.
- Owner decision
- Acknowledging team, fix or decision to pause the job.
- Outcome
- Next successful run and downstream jobs released.
The documented problem
Netflix describes Maestro as a general-purpose workflow orchestrator offered as a managed service to its data platform users. The project’s documentation says it serves thousands of users, among them data scientists, data engineers, machine learning engineers, content producers and business analysts.
According to Netflix, Maestro schedules hundreds of thousands of workflows and millions of jobs every day, and operates with a strict service-level objective even during traffic spikes. The source describes the scheduler, not the follow-through on an individual failed run. The OneShot workflow below is a proposed design for that follow-through; it is not part of Maestro, and Netflix is not described here as using it.
- Netflix Maestro project documentationPrimary source · Checked 2026-10-08
The OneShot solution
Keeping scheduled jobs healthy at scale
Keep every failed or late job moving to a rerun, an owner’s fix or an explicit decision, so no broken run sits unnoticed.
- 01
Open the incident record
Read the failed run, its logs and the last successful run of the same job. Record what changed between them and which downstream jobs are waiting.
- 02
Rerun within limits
The OneShot agent requests a rerun when the failure matches a cause the owning team has marked safe to retry, within the job’s approved compute budget. Anything else waits for the owner.
- 03
Bring in the owner
Send the owning team the failure evidence, what was already tried and which downstream work is blocked. Chase an unacknowledged incident before the next scheduled run.
- Email send
- Email Inbox
- Workflow schedules and job records
- 04
Record the resolution
Confirm that the next run succeeds. Store the cause, the fix and the owner’s decision against the job; reopen the record if the same failure returns.
What counts as complete
The job has a successful run after the fix, or its owner has recorded a decision to pause or retire it.
Decisions and exceptions
- A rerun that would exceed the approved budget waits for the owner’s approval.
- Failures with an unknown cause are never retried automatically.
- Changes to pipeline code stay with the owning team.
What to measure
- Time from failure to acknowledgement
- Reruns that succeed first time
- Repeat failures per job
- Downstream delay
Establish the baseline and review period before launch. Compare completed cases, unresolved work and reviewer corrections against the same scope.
The agent’s tool basket
OneShot gives its agents the tools to analyze records, communicate, research and act on authorized business data. The platform carries the workflow from the first action through to the recorded result.
Workflow schedules and job records
Read job runs, logs and ownership; request reruns the owning team has marked safe to retry.
Workflow pricing
Scope the full workflow around your volume, business systems and required outcome.
Discuss workflow pricing →