Temporal batch orchestration — SHIPPED (live-verified vs a real dev Temporal)
Status: implemented per this design and live-verified 2026-07-08 — worker SIGKILL mid-batch → a new worker replays the history and finishes the workflow (W1); control-plane SIGKILL mid-batch → the activities retry against the restarted CP, boot recovery respects workflow ownership ("batches resumed 1", no double-drive), 24/24 cases complete with zero loss or duplication (W2). Opt-in via EVERDICT_TEMPORAL_ADDRESS (default-on once set; EVERDICT_TEMPORAL_BATCHES=0 opts out); a failed workflow START degrades gracefully to the in-process loop.
Why (and why not yet)
Today a batch's driver loop (ScorecardService.track) lives in-process in one control plane.
The resilience layer (docs/architecture/batch-resilience.md) makes that survivable — results persist
per case, boot recovery resumes interrupted batches, retry-failed re-runs casualties — so a single
control plane now rides out its own restarts and cluster incidents. What in-process tracking cannot
give: more than one control plane (HA / rolling deploys with zero batch ownership gaps),
driver-loop survival independent of any process, and per-step observability of a batch as a
first-class workflow. That is Temporal's exact shape, and the worker/activity split already exists
(@everdict/orchestrator: workflow = deterministic, activity = dispatchCase).
The resilience layer was built first deliberately: it is the fallback story when Temporal is not deployed (self-hosters, dev), and the seeded track loop it produced is precisely the state machine the workflow needs.
Shape
One batch = one workflow (scorecardBatch), one case = one activity (dispatchCase).
scorecardBatchWorkflow(input: {scorecardId, tenant, dataset ref, harness ref, judges, judge,
runtimes[], concurrency, retries, seed caseIds})
├─ activity resolveBatch() → cases minus seeds (the same seeded-loop inputs, resolved fresh)
├─ for case in cases (bounded by concurrency, workflow-side semaphore):
│ activity dispatchCase(job) → CaseResult (retry policy: ONLY when classifyFailure().retryable —
│ the activity rethrows fatal classes as non-retryable ApplicationFailure)
│ activity settleCase(result) → child-run write-back + judge stream push + export push
└─ activity finalizeBatch() → aggregate/summarize/persist/notify
- Determinism: the workflow holds only ids and counters; every I/O (registry resolve, dispatch,
judge, persist) is an activity — same rule the repo already enforces for
workflows.ts. Failure classes map to Temporal retry policies— this half of the design was deliberately NOT implemented, and the reversal matters when justifying Temporal. Case-level retry stayed CP-side (workflow-batch-driver.ts, thefor (attempt…)loop overfailure.retryable): the workflow's generous activity retry is for TRANSPORT failures (the control plane unreachable), never for eval semantics. So "Temporal gives us retry" is NOT a reason this design holds — the reasons are durable timers, cross-restart progress, and (below) batch ownership across MORE THAN ONE control plane. Keeping the retry story here would have split the failure taxonomy across two engines.- Resume for free: workflow history replaces
orchestration+child-run reconstruction. Boot recovery stays for the non-Temporal deployment; whenEVERDICT_TEMPORAL_ADDRESSis set, submit starts the workflow instead ofvoid this.track(...)and recovery skips Temporal-owned batches (they own themselves). - Supersede / cancel → workflow cancellation (cooperative, same semantics as the AbortController).
- Sharding stays submit-side (per-case
placement.targetround-robin) — the workflow doesn't care where a case lands; the Scheduler/RuntimeDispatcher path is unchanged insidedispatchCase.
Increment plan (next slice)
scorecardBatchWorkflow+ activities in@everdict/orchestratorbeside the existing worker; reuseexecuteCase/ScoringServiceas activity bodies (no logic forks — the service methods are already transport-free).ScorecardService.submit:temporalDriver?dep — when configured, start the workflow with the same persistedorchestrationinputs;trackstays as the in-process driver otherwise.- Live e2e vs the real dev Temporal (
temporalio/temporal— the scheduled-evals harness already drives one:scripts/live/scheduled-pinch-temporal.mjs): kill the WORKER mid-batch → a new worker picks the workflow up with zero lost cases; kill the control plane → same. - Queue view: Temporal-owned batches surface
workflowIdon the record for deep-linking.
History budget — continue-as-new
One case ≈ a handful of history events (activity scheduled/started/completed × transport retries), so an
unbounded multi-thousand-case batch would walk into Temporal's per-execution history limits (50K events /
50MB). The workflow processes at most continueEvery cases per execution (default 500;
EVERDICT_TEMPORAL_BATCH_CONTINUE_EVERY on the CP feeds the start args) and then continueAsNews with the
same input. planBatch's idempotence (unfinished-only) is what makes this trivially correct — the continued
execution re-plans and receives exactly the remainder, with a fresh history, under the same workflowId.
Live e2e: 12-case batch with continueEvery=5 → plan steps "Running 12" → "Running 7 (5 kept)" →
"Running 2 (10 kept)"; temporal workflow list shows ContinuedAsNew ×2 → Completed on one workflowId;
12/12 pass.
Adaptive rotation (history pressure). The fixed case-count slice assumes ~a handful of events per case,
but activity transport retries inflate events-per-case — a flaky network can walk a 500-case slice into the
history limits anyway. The workflow therefore ALSO rotates when the server itself suggests it
(workflowInfo().continueAsNewSuggested) or when historyLength crosses a floor
(rotateAtHistoryLength, default 20 000; dial EVERDICT_TEMPORAL_BATCH_ROTATE_HISTORY): lanes stop TAKING
new cases, in-flight activities drain, and the execution continues-as-new — planBatch's idempotent re-plan
makes an early rotation harmless. Live: a 24-case batch with the floor at 40 ran as a 4-execution chain
(~56–59 events each) instead of one 158-event history, 24/24 succeeded.
Retry parity
retry-failed batches are workflow-owned too (a CP restart mid-retry must not lose them): the carried seeds
(passes + re-collected recoveries) are MATERIALIZED as succeeded child runs before the workflow starts, so the
idempotent planBatch naturally drives only the re-dispatch remainder and finalize aggregates everything. The
OOM escalation boost rides origin.memoryBoostMb → batch context → runBatchCase, identical to the in-process
loop. Start failure degrades to the in-process loop, same as submit. Live: an OOM retry chain ran entirely
workflow-owned (each retry its own Completed workflow), escalated 64 → 128 → 256 and passed.
Non-goals
- Moving judge/export streaming into Temporal (they are per-case activities already chained after dispatch — no barrier to remove).
- A second tenancy/fairness layer inside Temporal: WFQ/capacity stay in the Scheduler; Temporal owns durability of the driver loop, not placement.