Skip to content

[core] Detect wedged waits instead of wake-looping on them forever - #3541

Open
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection
Open

[core] Detect wedged waits instead of wake-looping on them forever#3541
pranaygp wants to merge 1 commit into
mainfrom
pgp/wait-wedge-detection

Conversation

@pranaygp

Copy link
Copy Markdown
Contributor

Root cause

On worlds that persist the wait entity and its event-log row in separate, non-transactional writes (world-vercel / DynamoDB), a request can commit the wait entity and then fail before the event-row insert (crash, dropped connection). Every retry of the event write then conflicts (409) against the committed entity, while the log stays permanently short one row. The SDK swallows the conflict as "my write already landed" — which is half-true: the entity landed, the row didn't.

The consequences are an invisible infinite loop, not an error:

  • sleep() resolves only from a wait_completed row (workflow/sleep.ts), and the elapsed-wait pass can only complete waits whose wait_created row it can read.
  • So the run replays into the same conflict forever. Once past resumeAt, every pass arms a fresh ~1s wake (the near-elapsed continuation key is second-bucketed, so dedup never collapses them — runtime/wait-continuation.ts).
  • The run sits in running forever, burning an invocation per second, with nothing but an info-level "already exists, skipping" log line.

Steps and runs had the same wedge class and got server-side recovery (workflow-server #704, #707); waits are the remaining unhealed sibling. A companion workflow-server PR makes wait writes transactional and backfills existing wedges; this PR is the SDK-side detection so the contradiction is loud while it persists and terminal once it is provable.

What this does

wait_completed (elapsed-wait pass, runtime.ts): when the create conflicts AND the follow-up reload still cannot produce the row — the server says "completed", the log says "pending" — log a warning and report workflow.wait.wedge_suspected on the invocation span. Once the clock is more than the threshold past the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG (same terminal path as the slot-gap check). The benign race (conflicting row IS readable after reload) stays silent exactly as before.

wait_created (suspension handler): resumeAt cannot anchor this site — an uncreated wait recomputes it from the live clock on every replay, so it always sits in the future. The anchor is the scheduling instant embedded in the wait's replay-stable correlation id (seeded RNG + replay clock ⇒ same ULID every replay). Within the threshold, behavior is unchanged (silent info — creation conflicts are the ordinary concurrent-suspension race). Past it, the contradiction is verified against a fresh event-log read before failing, so a concurrent writer's row landing after this replay's snapshot can never be mistaken for a wedge.

Why stateless, time-based escalation: every wake of the loop is a fresh queue message (fresh delivery attempt = 1), so there is no attempt counter to persist across invocations. "How long has this contradiction persisted against a replay-stable time anchor" is derivable on every observation, and a healthy wait completes within seconds of its target.

Threshold

WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS, default 600 (10 minutes), documented in docs/content/docs/v5/configuration/runtime-tuning.mdx next to the other wait tunables. The generous default means eventually-consistent read staleness cannot plausibly trigger a failure; the wedge, once real, is permanent — 10 minutes only bounds how long the loop burns invocations.

Failure shape

Reuses CorruptedEventLogErrorrun_failed with errorCode: CORRUPTED_EVENT_LOG (no new error code; the log genuinely cannot produce a row the World attests exists, which is this code's meaning, and it flows through existing classification, dashboards, and error docs). The suspension-handler throw required one gate change in runtime.ts: the suspension-error catch now routes CorruptedEventLogError to its terminal fail-the-run path alongside FatalError, instead of rethrowing for a redelivery that would replay into the same conflict.

Tests

  • runtime/wait-wedge.test.ts — unit: threshold classification + env override, ULID anchor decoding, fresh-read verification (found / missing / fail-open on read errors).
  • runtime/wait-wedge-detection.test.ts — drives the real queue handler with a fake World (same harness pattern as wait-completion-replay.test.ts) through both wedges: benign concurrent-winner races stay silent and the run completes; contradictions inside the threshold warn and keep retrying (wake continuation still armed); contradictions past the threshold fail the run with CORRUPTED_EVENT_LOG.
  • Full core suite: 97 files, 2135 passed, 3 expected-fail (no regressions).

🤖 Generated with Claude Code

On worlds that persist the wait entity and its event-log row in separate
writes (world-vercel), a request can commit the entity and then fail
before the row insert. Every retry of the event write then conflicts
(409) against the committed entity while the log stays permanently short
one row. sleep() resolves only from a wait_completed row and the
elapsed-wait pass can only complete waits whose wait_created row it can
read, so the run replays into the same conflict forever: a ~1s wake loop
that never errors and never completes.

Make the contradiction loud, and terminal past a generous threshold:

- wait_completed (elapsed-wait pass): when the create conflicts AND the
  follow-up reload still cannot produce the row, warn and report
  workflow.wait.wedge_suspected on the invocation span; once the clock
  is more than WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS (default 600) past
  the wait's resumeAt, fail the run as CORRUPTED_EVENT_LOG.
- wait_created (suspension handler): resumeAt cannot anchor this site
  (an uncreated wait recomputes it from the live clock every replay), so
  the anchor is the scheduling instant embedded in the wait's
  replay-stable correlation id. Past the threshold the contradiction is
  verified against a fresh event-log read before failing, so a
  concurrent writer's row landing after this replay's snapshot is never
  mistaken for a wedge.

Escalation is stateless on purpose: every wake is a fresh queue message,
so there is no attempt counter to persist — but "how long has this
contradiction persisted against a replay-stable anchor" is derivable on
every observation. Benign concurrent-handler races (the conflicting row
is readable) stay silent exactly as before.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
@pranaygp
pranaygp requested review from a team, fantix and msullivan as code owners August 14, 2026 01:08
Copilot AI lite review requested due to automatic review settings August 14, 2026 01:08
@changeset-bot

changeset-bot Bot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f81f331

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/core Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview Aug 14, 2026 1:11am
example-nextjs-workflow-webpack Ready Ready Preview Aug 14, 2026 1:11am
example-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-astro-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-express-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-fastify-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-hono-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-nestjs-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-nitro-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-nuxt-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-python-workflow Error Error Aug 14, 2026 1:11am
workbench-sveltekit-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-tanstack-start-workflow Ready Ready Preview Aug 14, 2026 1:11am
workbench-vite-workflow Ready Ready Preview Aug 14, 2026 1:11am
workflow-docs Ready Ready Preview, v0 Aug 14, 2026 1:11am
workflow-swc-playground Ready Ready Preview Aug 14, 2026 1:11am
workflow-tarballs Ready Ready Preview Aug 14, 2026 1:11am
workflow-web Ready Ready Preview Aug 14, 2026 1:11am

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3466 0 590 4056
✅ 💻 Local Development 3536 0 520 4056
✅ 📦 Local Production 3810 0 558 4368
✅ 🐘 Local Postgres 3810 0 558 4368
✅ 🪟 Windows 312 0 0 312
✅ vercel-multi-region 27 0 0 27
Total 14961 0 2226 17187
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 128 0 28
✅ astro-quickjs 128 0 28
✅ example-node 128 0 28
✅ example-quickjs 128 0 28
✅ express-node 128 0 28
✅ express-quickjs 128 0 28
✅ fastify-node 128 0 28
✅ fastify-quickjs 128 0 28
✅ hono-node 128 0 28
✅ hono-quickjs 128 0 28
✅ nest-node 128 0 28
✅ nest-quickjs 128 0 28
✅ nextjs-turbopack-node 153 0 3
✅ nextjs-turbopack-quickjs 153 0 3
✅ nextjs-webpack-node 153 0 3
✅ nextjs-webpack-quickjs 153 0 3
✅ nitro-node 128 0 28
✅ nitro-quickjs 128 0 28
✅ nuxt-node 128 0 28
✅ nuxt-quickjs 128 0 28
✅ sveltekit-node 147 0 9
✅ sveltekit-quickjs 147 0 9
✅ tanstack-start-node 128 0 28
✅ tanstack-start-quickjs 128 0 28
✅ vite-node 128 0 28
✅ vite-quickjs 128 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 156 0 0
✅ nextjs-turbopack-quickjs 156 0 0

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit f81f331 · Fri, 14 Aug 2026 01:30:20 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1076 (+182%) 🔻 1347 🔴 (+21%) 🔻 1375 🔴 (+21%) 🔻 1405 🔴 (-8.3%) 30
TTFS stream 1285 (+28%) 🔻 1340 🔴 (+27%) 🔻 1351 🔴 (+26%) 🔻 1471 🔴 (+33%) 🔻 30
TTFS hook + stream 1553 (+22%) 🔻 1635 🔴 (+18%) 🔻 1650 🔴 (+16%) 🔻 1703 🔴 (+5.1%) 30
Fan-out TTFS Promise.all(100 steps) 8809 (-1.2%) 10727 (+7.8%) 14678 (+46%) 🔻 15082 (+12%) 10
Fan-out TTLS Promise.all(100 steps) 17196 (-2.7%) 20436 (+8.3%) 23214 (+22%) 🔻 23631 (+0.8%) 10
STSO 1020 steps (inline) 124 (+0.8%) 174 (-8.9%) 197 (-14%) 391 (-33%) 💚 1019
WO 1020 steps 171617 (-12%) 171617 (-12%) 171617 (-12%) 171617 (-12%) 1
SL stream latency 78 (-1.3%) 105 🔴 (-4.5%) 117 🔴 (-9.3%) 129 🔴 (-62%) 💚 30
SO stream overhead (text) 97 (-13%) 141 (-22%) 💚 157 (-24%) 💚 314 (-48%) 💚 30
SO stream overhead (structured) 101 (+5.2%) 154 (-1.3%) 194 (+16%) 🔻 8083 🔴 (+4341%) 🔻 30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 171408ms (Δ -22960ms, -12%)

  100-150 ms  ███████░░░┃               main 180  this 295  +115
  150-200 ms  ███████████████████████┃  main 627  this 631    +4
  200-250 ms  █┃███                     main 134  this  57   -77
  250-300 ms  ┃                         main  29  this  14   -15
  300-350 ms  ┃                         main  15  this   7    -8
  350-400 ms  ┃                         main  11  this   5    -6
  400-450 ms  ┃                         main   4  this   5    +1
  450-500 ms  ┃                         main   5  this   3    -2
  500-550 ms  ┃                         main   3  this   0    -3
  550-600 ms  ┃                         main   1  this   2    +1
  600-650 ms  ┃                         main   5  this   0    -5
  650-700 ms  ┃                         main   1  this   0    -1
  750-800 ms  ┃                         main   1  this   0    -1
  800-850 ms  ┃                         main   1  this   0    -1
1100-1150 ms  ┃                         main   1  this   0    -1
4450-4500 ms  ┃                         main   1  this   0    -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds SDK-side detection for “wedged waits” (409 conflict on wait event writes where the corresponding event-log row is never readable), so runs stop silently wake-looping forever and instead warn for a configurable window before failing as CORRUPTED_EVENT_LOG. This fits into @workflow/core runtime durability/corruption detection, complementing server-side recovery for related wedge classes.

Changes:

  • Introduces stateless, time-anchored wait-wedge classification and error messaging (runtime/wait-wedge.ts) with a tunable threshold (WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS).
  • Adds runtime integration at both wedge sites (wait_completed in runtime.ts, wait_created in suspension-handler.ts), including telemetry reporting (workflow.wait.wedge_suspected).
  • Adds unit + queue-handler integration tests and documents the new environment variable.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
packages/core/src/telemetry/semantic-conventions.ts Adds the workflow.wait.wedge_suspected semantic convention for span reporting.
packages/core/src/runtime/wait-wedge.ts New wedge detection utilities: thresholding, ULID anchor decoding, fresh-read verification, shared error message.
packages/core/src/runtime/wait-wedge.test.ts Unit tests for classification, env override behavior, ULID decoding, and verification-read behavior.
packages/core/src/runtime/wait-wedge-detection.test.ts End-to-end-ish handler tests covering both wedge sites and benign concurrent-winner races.
packages/core/src/runtime/suspension-handler.ts Adds wedge detection/escalation on wait_created conflict path (suspension handler).
packages/core/src/runtime.ts Adds wedge detection/escalation on wait_completed conflict path (elapsed-wait pass) and routes CorruptedEventLogError to terminal handling.
docs/content/docs/v5/configuration/runtime-tuning.mdx Documents WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS behavior and default.
.changeset/wait-wedge-detection.md Changeset for the new runtime behavior (patch).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

runId,
queueItem.correlationId
));
if (suspectWedge) {
@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision failed 9 1.0m MISMATCH 1
in-flight-before-decision-counted failed 9 1.0m MISMATCH 1
in-flight-after-decision failed 9 1.0m MISMATCH 1
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-append-only.txt


/** Effective threshold. Override: `WORKFLOW_WAIT_WEDGE_FAIL_AFTER_SECONDS`. */
export const getWaitWedgeFailAfterSeconds = (): number =>
envNumber(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The wait_created wedge escalation uses the correlation-id ULID as a per-wait scheduling anchor, but that ULID encodes the run's creation time (a run-wide constant), so any run older than the threshold fails healthy waits with CORRUPTED_EVENT_LOG on a benign concurrent-suspension race.

Fix on Vercel

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants