From 468bfde6e55ecc7617a955632c1c5682786e0202 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 18:13:49 +0100 Subject: [PATCH 01/23] feat(workflow-engine): the contract of the workflow engine on the ledger @beonauto/workflow-engine holds the contract between the machine that runs a workflow and the adapters that store, time and execute for it, grouped by the engine's layers: machine, run log, timers, inbox, executor, dispatch, serialisation, settlement and engine. It knows workflows, the inputs of a run and an executor port, and nothing of brains, specs or models. - Inputs: started, timer fired, call answered, event received and cancel requested, each with the execution id and the time it arrived. - Events: one kind, input_applied, with the input's receipt, the change to the run's state as a JSON Patch, so evolve evaluates nothing, and the outputs it caused: arm or cancel a timer, start or cancel a call, settle. - State: plain JSON with a schema, so a snapshot is the state as it is; snapshots are due after 1,000 inputs or 1 MiB of events and stored in chunks of at most 65,536 UTF-16 code units, cut between characters. - Keys: timer ids, call keys of execution, reference and run, event ids and the execution id, with isStale saying which inputs a run has already taken. - Ports: run store, timers, executor, record store, dispatch watermark and run serialiser, each answering with an Effect; outputsAbove gives what a wake dispatches. The README states the invariants as sentences a reviewer can check, the limits (a run holds 4 MiB, an event 1.5 MiB, an input runs 100 tasks) and the open design points. A test checks that no source of the package uses a Node-only API, generates code or imports Temporal. The machine is the next step; the orchestration primitive still runs on Temporal. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/README.md | 133 ++++++++++ packages/workflow-engine/package.json | 23 ++ .../src/dispatch/dispatch-watermark.test.ts | 45 ++++ .../src/dispatch/dispatch-watermark.ts | 31 +++ .../src/dispatch/run-output.ts | 62 +++++ .../src/engine/portability.test.ts | 39 +++ .../src/engine/workflow-engine.ts | 33 +++ .../src/executor/call-key.test.ts | 23 ++ .../workflow-engine/src/executor/call-key.ts | 13 + .../src/executor/call-result.ts | 10 + .../workflow-engine/src/executor/executor.ts | 9 + .../src/inbox/received-event.ts | 20 ++ packages/workflow-engine/src/index.ts | 79 ++++++ .../src/machine/input-receipt.test.ts | 71 ++++++ .../src/machine/input-receipt.ts | 47 ++++ .../workflow-engine/src/machine/limits.ts | 7 + .../src/machine/run-decider.ts | 7 + .../workflow-engine/src/machine/run-input.ts | 79 ++++++ .../workflow-engine/src/machine/run-state.ts | 236 ++++++++++++++++++ .../src/run-log/run-event.test.ts | 81 ++++++ .../workflow-engine/src/run-log/run-event.ts | 19 ++ .../workflow-engine/src/run-log/run-store.ts | 32 +++ .../src/run-log/snapshot.test.ts | 46 ++++ .../workflow-engine/src/run-log/snapshot.ts | 54 ++++ .../src/run-log/state-patch.ts | 13 + .../src/serialisation/run-serialiser.ts | 5 + .../src/settlement/record-store.ts | 10 + .../src/settlement/run-settlement.ts | 13 + packages/workflow-engine/src/testing/runs.ts | 105 ++++++++ .../workflow-engine/src/timers/timer-id.ts | 16 ++ packages/workflow-engine/src/timers/timers.ts | 9 + packages/workflow-engine/tsconfig.json | 4 + packages/workflow-engine/vitest.config.ts | 5 + pnpm-lock.yaml | 16 ++ 34 files changed, 1395 insertions(+) create mode 100644 packages/workflow-engine/README.md create mode 100644 packages/workflow-engine/package.json create mode 100644 packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts create mode 100644 packages/workflow-engine/src/dispatch/dispatch-watermark.ts create mode 100644 packages/workflow-engine/src/dispatch/run-output.ts create mode 100644 packages/workflow-engine/src/engine/portability.test.ts create mode 100644 packages/workflow-engine/src/engine/workflow-engine.ts create mode 100644 packages/workflow-engine/src/executor/call-key.test.ts create mode 100644 packages/workflow-engine/src/executor/call-key.ts create mode 100644 packages/workflow-engine/src/executor/call-result.ts create mode 100644 packages/workflow-engine/src/executor/executor.ts create mode 100644 packages/workflow-engine/src/inbox/received-event.ts create mode 100644 packages/workflow-engine/src/index.ts create mode 100644 packages/workflow-engine/src/machine/input-receipt.test.ts create mode 100644 packages/workflow-engine/src/machine/input-receipt.ts create mode 100644 packages/workflow-engine/src/machine/limits.ts create mode 100644 packages/workflow-engine/src/machine/run-decider.ts create mode 100644 packages/workflow-engine/src/machine/run-input.ts create mode 100644 packages/workflow-engine/src/machine/run-state.ts create mode 100644 packages/workflow-engine/src/run-log/run-event.test.ts create mode 100644 packages/workflow-engine/src/run-log/run-event.ts create mode 100644 packages/workflow-engine/src/run-log/run-store.ts create mode 100644 packages/workflow-engine/src/run-log/snapshot.test.ts create mode 100644 packages/workflow-engine/src/run-log/snapshot.ts create mode 100644 packages/workflow-engine/src/run-log/state-patch.ts create mode 100644 packages/workflow-engine/src/serialisation/run-serialiser.ts create mode 100644 packages/workflow-engine/src/settlement/record-store.ts create mode 100644 packages/workflow-engine/src/settlement/run-settlement.ts create mode 100644 packages/workflow-engine/src/testing/runs.ts create mode 100644 packages/workflow-engine/src/timers/timer-id.ts create mode 100644 packages/workflow-engine/src/timers/timers.ts create mode 100644 packages/workflow-engine/tsconfig.json create mode 100644 packages/workflow-engine/vitest.config.ts diff --git a/packages/workflow-engine/README.md b/packages/workflow-engine/README.md new file mode 100644 index 000000000..1903b5b85 --- /dev/null +++ b/packages/workflow-engine/README.md @@ -0,0 +1,133 @@ +# @beonauto/workflow-engine + +The core of the workflow engine that runs on the ledger: the contract between the machine that runs a workflow and the adapters that store, time and execute for it. It knows workflows, the inputs a run takes and an executor that performs calls. It does not know brains, prompts, models or specs: whatever an adapter needs to know about the run, such as who started it, it passes as opaque `attributes` and gets back with every output. + +The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. This package holds the contract: the types, the ports, the idempotency keys, the dispatch watermark and the invariants below. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, unchanged. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. + +## How a run moves + +1. An adapter submits an input for an execution, with `at`, the time on its own clock. The machine never reads a clock. +2. Holding the run's serialisation, the engine loads the run: its latest snapshot and the events after it, folded with `evolve`. +3. It asks the decider. A stale input, or one that changes nothing, appends nothing. Otherwise the decision is one event, appended with the version the engine read as the expected version. This is the load-decide-append loop of `@beonauto/ledger`: a version conflict loads the run again and decides again, up to three more times. +4. When a snapshot is due, the engine saves one. +5. It dispatches the outputs of every event above the run's dispatch watermark, in the order of the stream, and then moves the watermark up to the last event whose outputs were all dispatched. +6. `wake(executionId)` does step 5 again. An adapter wakes a run after a crash, on a restart, and from a sweep that finds runs whose watermark is behind their stream. + +## Layers and ports + +| Layer | Folder | What it holds | Port | +| ------------- | ------------------- | ----------------------------------------------------- | --------------------------------------- | +| machine | `src/machine` | inputs, state, staleness, limits, the decider's type | none: pure | +| run log | `src/run-log` | the run's events, its state patches, snapshots | `RunStore` (Emmett, one stream per run) | +| timers | `src/timers` | timer ids and what each timer is for | `Timers` | +| inbox | `src/inbox` | the external events a run receives, and their limits | none: events arrive as `event_received` | +| executor | `src/executor` | call keys and call results | `Executor` | +| dispatch | `src/dispatch` | outputs and the dispatch watermark | `DispatchWatermark` | +| serialisation | `src/serialisation` | one input at a time for each run | `RunSerialiser` | +| settlement | `src/settlement` | the settlement of a run and the store that records it | `RecordStore` (the brain's ledger) | +| engine | `src/engine` | the ports together and the engine's own interface | `WorkflowEngine` | + +Every port answers with an Effect. None of them is a clock: time comes in with the inputs. + +## Inputs + +| Input | Carries | Stale when | +| ------------------ | ------------------------------------------------------------------------------- | ------------------------------------------------------------- | +| `started` | the document, the input, the limits, the attributes and a seed for random draws | the run has started before | +| `timer_fired` | the timer id | the timer is not armed: never armed, already fired, cancelled | +| `call_answered` | the call key and the result: succeeded, rejected, failed or unreachable | the call is not open: never started, answered, cancelled | +| `event_received` | the event, with an `id` of 1 to 256 characters and a `type` | the run has received an event with that id | +| `cancel_requested` | nothing more | a cancel was requested before | + +Every input carries the execution id and `at`. Any input but `started` is stale for a run that has not started, has ended, or is another execution. `isStale(state, input)` is this table. + +## Events + +A run's stream holds one kind of event, `input_applied`: + +- `receipt`: the kind of the input, the key it is deduplicated by and its `at`. The input's payload is not stored again: what it changed is in the patch. +- `patch`: the change the input made to the run's state, as JSON Patch operations (RFC 6902: `add`, `replace`, `remove`) addressed by JSON Pointer. +- `outputs`: what the engine must do because of this input. + +`evolve` applies the patch and nothing else, so replaying a run never evaluates an expression and never depends on the version of the machine that decided it. + +## Outputs + +| Output | Carries | Idempotent by | Port | +| -------------- | ---------------------------------------------------------------------------------- | ------------- | -------------------- | +| `arm_timer` | timer id, due time, purpose | timer id | `Timers.arm` | +| `cancel_timer` | timer id | timer id | `Timers.cancel` | +| `start_call` | call key, the call as the document names it, its arguments, the longest it may run | call key | `Executor.start` | +| `cancel_call` | call key | call key | `Executor.cancel` | +| `settle` | execution id and settlement: succeeded, rejected or failed | execution id | `RecordStore.settle` | + +Answers come back as inputs: a fired timer as `timer_fired`, a finished call as `call_answered`. `RecordStore.settle` answers `recorded`, `already_recorded`, `settled_otherwise` or `unknown_execution`; each of them is final, and only a failure to answer leaves the output to be dispatched again. + +## Idempotency keys + +| Key | Made of | Where it is deduplicated | +| ------------ | --------------------------------------------------------- | ---------------------------------------------------- | +| execution id | given by the adapter that starts the run | the run's status; `RecordStore` by execution id | +| timer id | `/timers/`, n counting up within the run | `state.timers.armed`; `Timers` by id | +| call key | execution id, the task's reference, the run of that task | `state.calls.open`; `Executor` by `callKeyText(key)` | +| event id | the external event's own `id` | `state.inbox.receivedIds` | +| snapshot | execution id and stream version | the run store keeps only the latest | +| watermark | execution id | the watermark store | + +The event store does not deduplicate anything: Emmett appends a message with an id it has seen before as a new message, on SQLite and on D1. Deduplication lives in the run's state. + +## The dispatch watermark + +The watermark of a run is a stream version. Every output of every event at or below it has been dispatched at least once. Outputs above it are dispatched on the next submit or wake, in the order of the stream, and the watermark then moves up to the last event whose outputs were all dispatched. A crash between the append and the dispatch loses nothing: the outputs are in the event, and the next wake dispatches them. + +## State and snapshots + +A run's state is plain JSON: no `Map`, `Set`, `Date`, `undefined`, class or function, so `JSON.parse(JSON.stringify(state))` is the state. It holds the run's definition and attributes, the machine's frames and context, the armed timers and open calls, the inbox, the random draws made so far, the cancel request and, once the run ends, its outcome. + +A snapshot is `{ format, executionId, version, state }`, the state folded from the events up to `version`. It is due after 1,000 inputs or 1 MiB of events since the last one, whichever comes first. It is stored in chunks of at most 65,536 UTF-16 code units, at most 192 KiB in UTF-8 each, cut between characters, never inside one; the run store keeps only the latest snapshot. + +## Limits + +| Limit | Value | Why | +| --------------------------- | ------------ | -------------------------------------------------------------------------------- | +| data a run holds | 4 MiB | measured by size, not by object identity; it bounds the state, so every snapshot | +| one event | 1.5 MiB | a call may answer with 1 MiB; D1 and Durable Object rows take up to 2 MB | +| tasks in one input | 100 | then the machine arms a timer due at once and goes on when it fires | +| tasks without waiting | 10,000 | as the interpreter has it | +| events waiting in the inbox | 64, 1 MiB | as the interpreter has it | +| events a run receives | 1,024, 4 MiB | as the interpreter has it; it also bounds the event ids kept for deduplication | + +The 16 MiB a run could hold on Temporal comes down to 4 MiB, and snapshots are chunked: a snapshot is then at most about 5 MiB in chunks of at most 192 KiB. Bringing the limit under 1 MiB instead would break what callers rely on today: a call answers with up to 1 MiB, and a workflow's output may take 1 MiB. Temporal's limits on history, 40,000 events and 8 MiB, go: a run is loaded from its snapshot, so the length of its history no longer costs time. + +## Invariants + +Each sentence is something a reviewer can check against the code or a test. + +1. A run has exactly one stream, and the stream's version is the number of inputs the run applied. +2. `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source and no storage, and the same state and input give the same events. +3. Time in a run is only ever an input's `at`; random draws come from the seed in `started` and the number of draws in the state. +4. `evolve(state, event)` applies the event's patch and does nothing else. +5. An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version. +6. A stale input appends nothing, and neither does an input that would change nothing. +7. A late answer, a duplicate answer, a second delivery of an event and the fire of a cancelled timer are stale inputs. +8. The deduplication state is bounded: armed timers and open calls are what is outstanding, and a run keeps at most 1,024 event ids. +9. Timer ids are never reused within a run, and a call key's run counts up for each reference. +10. The outputs of a run are exactly the `outputs` of its events, and an output is dispatched only after the event that holds it is appended. +11. Every output is idempotent by its key, so dispatching it twice has the effect of dispatching it once. +12. The watermark never goes down, and every output of every event at or below it has been dispatched at least once. +13. A run's outcome is in its stream before the record store is asked to record it, and the `settle` output is dispatched again until the record store answers. +14. The run log and the record store may be different stores: the contract assumes two writes, each idempotent by execution id, the second retried. An adapter whose two stores are one database may make them one transaction. +15. Loading a run from its latest snapshot and the events after it gives the same state as folding its whole stream. +16. Only the latest snapshot of a run is kept, and no stored chunk of it is longer than 65,536 UTF-16 code units. +17. The state of a run is plain JSON, and the data it holds, measured by size, stays at or under 4 MiB. +18. No event is larger than 1.5 MiB as JSON, and no input runs more than 100 tasks. +19. The machine never assumes it is the only writer: it relies only on the expected version of each append. Serialising the inputs of one run is the adapter's job. +20. Nothing in this package uses a Node-only API, generates code, or imports Temporal (`src/engine/portability.test.ts`). +21. No module of the engine keeps a cache that grows with the history of a run: what an input costs in memory is bounded by the input and the state. + +## Open design points + +- Whether an event may reach a run before its `started` input: the contract calls it stale, so an adapter must deliver events only to a started run. +- Whether the arguments of a `call` are checked by the machine, as the interpreter checks a spec call today, or by the executor, which would answer `rejected`. The engine itself does not know specs. +- The 1 MiB snapshot interval with a 4 MiB state can write up to four bytes of snapshot for each byte of history. +- What replaces Temporal's limits on history, if anything, now that a run's history no longer costs time to load. diff --git a/packages/workflow-engine/package.json b/packages/workflow-engine/package.json new file mode 100644 index 000000000..27d80d8a7 --- /dev/null +++ b/packages/workflow-engine/package.json @@ -0,0 +1,23 @@ +{ + "name": "@beonauto/workflow-engine", + "version": "0.0.0", + "private": true, + "license": "Elastic-2.0", + "type": "module", + "exports": { + ".": "./src/index.ts" + }, + "scripts": { + "lint": "oxlint -c ../../.oxlintrc.json", + "typecheck": "tsc", + "test": "vitest run --coverage" + }, + "dependencies": { + "@beonauto/operations": "workspace:*", + "effect": "catalog:" + }, + "devDependencies": { + "@vitest/coverage-v8": "catalog:", + "vitest": "catalog:" + } +} diff --git a/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts new file mode 100644 index 000000000..b7fb0e4f7 --- /dev/null +++ b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts @@ -0,0 +1,45 @@ +import { describe, expect, it } from 'vitest'; + +import { DispatchFailed, outputsAbove, type PositionedEvent, type RunOutput } from '../index.ts'; +import { at, executionId, openCall } from '../testing/runs.ts'; + +const arm: RunOutput = { + kind: 'arm_timer', + executionId, + timerId: `${executionId}/timers/1`, + dueAt: at + 60_000, + purpose: 'timeout', +}; + +const start: RunOutput = { kind: 'start_call', key: openCall, function: 'executeSpec', arguments: {}, longestMs: 600 }; + +const settle: RunOutput = { kind: 'settle', executionId, settlement: { status: 'failed' } }; + +function applied(version: number, outputs: readonly RunOutput[]): PositionedEvent { + return { + version, + event: { type: 'input_applied', receipt: { kind: 'timer_fired', key: 'k', at }, patch: [], outputs }, + }; +} + +describe('the outputs above a watermark', () => { + it('are those of the events after it, in the order of the stream', () => { + const events = [applied(3, [settle]), applied(1, [start]), applied(2, [arm, start])]; + + expect(outputsAbove(1, events)).toEqual([ + { version: 2, output: arm }, + { version: 2, output: start }, + { version: 3, output: settle }, + ]); + expect(outputsAbove(3, events)).toEqual([]); + }); +}); + +describe('a failed dispatch', () => { + it('names the kind of output that could not be dispatched', () => { + expect(new DispatchFailed({ output: 'settle', detail: 'the record store did not answer' })).toMatchObject({ + _tag: 'dispatch_failed', + output: 'settle', + }); + }); +}); diff --git a/packages/workflow-engine/src/dispatch/dispatch-watermark.ts b/packages/workflow-engine/src/dispatch/dispatch-watermark.ts new file mode 100644 index 000000000..bf84d850c --- /dev/null +++ b/packages/workflow-engine/src/dispatch/dispatch-watermark.ts @@ -0,0 +1,31 @@ +import { Data, type Effect, type Schema } from 'effect'; + +import type { PositionedEvent } from '../run-log/run-event.ts'; +import type { RunOutput } from './run-output.ts'; + +export interface DispatchWatermark { + readonly read: (executionId: string) => Effect.Effect; + readonly advance: (executionId: string, through: number) => Effect.Effect; +} + +export interface RunContext { + readonly executionId: string; + readonly attributes: Schema.JsonObject; +} + +export interface PositionedOutput { + readonly version: number; + readonly output: RunOutput; +} + +export class DispatchFailed extends Data.TaggedError('dispatch_failed')<{ + readonly output: RunOutput['kind']; + readonly detail: string; +}> {} + +export function outputsAbove(watermark: number, events: readonly PositionedEvent[]): readonly PositionedOutput[] { + return events + .filter(({ version }) => version > watermark) + .toSorted((first, second) => first.version - second.version) + .flatMap(({ version, event }) => event.outputs.map((output) => ({ version, output }))); +} diff --git a/packages/workflow-engine/src/dispatch/run-output.ts b/packages/workflow-engine/src/dispatch/run-output.ts new file mode 100644 index 000000000..f9ae948bb --- /dev/null +++ b/packages/workflow-engine/src/dispatch/run-output.ts @@ -0,0 +1,62 @@ +import { Schema } from 'effect'; + +import { CallKeySchema } from '../executor/call-key.ts'; +import { RunSettlementSchema } from '../settlement/run-settlement.ts'; +import { TimerPurposeSchema } from '../timers/timer-id.ts'; + +const ExecutionIdSchema = Schema.NonEmptyString; + +const MillisecondsSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); + +const ArmTimerSchema = Schema.Struct({ + kind: Schema.Literal('arm_timer'), + executionId: ExecutionIdSchema, + timerId: Schema.NonEmptyString, + dueAt: MillisecondsSchema, + purpose: TimerPurposeSchema, +}); + +const CancelTimerSchema = Schema.Struct({ + kind: Schema.Literal('cancel_timer'), + executionId: ExecutionIdSchema, + timerId: Schema.NonEmptyString, +}); + +const StartCallSchema = Schema.Struct({ + kind: Schema.Literal('start_call'), + key: CallKeySchema, + function: Schema.NonEmptyString, + arguments: Schema.Json, + longestMs: MillisecondsSchema, +}); + +const CancelCallSchema = Schema.Struct({ + kind: Schema.Literal('cancel_call'), + key: CallKeySchema, +}); + +const SettleSchema = Schema.Struct({ + kind: Schema.Literal('settle'), + executionId: ExecutionIdSchema, + settlement: RunSettlementSchema, +}); + +export const RunOutputSchema = Schema.Union([ + ArmTimerSchema, + CancelTimerSchema, + StartCallSchema, + CancelCallSchema, + SettleSchema, +]); + +export type ArmTimer = typeof ArmTimerSchema.Type; + +export type CancelTimer = typeof CancelTimerSchema.Type; + +export type StartCall = typeof StartCallSchema.Type; + +export type CancelCall = typeof CancelCallSchema.Type; + +export type Settle = typeof SettleSchema.Type; + +export type RunOutput = typeof RunOutputSchema.Type; diff --git a/packages/workflow-engine/src/engine/portability.test.ts b/packages/workflow-engine/src/engine/portability.test.ts new file mode 100644 index 000000000..69b876e15 --- /dev/null +++ b/packages/workflow-engine/src/engine/portability.test.ts @@ -0,0 +1,39 @@ +import { readdirSync, readFileSync } from 'node:fs'; +import { join } from 'node:path'; +import { fileURLToPath } from 'node:url'; + +import { describe, expect, it } from 'vitest'; + +interface Forbidden { + readonly what: string; + readonly pattern: Readonly; +} + +const source = fileURLToPath(new URL('..', import.meta.url)); + +const nodeOnly: readonly Forbidden[] = [ + { what: 'a node: module', pattern: /from ['"]node:/u }, + { what: 'a Node global', pattern: /\b(?:process|Buffer|require|setImmediate|__dirname)\b/u }, + { what: 'code generation', pattern: /\beval\(|new Function\(/u }, + { what: 'Temporal', pattern: /@temporalio\//u }, +]; + +function isProductionSource(file: string): boolean { + return file.endsWith('.ts') && !file.endsWith('.test.ts'); +} + +function findingsIn(file: string): readonly string[] { + const text = readFileSync(join(source, file), 'utf8'); + return nodeOnly + .filter(({ pattern }: Forbidden) => pattern.test(text)) + .map(({ what }: Forbidden) => `${file}: ${what}`); +} + +describe('the engine core', () => { + it('uses no Node-only API, no code generation and no Temporal, so it runs in workerd as it runs in Node', () => { + const files = readdirSync(source, { recursive: true, encoding: 'utf8' }).filter((file) => isProductionSource(file)); + + expect(files.length).toBeGreaterThan(10); + expect(files.flatMap((file) => findingsIn(file))).toEqual([]); + }); +}); diff --git a/packages/workflow-engine/src/engine/workflow-engine.ts b/packages/workflow-engine/src/engine/workflow-engine.ts new file mode 100644 index 000000000..4b90f9318 --- /dev/null +++ b/packages/workflow-engine/src/engine/workflow-engine.ts @@ -0,0 +1,33 @@ +import type { Effect } from 'effect'; + +import type { DispatchWatermark } from '../dispatch/dispatch-watermark.ts'; +import type { Executor } from '../executor/executor.ts'; +import type { RunInput } from '../machine/run-input.ts'; +import type { RunStore, VersionConflict } from '../run-log/run-store.ts'; +import type { RunSerialiser } from '../serialisation/run-serialiser.ts'; +import type { RecordStore } from '../settlement/record-store.ts'; +import type { Timers } from '../timers/timers.ts'; + +export interface EnginePorts { + readonly runStore: RunStore; + readonly watermark: DispatchWatermark; + readonly timers: Timers; + readonly executor: Executor; + readonly recordStore: RecordStore; + readonly serialiser: RunSerialiser; +} + +export interface Submission { + readonly applied: boolean; + readonly version: number; +} + +export interface Wake { + readonly version: number; + readonly dispatchedThrough: number; +} + +export interface WorkflowEngine { + readonly submit: (input: RunInput) => Effect.Effect; + readonly wake: (executionId: string) => Effect.Effect; +} diff --git a/packages/workflow-engine/src/executor/call-key.test.ts b/packages/workflow-engine/src/executor/call-key.test.ts new file mode 100644 index 000000000..e1a413aa5 --- /dev/null +++ b/packages/workflow-engine/src/executor/call-key.test.ts @@ -0,0 +1,23 @@ +import { describe, expect, it } from 'vitest'; + +import { callKeyText, timerIdOf, type CallKey } from '../index.ts'; + +describe('the text of a call key', () => { + it('is the same for the same execution, reference and run, and different whenever one differs', () => { + const keys = [ + { executionId: 'e', reference: '/do/0', run: 1 }, + { executionId: 'e', reference: '/do/0', run: 2 }, + { executionId: 'e/do', reference: '/0', run: 1 }, + { executionId: 'e', reference: '/do/0","x', run: 1 }, + ].map((key: CallKey) => callKeyText(key)); + + expect(new Set(keys).size).toBe(4); + expect(callKeyText({ executionId: 'e', reference: '/do/0', run: 1 })).toBe(keys[0]); + }); +}); + +describe('the id of a timer', () => { + it('is the execution and the sequence number the run gave it', () => { + expect([timerIdOf('e', 1), timerIdOf('e', 2)]).toEqual(['e/timers/1', 'e/timers/2']); + }); +}); diff --git a/packages/workflow-engine/src/executor/call-key.ts b/packages/workflow-engine/src/executor/call-key.ts new file mode 100644 index 000000000..80ec7b0bf --- /dev/null +++ b/packages/workflow-engine/src/executor/call-key.ts @@ -0,0 +1,13 @@ +import { Schema } from 'effect'; + +export const CallKeySchema = Schema.Struct({ + executionId: Schema.NonEmptyString, + reference: Schema.String, + run: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), +}); + +export type CallKey = typeof CallKeySchema.Type; + +export function callKeyText({ executionId, reference, run }: CallKey): string { + return JSON.stringify([executionId, reference, run]); +} diff --git a/packages/workflow-engine/src/executor/call-result.ts b/packages/workflow-engine/src/executor/call-result.ts new file mode 100644 index 000000000..5e174735c --- /dev/null +++ b/packages/workflow-engine/src/executor/call-result.ts @@ -0,0 +1,10 @@ +import { Schema } from 'effect'; + +export const CallResultSchema = Schema.Union([ + Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), + Schema.Struct({ status: Schema.Literal('rejected'), reason: Schema.String, detail: Schema.String }), + Schema.Struct({ status: Schema.Literal('failed'), detail: Schema.String }), + Schema.Struct({ status: Schema.Literal('unreachable'), detail: Schema.String }), +]); + +export type CallResult = typeof CallResultSchema.Type; diff --git a/packages/workflow-engine/src/executor/executor.ts b/packages/workflow-engine/src/executor/executor.ts new file mode 100644 index 000000000..fc5c54323 --- /dev/null +++ b/packages/workflow-engine/src/executor/executor.ts @@ -0,0 +1,9 @@ +import type { Effect } from 'effect'; + +import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; +import type { CancelCall, StartCall } from '../dispatch/run-output.ts'; + +export interface Executor { + readonly start: (call: StartCall, run: RunContext) => Effect.Effect; + readonly cancel: (call: CancelCall, run: RunContext) => Effect.Effect; +} diff --git a/packages/workflow-engine/src/inbox/received-event.ts b/packages/workflow-engine/src/inbox/received-event.ts new file mode 100644 index 000000000..952de4281 --- /dev/null +++ b/packages/workflow-engine/src/inbox/received-event.ts @@ -0,0 +1,20 @@ +import { Schema } from 'effect'; + +export const mostEventIdLength = 256; + +export const mostWaitingEvents = 64; + +export const mostWaitingEventBytes = 1_048_576; + +export const mostReceivedEvents = 1024; + +export const mostReceivedEventBytes = 4_194_304; + +const EventIdSchema = Schema.String.check(Schema.isMinLength(1), Schema.isMaxLength(mostEventIdLength)); + +export const ReceivedEventSchema = Schema.StructWithRest( + Schema.Struct({ id: EventIdSchema, type: Schema.NonEmptyString }), + [Schema.Record(Schema.String, Schema.Json)], +); + +export type ReceivedEvent = typeof ReceivedEventSchema.Type; diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts new file mode 100644 index 000000000..94439cca3 --- /dev/null +++ b/packages/workflow-engine/src/index.ts @@ -0,0 +1,79 @@ +export { + DispatchFailed, + outputsAbove, + type DispatchWatermark, + type PositionedOutput, + type RunContext, +} from './dispatch/dispatch-watermark.ts'; +export { + RunOutputSchema, + type ArmTimer, + type CancelCall, + type CancelTimer, + type RunOutput, + type Settle, + type StartCall, +} from './dispatch/run-output.ts'; +export type { EnginePorts, Submission, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; +export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; +export { CallResultSchema, type CallResult } from './executor/call-result.ts'; +export type { Executor } from './executor/executor.ts'; +export { + ReceivedEventSchema, + mostEventIdLength, + mostReceivedEventBytes, + mostReceivedEvents, + mostWaitingEventBytes, + mostWaitingEvents, + type ReceivedEvent, +} from './inbox/received-event.ts'; +export { InputReceiptSchema, isStale, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; +export { mostEventBytes, mostHeldBytes, mostStepsWithoutWaiting, mostTasksPerInput } from './machine/limits.ts'; +export type { RunDecider } from './machine/run-decider.ts'; +export { + RunInputSchema, + type CallAnswered, + type CancelRequested, + type EventReceived, + type RunInput, + type RunInputKind, + type RunLimits, + type Started, + type TimerFired, +} from './machine/run-input.ts'; +export { + RunStateSchema, + newRun, + type ArmedTimer, + type Branch, + type DslError, + type FrameBody, + type InboxState, + type ListCursor, + type MachineState, + type RunOutcome, + type RunState, + type TaskFrame, + type TryPhase, + type Variables, + type WaitingEvent, +} from './machine/run-state.ts'; +export { RunEventSchema, type PositionedEvent, type RunEvent } from './run-log/run-event.ts'; +export { VersionConflict, type AppendedEvent, type LoadedRun, type RunStore } from './run-log/run-store.ts'; +export { + SnapshotSchema, + isSnapshotDue, + snapshotChunkLength, + snapshotChunks, + snapshotEveryBytes, + snapshotEveryInputs, + snapshotFromChunks, + type SinceSnapshot, + type Snapshot, +} from './run-log/snapshot.ts'; +export { PatchOperationSchema, type PatchOperation, type StatePatch } from './run-log/state-patch.ts'; +export type { RunSerialiser } from './serialisation/run-serialiser.ts'; +export type { RecordStore, SettleReceipt } from './settlement/record-store.ts'; +export { RunSettlementSchema, type RunSettlement } from './settlement/run-settlement.ts'; +export { TimerPurposeSchema, timerIdOf, type TimerPurpose } from './timers/timer-id.ts'; +export type { Timers } from './timers/timers.ts'; diff --git a/packages/workflow-engine/src/machine/input-receipt.test.ts b/packages/workflow-engine/src/machine/input-receipt.test.ts new file mode 100644 index 000000000..44e428581 --- /dev/null +++ b/packages/workflow-engine/src/machine/input-receipt.test.ts @@ -0,0 +1,71 @@ +import { describe, expect, it } from 'vitest'; + +import { callKeyText, isStale, newRun, receiptOf, type RunInput } from '../index.ts'; +import { armedTimer, at, executionId, openCall, runningState, started } from '../testing/runs.ts'; + +const fired = (timerId: string): RunInput => ({ kind: 'timer_fired', executionId, at, timerId }); + +const answered = (run: number): RunInput => ({ + kind: 'call_answered', + executionId, + at, + key: { ...openCall, run }, + result: { status: 'succeeded', output: { approved: true } }, +}); + +const received = (id: string): RunInput => ({ + kind: 'event_received', + executionId, + at, + event: { id, type: 'com.acme.approval' }, +}); + +const cancelled: RunInput = { kind: 'cancel_requested', executionId, at }; + +describe('the receipt of an input', () => { + it('names the input by the key it is deduplicated by', () => { + expect( + [started, fired(armedTimer), answered(1), received('event-9'), cancelled].map((input) => receiptOf(input)), + ).toEqual([ + { kind: 'started', key: executionId, at }, + { kind: 'timer_fired', key: armedTimer, at }, + { kind: 'call_answered', key: callKeyText(openCall), at }, + { kind: 'event_received', key: 'event-9', at }, + { kind: 'cancel_requested', key: executionId, at }, + ]); + }); +}); + +describe('an input that can still change a run', () => { + it('is a start of a new run, the fire of an armed timer, the answer of an open call, a new event or the first cancel', () => { + expect( + [ + isStale(newRun, started), + isStale(runningState, fired(armedTimer)), + isStale(runningState, answered(1)), + isStale(runningState, received('event-3')), + isStale(runningState, cancelled), + ].every((stale) => !stale), + ).toBe(true); + }); +}); + +describe('a stale input', () => { + it('is a second start, a fire of a timer not armed, an answer of a call not open, an event seen before or a second cancel', () => { + expect([ + isStale(runningState, started), + isStale(runningState, fired(`${executionId}/timers/9`)), + isStale(runningState, answered(2)), + isStale(runningState, received('event-1')), + isStale({ ...runningState, cancelRequested: true }, cancelled), + ]).toEqual([true, true, true, true, true]); + }); + + it('is anything but a start for a run that has not started, has ended, or is another run', () => { + expect([ + isStale(newRun, fired(armedTimer)), + isStale({ ...runningState, status: 'ended' }, answered(1)), + isStale({ ...runningState, executionId: 'another' }, received('event-3')), + ]).toEqual([true, true, true]); + }); +}); diff --git a/packages/workflow-engine/src/machine/input-receipt.ts b/packages/workflow-engine/src/machine/input-receipt.ts new file mode 100644 index 000000000..9a1b1b359 --- /dev/null +++ b/packages/workflow-engine/src/machine/input-receipt.ts @@ -0,0 +1,47 @@ +import { Schema } from 'effect'; + +import { callKeyText } from '../executor/call-key.ts'; +import type { RunInput } from './run-input.ts'; +import type { RunState } from './run-state.ts'; + +export const InputReceiptSchema = Schema.Struct({ + kind: Schema.Literals(['started', 'timer_fired', 'call_answered', 'event_received', 'cancel_requested']), + key: Schema.String, + at: Schema.Int, +}); + +export type InputReceipt = typeof InputReceiptSchema.Type; + +function keyOf(input: RunInput): string { + if (input.kind === 'timer_fired') { + return input.timerId; + } + if (input.kind === 'call_answered') { + return callKeyText(input.key); + } + return input.kind === 'event_received' ? input.event.id : input.executionId; +} + +export function receiptOf(input: RunInput): InputReceipt { + return { kind: input.kind, key: keyOf(input), at: input.at }; +} + +function isSpent(state: RunState, input: RunInput): boolean { + if (input.kind === 'started') { + return state.status !== 'new'; + } + if (input.kind === 'timer_fired') { + return !Object.hasOwn(state.timers.armed, input.timerId); + } + if (input.kind === 'call_answered') { + return !Object.hasOwn(state.calls.open, callKeyText(input.key)); + } + return input.kind === 'event_received' ? state.inbox.receivedIds.includes(input.event.id) : state.cancelRequested; +} + +export function isStale(state: RunState, input: RunInput): boolean { + if (input.kind !== 'started' && (state.status !== 'running' || input.executionId !== state.executionId)) { + return true; + } + return isSpent(state, input); +} diff --git a/packages/workflow-engine/src/machine/limits.ts b/packages/workflow-engine/src/machine/limits.ts new file mode 100644 index 000000000..64b2f9827 --- /dev/null +++ b/packages/workflow-engine/src/machine/limits.ts @@ -0,0 +1,7 @@ +export const mostHeldBytes = 4_194_304; + +export const mostEventBytes = 1_572_864; + +export const mostTasksPerInput = 100; + +export const mostStepsWithoutWaiting = 10_000; diff --git a/packages/workflow-engine/src/machine/run-decider.ts b/packages/workflow-engine/src/machine/run-decider.ts new file mode 100644 index 000000000..873a407f2 --- /dev/null +++ b/packages/workflow-engine/src/machine/run-decider.ts @@ -0,0 +1,7 @@ +import type { Decider } from '@beonauto/operations'; + +import type { RunEvent } from '../run-log/run-event.ts'; +import type { RunInput } from './run-input.ts'; +import type { RunState } from './run-state.ts'; + +export type RunDecider = Decider; diff --git a/packages/workflow-engine/src/machine/run-input.ts b/packages/workflow-engine/src/machine/run-input.ts new file mode 100644 index 000000000..c6a03c571 --- /dev/null +++ b/packages/workflow-engine/src/machine/run-input.ts @@ -0,0 +1,79 @@ +import { Schema } from 'effect'; + +import { CallKeySchema } from '../executor/call-key.ts'; +import { CallResultSchema } from '../executor/call-result.ts'; +import { ReceivedEventSchema } from '../inbox/received-event.ts'; + +const ExecutionIdSchema = Schema.NonEmptyString; + +const InstantSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); + +const PositiveMillisecondsSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)); + +export const RunLimitsSchema = Schema.Struct({ + mostDurationMs: PositiveMillisecondsSchema, + longestCallMs: PositiveMillisecondsSchema, +}); + +const StartedSchema = Schema.Struct({ + kind: Schema.Literal('started'), + executionId: ExecutionIdSchema, + at: InstantSchema, + document: Schema.JsonObject, + input: Schema.Json, + limits: RunLimitsSchema, + attributes: Schema.JsonObject, + seed: Schema.Int, +}); + +const TimerFiredSchema = Schema.Struct({ + kind: Schema.Literal('timer_fired'), + executionId: ExecutionIdSchema, + at: InstantSchema, + timerId: Schema.NonEmptyString, +}); + +const CallAnsweredSchema = Schema.Struct({ + kind: Schema.Literal('call_answered'), + executionId: ExecutionIdSchema, + at: InstantSchema, + key: CallKeySchema, + result: CallResultSchema, +}); + +const EventReceivedSchema = Schema.Struct({ + kind: Schema.Literal('event_received'), + executionId: ExecutionIdSchema, + at: InstantSchema, + event: ReceivedEventSchema, +}); + +const CancelRequestedSchema = Schema.Struct({ + kind: Schema.Literal('cancel_requested'), + executionId: ExecutionIdSchema, + at: InstantSchema, +}); + +export const RunInputSchema = Schema.Union([ + StartedSchema, + TimerFiredSchema, + CallAnsweredSchema, + EventReceivedSchema, + CancelRequestedSchema, +]); + +export type RunLimits = typeof RunLimitsSchema.Type; + +export type Started = typeof StartedSchema.Type; + +export type TimerFired = typeof TimerFiredSchema.Type; + +export type CallAnswered = typeof CallAnsweredSchema.Type; + +export type EventReceived = typeof EventReceivedSchema.Type; + +export type CancelRequested = typeof CancelRequestedSchema.Type; + +export type RunInput = typeof RunInputSchema.Type; + +export type RunInputKind = RunInput['kind']; diff --git a/packages/workflow-engine/src/machine/run-state.ts b/packages/workflow-engine/src/machine/run-state.ts new file mode 100644 index 000000000..670de117a --- /dev/null +++ b/packages/workflow-engine/src/machine/run-state.ts @@ -0,0 +1,236 @@ +import { Schema } from 'effect'; + +import { CallKeySchema, type CallKey } from '../executor/call-key.ts'; +import { ReceivedEventSchema, type ReceivedEvent } from '../inbox/received-event.ts'; +import { TimerPurposeSchema, type TimerPurpose } from '../timers/timer-id.ts'; +import { RunLimitsSchema, type RunLimits } from './run-input.ts'; + +export interface DslError { + readonly type: string; + readonly status: number; + readonly instance: string; + readonly title?: string; + readonly detail?: string; +} + +export type Variables = Readonly>; + +export interface ListCursor { + readonly pointer: string; + readonly position: number; + readonly data: Schema.Json; + readonly variables: Variables; + readonly current: TaskFrame | null; +} + +export type Branch = + | { readonly state: 'running'; readonly task: TaskFrame } + | { readonly state: 'finished'; readonly output: Schema.Json; readonly flow: string } + | { readonly state: 'failed'; readonly error: DslError }; + +export type TryPhase = + | { readonly kind: 'trying'; readonly list: ListCursor; readonly attemptLimit: string | null } + | { readonly kind: 'backing_off'; readonly timer: string; readonly error: DslError } + | { readonly kind: 'recovering'; readonly list: ListCursor }; + +export type FrameBody = + | { readonly kind: 'list'; readonly list: ListCursor } + | { + readonly kind: 'for'; + readonly items: readonly Schema.Json[]; + readonly index: number; + readonly data: Schema.Json; + readonly list: ListCursor | null; + } + | { readonly kind: 'fork'; readonly compete: boolean; readonly branches: readonly Branch[] } + | { readonly kind: 'try'; readonly attempt: number; readonly startedAt: number; readonly phase: TryPhase } + | { readonly kind: 'wait'; readonly timer: string } + | { readonly kind: 'call'; readonly key: CallKey } + | { readonly kind: 'listen'; readonly consumed: readonly Schema.JsonObject[] } + | { readonly kind: 'yield'; readonly timer: string }; + +export interface TaskFrame { + readonly reference: string; + readonly run: number; + readonly rawInput: Schema.Json; + readonly input: Schema.Json; + readonly variables: Variables; + readonly timeout: string | null; + readonly body: FrameBody; +} + +export interface MachineState { + readonly context: Schema.Json; + readonly root: TaskFrame | null; +} + +export type RunOutcome = + | { readonly kind: 'completed'; readonly output: Schema.Json } + | { readonly kind: 'raised'; readonly error: DslError } + | { readonly kind: 'cancelled' } + | { readonly kind: 'broken'; readonly reason: string } + | { readonly kind: 'oversized'; readonly bytes: number; readonly most: number } + | { readonly kind: 'overran'; readonly milliseconds: number }; + +export interface ArmedTimer { + readonly purpose: TimerPurpose; + readonly reference: string; +} + +export interface WaitingEvent { + readonly event: ReceivedEvent; + readonly bytes: number; +} + +export interface InboxState { + readonly waiting: readonly WaitingEvent[]; + readonly waitingBytes: number; + readonly receivedIds: readonly string[]; + readonly received: number; + readonly receivedBytes: number; + readonly overflow: DslError | null; +} + +export interface RunState { + readonly executionId: string; + readonly status: 'new' | 'running' | 'ended'; + readonly workflow: { readonly document: Schema.JsonObject; readonly input: Schema.Json } | null; + readonly attributes: Schema.JsonObject; + readonly limits: RunLimits; + readonly startedAt: number; + readonly lastInputAt: number; + readonly random: { readonly seed: number; readonly draws: number }; + readonly timers: { readonly next: number; readonly armed: Readonly> }; + readonly calls: { + readonly runs: Readonly>; + readonly open: Readonly>; + }; + readonly inbox: InboxState; + readonly heldBytes: number; + readonly stepsWithoutWaiting: number; + readonly cancelRequested: boolean; + readonly machine: MachineState; + readonly outcome: RunOutcome | null; +} + +const IntSchema = Schema.Int; + +const VariablesSchema = Schema.Record(Schema.String, Schema.Json); + +const DslErrorSchema = Schema.Struct({ + type: Schema.String, + status: IntSchema, + instance: Schema.String, + title: Schema.optionalKey(Schema.String), + detail: Schema.optionalKey(Schema.String), +}); + +const ListCursorSchema: Schema.Codec = Schema.Struct({ + pointer: Schema.String, + position: IntSchema, + data: Schema.Json, + variables: VariablesSchema, + current: Schema.NullOr(Schema.suspend((): Schema.Codec => TaskFrameSchema)), +}); + +const BranchSchema: Schema.Codec = Schema.Union([ + Schema.Struct({ + state: Schema.Literal('running'), + task: Schema.suspend((): Schema.Codec => TaskFrameSchema), + }), + Schema.Struct({ state: Schema.Literal('finished'), output: Schema.Json, flow: Schema.String }), + Schema.Struct({ state: Schema.Literal('failed'), error: DslErrorSchema }), +]); + +const TryPhaseSchema: Schema.Codec = Schema.Union([ + Schema.Struct({ kind: Schema.Literal('trying'), list: ListCursorSchema, attemptLimit: Schema.NullOr(Schema.String) }), + Schema.Struct({ kind: Schema.Literal('backing_off'), timer: Schema.String, error: DslErrorSchema }), + Schema.Struct({ kind: Schema.Literal('recovering'), list: ListCursorSchema }), +]); + +const FrameBodySchema: Schema.Codec = Schema.Union([ + Schema.Struct({ kind: Schema.Literal('list'), list: ListCursorSchema }), + Schema.Struct({ + kind: Schema.Literal('for'), + items: Schema.Array(Schema.Json), + index: IntSchema, + data: Schema.Json, + list: Schema.NullOr(ListCursorSchema), + }), + Schema.Struct({ kind: Schema.Literal('fork'), compete: Schema.Boolean, branches: Schema.Array(BranchSchema) }), + Schema.Struct({ kind: Schema.Literal('try'), attempt: IntSchema, startedAt: IntSchema, phase: TryPhaseSchema }), + Schema.Struct({ kind: Schema.Literal('wait'), timer: Schema.String }), + Schema.Struct({ kind: Schema.Literal('call'), key: CallKeySchema }), + Schema.Struct({ kind: Schema.Literal('listen'), consumed: Schema.Array(Schema.JsonObject) }), + Schema.Struct({ kind: Schema.Literal('yield'), timer: Schema.String }), +]); + +const TaskFrameSchema: Schema.Codec = Schema.Struct({ + reference: Schema.String, + run: IntSchema, + rawInput: Schema.Json, + input: Schema.Json, + variables: VariablesSchema, + timeout: Schema.NullOr(Schema.String), + body: FrameBodySchema, +}); + +const RunOutcomeSchema: Schema.Codec = Schema.Union([ + Schema.Struct({ kind: Schema.Literal('completed'), output: Schema.Json }), + Schema.Struct({ kind: Schema.Literal('raised'), error: DslErrorSchema }), + Schema.Struct({ kind: Schema.Literal('cancelled') }), + Schema.Struct({ kind: Schema.Literal('broken'), reason: Schema.String }), + Schema.Struct({ kind: Schema.Literal('oversized'), bytes: IntSchema, most: IntSchema }), + Schema.Struct({ kind: Schema.Literal('overran'), milliseconds: IntSchema }), +]); + +export const RunStateSchema: Schema.Codec = Schema.Struct({ + executionId: Schema.String, + status: Schema.Literals(['new', 'running', 'ended']), + workflow: Schema.NullOr(Schema.Struct({ document: Schema.JsonObject, input: Schema.Json })), + attributes: Schema.JsonObject, + limits: RunLimitsSchema, + startedAt: IntSchema, + lastInputAt: IntSchema, + random: Schema.Struct({ seed: IntSchema, draws: IntSchema }), + timers: Schema.Struct({ + next: IntSchema, + armed: Schema.Record(Schema.String, Schema.Struct({ purpose: TimerPurposeSchema, reference: Schema.String })), + }), + calls: Schema.Struct({ + runs: Schema.Record(Schema.String, IntSchema), + open: Schema.Record(Schema.String, CallKeySchema), + }), + inbox: Schema.Struct({ + waiting: Schema.Array(Schema.Struct({ event: ReceivedEventSchema, bytes: IntSchema })), + waitingBytes: IntSchema, + receivedIds: Schema.Array(Schema.String), + received: IntSchema, + receivedBytes: IntSchema, + overflow: Schema.NullOr(DslErrorSchema), + }), + heldBytes: IntSchema, + stepsWithoutWaiting: IntSchema, + cancelRequested: Schema.Boolean, + machine: Schema.Struct({ context: Schema.Json, root: Schema.NullOr(TaskFrameSchema) }), + outcome: Schema.NullOr(RunOutcomeSchema), +}); + +export const newRun: RunState = { + executionId: '', + status: 'new', + workflow: null, + attributes: {}, + limits: { mostDurationMs: 1, longestCallMs: 1 }, + startedAt: 0, + lastInputAt: 0, + random: { seed: 0, draws: 0 }, + timers: { next: 1, armed: {} }, + calls: { runs: {}, open: {} }, + inbox: { waiting: [], waitingBytes: 0, receivedIds: [], received: 0, receivedBytes: 0, overflow: null }, + heldBytes: 0, + stepsWithoutWaiting: 0, + cancelRequested: false, + machine: { context: {}, root: null }, + outcome: null, +}; diff --git a/packages/workflow-engine/src/run-log/run-event.test.ts b/packages/workflow-engine/src/run-log/run-event.test.ts new file mode 100644 index 000000000..65ebd67e4 --- /dev/null +++ b/packages/workflow-engine/src/run-log/run-event.test.ts @@ -0,0 +1,81 @@ +import { Result, Schema } from 'effect'; +import { describe, expect, it } from 'vitest'; + +import { newRun, RunEventSchema, RunInputSchema, RunStateSchema, type RunEvent, type RunInput } from '../index.ts'; +import { at, executionId, openCall, runningState, started } from '../testing/runs.ts'; + +function asStored>(schema: S, value: S['Type']): unknown { + return JSON.parse(JSON.stringify(Schema.encodeUnknownSync(Schema.toCodecJson(schema))(value))); +} + +function readBack>( + schema: S, + stored: unknown, +): Result.Result { + return Schema.decodeUnknownResult(Schema.toCodecJson(schema))(stored); +} + +const event: RunEvent = { + type: 'input_applied', + receipt: { kind: 'call_answered', key: 'k', at }, + patch: [ + { op: 'replace', path: '/machine/context', value: { approved: true } }, + { op: 'add', path: '/timers/armed/e~1timers~13', value: { purpose: 'wait', reference: '/do/2' } }, + { op: 'remove', path: '/calls/open/k' }, + ], + outputs: [ + { kind: 'cancel_timer', executionId, timerId: `${executionId}/timers/2` }, + { kind: 'arm_timer', executionId, timerId: `${executionId}/timers/3`, dueAt: at + 1000, purpose: 'wait' }, + { kind: 'cancel_call', key: openCall }, + { kind: 'settle', executionId, settlement: { status: 'succeeded', output: { approved: true } } }, + ], +}; + +const inputs: readonly RunInput[] = [ + started, + { kind: 'timer_fired', executionId, at, timerId: `${executionId}/timers/1` }, + { + kind: 'call_answered', + executionId, + at, + key: openCall, + result: { status: 'rejected', reason: 'conflict', detail: 'x' }, + }, + { kind: 'event_received', executionId, at, event: { id: 'event-3', type: 'com.acme.approval', data: [1, 2] } }, + { kind: 'cancel_requested', executionId, at }, +]; + +describe('a run event', () => { + it('is stored as JSON and read back as it was decided', () => { + expect(readBack(RunEventSchema, asStored(RunEventSchema, event))).toEqual(Result.succeed(event)); + }); + + it('is refused when it holds an output the engine does not dispatch', () => { + const stored = { ...event, outputs: [{ kind: 'send_email', to: 'someone' }] }; + + expect(Result.isFailure(readBack(RunEventSchema, stored))).toBe(true); + }); +}); + +describe('a run input', () => { + it('of each kind is stored as JSON and read back as it was given', () => { + expect(inputs.map((input) => readBack(RunInputSchema, asStored(RunInputSchema, input)))).toEqual( + inputs.map((input) => Result.succeed(input)), + ); + }); + + it('is refused for an event without an id, or a time before the epoch', () => { + expect([ + Result.isFailure(readBack(RunInputSchema, { kind: 'event_received', executionId, at, event: { type: 't' } })), + Result.isFailure(readBack(RunInputSchema, { kind: 'cancel_requested', executionId, at: -1 })), + ]).toEqual([true, true]); + }); +}); + +describe('the state of a run', () => { + it('is plain JSON, new or running, so a snapshot is the state as it is', () => { + expect(JSON.parse(JSON.stringify(newRun))).toEqual(newRun); + expect(JSON.parse(JSON.stringify(runningState))).toEqual(runningState); + expect(readBack(RunStateSchema, asStored(RunStateSchema, runningState))).toEqual(Result.succeed(runningState)); + }); +}); diff --git a/packages/workflow-engine/src/run-log/run-event.ts b/packages/workflow-engine/src/run-log/run-event.ts new file mode 100644 index 000000000..1fc428c96 --- /dev/null +++ b/packages/workflow-engine/src/run-log/run-event.ts @@ -0,0 +1,19 @@ +import { Schema } from 'effect'; + +import { RunOutputSchema } from '../dispatch/run-output.ts'; +import { InputReceiptSchema } from '../machine/input-receipt.ts'; +import { PatchOperationSchema } from './state-patch.ts'; + +export const RunEventSchema = Schema.Struct({ + type: Schema.Literal('input_applied'), + receipt: InputReceiptSchema, + patch: Schema.Array(PatchOperationSchema), + outputs: Schema.Array(RunOutputSchema), +}); + +export type RunEvent = typeof RunEventSchema.Type; + +export interface PositionedEvent { + readonly version: number; + readonly event: RunEvent; +} diff --git a/packages/workflow-engine/src/run-log/run-store.ts b/packages/workflow-engine/src/run-log/run-store.ts new file mode 100644 index 000000000..188c52ad9 --- /dev/null +++ b/packages/workflow-engine/src/run-log/run-store.ts @@ -0,0 +1,32 @@ +import { Data, type Effect } from 'effect'; + +import type { RunState } from '../machine/run-state.ts'; +import type { PositionedEvent, RunEvent } from './run-event.ts'; +import type { SinceSnapshot, Snapshot } from './snapshot.ts'; + +export interface LoadedRun { + readonly state: RunState; + readonly version: number; + readonly sinceSnapshot: SinceSnapshot; +} + +export interface AppendedEvent { + readonly version: number; + readonly bytes: number; +} + +export class VersionConflict extends Data.TaggedError('version_conflict')<{ + readonly executionId: string; + readonly expectedVersion: number; +}> {} + +export interface RunStore { + readonly load: (executionId: string) => Effect.Effect; + readonly append: ( + executionId: string, + event: RunEvent, + expectedVersion: number, + ) => Effect.Effect; + readonly eventsAfter: (executionId: string, version: number) => Effect.Effect; + readonly saveSnapshot: (snapshot: Snapshot) => Effect.Effect; +} diff --git a/packages/workflow-engine/src/run-log/snapshot.test.ts b/packages/workflow-engine/src/run-log/snapshot.test.ts new file mode 100644 index 000000000..ff13befff --- /dev/null +++ b/packages/workflow-engine/src/run-log/snapshot.test.ts @@ -0,0 +1,46 @@ +import { Result } from 'effect'; +import { describe, expect, it } from 'vitest'; + +import { isSnapshotDue, snapshotChunkLength, snapshotChunks, snapshotFromChunks, type Snapshot } from '../index.ts'; +import { executionId, runningState } from '../testing/runs.ts'; + +const highSurrogate = /[\uD800-\uDBFF]$/u; + +const lowSurrogate = /^[\uDC00-\uDFFF]/u; + +function snapshotHolding(text: string): Snapshot { + return { format: 1, executionId, version: 4000, state: { ...runningState, machine: { context: text, root: null } } }; +} + +describe('a snapshot', () => { + it('is due after 1,000 inputs or 1 MiB of history since the last one, whichever comes first', () => { + expect([ + isSnapshotDue({ inputs: 999, bytes: 1_048_575 }), + isSnapshotDue({ inputs: 1000, bytes: 0 }), + isSnapshotDue({ inputs: 1, bytes: 1_048_576 }), + ]).toEqual([false, true, true]); + }); + + it('round-trips a running state through its chunks', () => { + const snapshot: Snapshot = { format: 1, executionId, version: 1000, state: runningState }; + + expect(snapshotFromChunks(snapshotChunks(snapshot))).toEqual(Result.succeed(snapshot)); + }); + + it('is stored in chunks of at most 65,536 UTF-16 code units that never split a character', () => { + const snapshot = snapshotHolding(`${'😀'.repeat(70_000)}x${'😀'.repeat(70_000)}é`); + + const chunks = snapshotChunks(snapshot); + + expect(chunks.length).toBeGreaterThan(2); + expect(chunks.every((chunk) => chunk.length <= snapshotChunkLength)).toBe(true); + expect(chunks.filter((chunk) => highSurrogate.test(chunk))).toEqual([]); + expect(chunks.filter((chunk) => lowSurrogate.test(chunk))).toEqual([]); + expect(Math.max(...chunks.map((chunk) => new TextEncoder().encode(chunk).byteLength))).toBeLessThanOrEqual(196_608); + expect(snapshotFromChunks(chunks)).toEqual(Result.succeed(snapshot)); + }); + + it('is refused when its chunks do not make a snapshot of this format', () => { + expect(Result.isFailure(snapshotFromChunks(['{"format":2}']))).toBe(true); + }); +}); diff --git a/packages/workflow-engine/src/run-log/snapshot.ts b/packages/workflow-engine/src/run-log/snapshot.ts new file mode 100644 index 000000000..4467d5702 --- /dev/null +++ b/packages/workflow-engine/src/run-log/snapshot.ts @@ -0,0 +1,54 @@ +import { Schema } from 'effect'; + +import { RunStateSchema } from '../machine/run-state.ts'; + +export const snapshotEveryInputs = 1000; + +export const snapshotEveryBytes = 1_048_576; + +export const snapshotChunkLength = 65_536; + +export const SnapshotSchema = Schema.Struct({ + format: Schema.Literal(1), + executionId: Schema.NonEmptyString, + version: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), + state: RunStateSchema, +}); + +export type Snapshot = typeof SnapshotSchema.Type; + +export interface SinceSnapshot { + readonly inputs: number; + readonly bytes: number; +} + +const SnapshotTextSchema = Schema.fromJsonString(SnapshotSchema); + +const encodeSnapshot = Schema.encodeSync(SnapshotTextSchema); + +export const decodeSnapshot = Schema.decodeUnknownResult(SnapshotTextSchema); + +export function isSnapshotDue({ inputs, bytes }: SinceSnapshot): boolean { + return inputs >= snapshotEveryInputs || bytes >= snapshotEveryBytes; +} + +const lastSingleUnitCodePoint = 0xff_ff; + +function chunkEnd(text: string, start: number): number { + const end = Math.min(start + snapshotChunkLength, text.length); + const splitsAPair = Number(text.codePointAt(end - 1)) > lastSingleUnitCodePoint; + return splitsAPair ? end - 1 : end; +} + +export function snapshotChunks(snapshot: Snapshot): readonly string[] { + const text = encodeSnapshot(snapshot); + const chunks: string[] = []; + for (let start = 0; start < text.length; start = chunkEnd(text, start)) { + chunks.push(text.slice(start, chunkEnd(text, start))); + } + return chunks; +} + +export function snapshotFromChunks(chunks: readonly string[]): ReturnType { + return decodeSnapshot(chunks.join('')); +} diff --git a/packages/workflow-engine/src/run-log/state-patch.ts b/packages/workflow-engine/src/run-log/state-patch.ts new file mode 100644 index 000000000..31e0e1ca5 --- /dev/null +++ b/packages/workflow-engine/src/run-log/state-patch.ts @@ -0,0 +1,13 @@ +import { Schema } from 'effect'; + +const PointerSchema = Schema.String; + +export const PatchOperationSchema = Schema.Union([ + Schema.Struct({ op: Schema.Literal('add'), path: PointerSchema, value: Schema.Json }), + Schema.Struct({ op: Schema.Literal('replace'), path: PointerSchema, value: Schema.Json }), + Schema.Struct({ op: Schema.Literal('remove'), path: PointerSchema }), +]); + +export type PatchOperation = typeof PatchOperationSchema.Type; + +export type StatePatch = readonly PatchOperation[]; diff --git a/packages/workflow-engine/src/serialisation/run-serialiser.ts b/packages/workflow-engine/src/serialisation/run-serialiser.ts new file mode 100644 index 000000000..05d6b453f --- /dev/null +++ b/packages/workflow-engine/src/serialisation/run-serialiser.ts @@ -0,0 +1,5 @@ +import type { Effect } from 'effect'; + +export interface RunSerialiser { + readonly serialise: (executionId: string, work: Effect.Effect) => Effect.Effect; +} diff --git a/packages/workflow-engine/src/settlement/record-store.ts b/packages/workflow-engine/src/settlement/record-store.ts new file mode 100644 index 000000000..86838d88a --- /dev/null +++ b/packages/workflow-engine/src/settlement/record-store.ts @@ -0,0 +1,10 @@ +import type { Effect } from 'effect'; + +import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; +import type { Settle } from '../dispatch/run-output.ts'; + +export type SettleReceipt = 'recorded' | 'already_recorded' | 'settled_otherwise' | 'unknown_execution'; + +export interface RecordStore { + readonly settle: (settle: Settle, run: RunContext) => Effect.Effect; +} diff --git a/packages/workflow-engine/src/settlement/run-settlement.ts b/packages/workflow-engine/src/settlement/run-settlement.ts new file mode 100644 index 000000000..8d9305490 --- /dev/null +++ b/packages/workflow-engine/src/settlement/run-settlement.ts @@ -0,0 +1,13 @@ +import { Schema } from 'effect'; + +export const RunSettlementSchema = Schema.Union([ + Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), + Schema.Struct({ + status: Schema.Literal('rejected'), + reason: Schema.Literals(['invalid_input', 'unavailable']), + detail: Schema.String, + }), + Schema.Struct({ status: Schema.Literal('failed') }), +]); + +export type RunSettlement = typeof RunSettlementSchema.Type; diff --git a/packages/workflow-engine/src/testing/runs.ts b/packages/workflow-engine/src/testing/runs.ts new file mode 100644 index 000000000..8816ef130 --- /dev/null +++ b/packages/workflow-engine/src/testing/runs.ts @@ -0,0 +1,105 @@ +import { callKeyText, type CallKey } from '../executor/call-key.ts'; +import type { RunInput } from '../machine/run-input.ts'; +import { newRun, type RunState, type TaskFrame } from '../machine/run-state.ts'; + +export const executionId = '0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a'; + +export const openCall: CallKey = { executionId, reference: '/do/1/fork/branches/0/ask', run: 1 }; + +export const armedTimer = `${executionId}/timers/1`; + +const waiting: TaskFrame = { + reference: '/do/1/fork/branches/1/pause', + run: 1, + rawInput: { ticket: 7 }, + input: { ticket: 7 }, + variables: {}, + timeout: null, + body: { kind: 'wait', timer: armedTimer }, +}; + +const asking: TaskFrame = { + reference: openCall.reference, + run: 1, + rawInput: { ticket: 7 }, + input: { ticket: 7 }, + variables: { attempt: 1 }, + timeout: `${executionId}/timers/2`, + body: { kind: 'call', key: openCall }, +}; + +const forking: TaskFrame = { + reference: '/do/1', + run: 1, + rawInput: { ticket: 7 }, + input: { ticket: 7 }, + variables: {}, + timeout: null, + body: { + kind: 'fork', + compete: true, + branches: [ + { state: 'running', task: asking }, + { state: 'running', task: waiting }, + { state: 'failed', error: { type: 'runtime', status: 500, instance: '/do/1/fork/branches/2' } }, + ], + }, +}; + +export const runningState: RunState = { + ...newRun, + executionId, + status: 'running', + workflow: { document: { document: { dsl: '1.0.3' }, do: [] }, input: { ticket: 7 } }, + attributes: { owner: 'tests' }, + limits: { mostDurationMs: 2_592_000_000, longestCallMs: 600_000 }, + startedAt: 1_791_100_000_000, + lastInputAt: 1_791_100_000_000, + random: { seed: 42, draws: 1 }, + timers: { + next: 3, + armed: { + [armedTimer]: { purpose: 'wait', reference: waiting.reference }, + [`${executionId}/timers/2`]: { purpose: 'timeout', reference: asking.reference }, + }, + }, + calls: { runs: { [openCall.reference]: 1 }, open: { [callKeyText(openCall)]: openCall } }, + inbox: { + waiting: [{ event: { id: 'event-2', type: 'com.acme.approval', data: { approved: true } }, bytes: 80 }], + waitingBytes: 80, + receivedIds: ['event-1', 'event-2'], + received: 2, + receivedBytes: 160, + overflow: null, + }, + heldBytes: 9000, + stepsWithoutWaiting: 0, + machine: { + context: { seen: 1 }, + root: { + reference: '/', + run: 1, + rawInput: { ticket: 7 }, + input: { ticket: 7 }, + variables: {}, + timeout: null, + body: { + kind: 'list', + list: { pointer: '/do', position: 1, data: { ticket: 7 }, variables: {}, current: forking }, + }, + }, + }, +}; + +export const at = 1_791_100_060_000; + +export const started: RunInput = { + kind: 'started', + executionId, + at, + document: { document: { dsl: '1.0.3' }, do: [] }, + input: { ticket: 7 }, + limits: { mostDurationMs: 2_592_000_000, longestCallMs: 600_000 }, + attributes: { owner: 'tests' }, + seed: 42, +}; diff --git a/packages/workflow-engine/src/timers/timer-id.ts b/packages/workflow-engine/src/timers/timer-id.ts new file mode 100644 index 000000000..85c8d202a --- /dev/null +++ b/packages/workflow-engine/src/timers/timer-id.ts @@ -0,0 +1,16 @@ +import { Schema } from 'effect'; + +export const TimerPurposeSchema = Schema.Literals([ + 'wait', + 'timeout', + 'retry_delay', + 'attempt_limit', + 'deadline', + 'yield', +]); + +export type TimerPurpose = typeof TimerPurposeSchema.Type; + +export function timerIdOf(executionId: string, sequence: number): string { + return `${executionId}/timers/${sequence}`; +} diff --git a/packages/workflow-engine/src/timers/timers.ts b/packages/workflow-engine/src/timers/timers.ts new file mode 100644 index 000000000..db4f9cf27 --- /dev/null +++ b/packages/workflow-engine/src/timers/timers.ts @@ -0,0 +1,9 @@ +import type { Effect } from 'effect'; + +import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; +import type { ArmTimer, CancelTimer } from '../dispatch/run-output.ts'; + +export interface Timers { + readonly arm: (timer: ArmTimer, run: RunContext) => Effect.Effect; + readonly cancel: (timer: CancelTimer, run: RunContext) => Effect.Effect; +} diff --git a/packages/workflow-engine/tsconfig.json b/packages/workflow-engine/tsconfig.json new file mode 100644 index 000000000..000610b26 --- /dev/null +++ b/packages/workflow-engine/tsconfig.json @@ -0,0 +1,4 @@ +{ + "extends": "../../tsconfig.base.json", + "include": ["*.ts", "src"] +} diff --git a/packages/workflow-engine/vitest.config.ts b/packages/workflow-engine/vitest.config.ts new file mode 100644 index 000000000..64c238962 --- /dev/null +++ b/packages/workflow-engine/vitest.config.ts @@ -0,0 +1,5 @@ +import { defineConfig, mergeConfig } from 'vitest/config'; + +import { sharedConfig } from '../../vitest.shared.ts'; + +export default mergeConfig(sharedConfig, defineConfig({ test: { name: 'workflow-engine' } })); diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 777edd9ca..115498ede 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -446,6 +446,22 @@ importers: specifier: 'catalog:' version: 5.0.2(@opentelemetry/api@1.9.1)(@types/node@26.6.3)(@vitest/coverage-v8@5.0.2)(vite@8.3.1(@types/node@26.6.3)(esbuild@0.28.1)(jiti@2.7.0)(terser@5.51.2)(yaml@2.9.1)) + packages/workflow-engine: + dependencies: + '@beonauto/operations': + specifier: workspace:* + version: link:../operations + effect: + specifier: 'catalog:' + version: 4.0.0 + devDependencies: + '@vitest/coverage-v8': + specifier: 'catalog:' + version: 5.0.2(vitest@5.0.2) + vitest: + specifier: 'catalog:' + version: 5.0.2(@opentelemetry/api@1.9.1)(@types/node@26.6.3)(@vitest/coverage-v8@5.0.2)(vite@8.3.1(@types/node@26.6.3)(esbuild@0.28.1)(jiti@2.7.0)(terser@5.51.2)(yaml@2.9.1)) + primitives/inference: dependencies: '@ai-sdk/amazon-bedrock': From 2001985310b2dab023d8b049469f3e3cf707472e Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 18:14:10 +0100 Subject: [PATCH 02/23] docs(global): record the decision to run workflows on the ledger docs/decisions/0001-workflow-engine-on-the-ledger.md is the first decision record: why workflows move from Temporal to an engine on the ledger, the shape of that engine, what we give up and must keep correct, and the measurements of the two spikes it rests on, cited by branch and file. Co-Authored-By: Claude Opus 5.5 --- .../0001-workflow-engine-on-the-ledger.md | 57 +++++++++++++++++++ 1 file changed, 57 insertions(+) create mode 100644 docs/decisions/0001-workflow-engine-on-the-ledger.md diff --git a/docs/decisions/0001-workflow-engine-on-the-ledger.md b/docs/decisions/0001-workflow-engine-on-the-ledger.md new file mode 100644 index 000000000..373a2cee7 --- /dev/null +++ b/docs/decisions/0001-workflow-engine-on-the-ledger.md @@ -0,0 +1,57 @@ +# 1. Run workflows on an engine on the ledger, not on Temporal + +- Status: accepted +- Date: 2026-10-04 + +## Context + +Workflow specs run on Temporal today, the wrong place for them: + +- Temporal cannot run on Cloudflare Workers, where Auto's cloud hosting runs auto-brain: there is no Temporal worker for workerd. +- Nobody asked for Temporal. People ask for workflows that wait, retry and survive a restart, and a self-hosted server must run a Temporal service beside it to get them. +- Tenant data is stored twice: Temporal's history holds each workflow's document, input, the outputs of its calls and its events, unencrypted, beside the brain's ledger. +- We already have the store: the ledger keeps every brain's events with Emmett on SQLite, on the sqlite3, D1 and Durable Object drivers. + +## Decision + +A run is a decider in Emmett's workflow shape, on the load-decide-append loop the ledger already uses: + +- `decide(input, state)` says what happened and `evolve(state, event)` folds it. The inputs are a start, a timer fired, a call answered, an event received and a cancel request, each with the time it arrived. +- One stream per run. An applied input appends one event, with the version it was decided on as the expected version. The event holds the change to the state as a JSON Patch, so replay evaluates nothing, and the outputs: arm or cancel a timer, start or cancel a call, settle. +- Outputs are dispatched after the append, behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. +- Deduplication lives in the run's state. Emmett is the store, never the engine: we use neither its workflow handler, which folds the whole stream for every input, nor its processors. +- A snapshot follows every 1,000 inputs or 1 MiB of events, whichever comes first; only the latest is kept, in chunks of at most 192 KiB. A run may hold 4 MiB instead of 16. +- Four adapters sit behind small ports: run store, timers, executor and record store, with the watermark and per-run serialisation beside them. +- On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes runs that fell behind. +- Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process. + +## Consequences + +We give up Temporal's durable timers, deduplicated delivery, replay, web UI and operator tools. We must build and keep correct: + +- Timers that fire at least once, on a table in Node and on the object's alarm on Cloudflare; a fire of a timer no longer armed changes nothing. +- Deduplication in state, every key bounded, or snapshots grow with the run. +- The outbox watermark: a crash between append and dispatch loses nothing. +- The sweep on Cloudflare, for alarms that fire late after eviction or give up after their retries. +- Serialisation per run: an in-process lock in Node, the object's thread on Cloudflare. +- Settlement in two stores, the run's stream and the brain's record, each idempotent by execution id, the second retried. + +Tenant data is stored once, and a workflow needs no service beyond the server. + +## Evidence + +Branch `spike/engine-node`: + +- `spikes/node/results/replay.json`: the interpreter as it runs on Temporal replays 40,000 inputs in 4.6 s and retains up to 103 MiB. +- `spikes/node/results/message-id-probe.json`: Emmett appends a message with an id it has seen as a new message. +- `spikes/node/results/executor-virtual.json`, `executor-real.json`: a result delivered four times settles once, a result after its timeout is ignored, a crash after the append is recovered on wake. +- `spikes/node/results/timers-precision.json`, `timers-recovery.json`: timers fire 3.7 ms late at p99 when idle; after a killed scheduler all 200 fire, none twice. +- `spikes/node/results/timers-two-processes.json`: two processes double-fire 227 of 300 timers unless each claims a timer first. +- `spikes/node/results/lost-write-repeat.json`: timers written through a second SQLite library to the ledger's file lost committed cancels in three runs of three. + +Branch `spike/engine-cloudflare`: + +- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. +- `spikes/cloudflare/results/timers.json`: alarms fire 5 ms late at p99; an evicted object's alarm fired 15.6 s late; a sweep re-armed one that had given up. +- `spikes/cloudflare/results/settlement.json`: the record was written exactly once, or given up as intended, under every injected D1 fault and crash; D1 refuses eleven events in one append. +- `spikes/cloudflare/results/portability.json`: the interpreter, the DSL policy, jq and both Cloudflare ledger drivers run in workerd; the workflow SDK's validators run once precompiled. From 23d8c20a23bd7303911f8f02996b444872c3a373 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 18:46:30 +0100 Subject: [PATCH 03/23] feat(ledger): read a stream after a version and share the decision loop The workflow engine keeps one stream per run and loads it from its latest snapshot, so it must read only the events after a version, and it must decide and append the way the ledger does rather than through a copy. EventStore.read takes an optional version to read after; Emmett answers a read past the end of a stream with version 0, so read answers with that version instead. The retries after a version conflict move into retriedOnVersionConflict, which the ledger's own execute now uses, and the index exports it with VersionConflict, the event store, its appender and its codec, and sqliteEventStore, which opens the store on any of Emmett's SQLite drivers without the layer. Co-Authored-By: Claude Opus 5.5 --- packages/ledger/README.md | 11 +++++- packages/ledger/src/event-store.test.ts | 42 +++++++++++++++++++++++ packages/ledger/src/event-store.ts | 4 +-- packages/ledger/src/index.ts | 6 +++- packages/ledger/src/ledger-service.ts | 13 ++----- packages/ledger/src/sqlite-event-store.ts | 10 +++--- packages/ledger/src/version-conflict.ts | 14 +++++++- 7 files changed, 79 insertions(+), 21 deletions(-) create mode 100644 packages/ledger/src/event-store.test.ts diff --git a/packages/ledger/README.md b/packages/ledger/README.md index 7f346e1db..033db9b41 100644 --- a/packages/ledger/README.md +++ b/packages/ledger/README.md @@ -20,6 +20,15 @@ This package is the event store behind the `Ledger` port of `@beonauto/operation - A decision of more than eight events is a defect. - Any other failure of the database is a defect, never a `Conflict` or a rejection. +## Building on the ledger's loop + +Code that keeps its own streams, such as `@beonauto/workflow-engine`, uses the same pieces as the ledger itself rather than a copy of them: + +- `sqliteEventStore(optionsOf)` opens the event store on any of Emmett's SQLite drivers, without the layer. +- `EventStore.read(stream, after)` gives the events after version `after` and the version of the whole stream, so a reader that holds a snapshot at version `after` reads only the tail. Emmett answers a read past the end of a stream with version 0; `read` answers with `after` instead. +- `eventAppenderOf(store)` encodes and appends events with an expected version, at most eight in one append, and fails with `VersionConflict` when another writer appended first. +- `retriedOnVersionConflict(attempt)` runs a load-decide-append attempt again after a version conflict, up to three more times, and then fails with `Conflict`. + ## Creating the layer ```ts @@ -36,7 +45,7 @@ Each SQLite connection may cache up to 8 MiB of pages and maps none of the file The same ledger will run on Cloudflare D1, so it follows these rules: -- It uses only the event store's own operations: read a stream, append with an expected version, migrate, close. It writes no SQL and registers no projections or consumers. +- It uses only the event store's own operations: read a stream, or its tail after a version, append with an expected version, migrate, close. It writes no SQL and registers no projections or consumers. - It does not rely on transactions or rollback: each command makes at most one append. - An append carries at most eight events. Emmett binds ten parameters for each event it inserts, and D1 accepts at most 100 bound parameters in one query. - Only `src/open-event-store.ts` knows which SQLite driver is in use and that the database is a file. diff --git a/packages/ledger/src/event-store.test.ts b/packages/ledger/src/event-store.test.ts new file mode 100644 index 000000000..18e712cbe --- /dev/null +++ b/packages/ledger/src/event-store.test.ts @@ -0,0 +1,42 @@ +import { sqlite3EventStoreDriver } from '@event-driven-io/emmett-sqlite/sqlite3'; +import { afterEach, describe, expect, it } from 'vitest'; + +import { sqliteEventStore, type EventStore } from './index.ts'; + +const opened: EventStore[] = []; + +async function aStore(): Promise { + const store = sqliteEventStore(() => ({ driver: sqlite3EventStoreDriver, fileName: ':memory:' })); + opened.push(store); + await store.migrate(); + return store; +} + +afterEach(async () => { + await Promise.all(opened.splice(0).map((store) => store.close())); +}); + +const run = 'run/0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a'; + +function numbered(from: number, count: number) { + return Array.from({ length: count }, (_, index) => ({ type: 'counted', data: { n: from + index } })); +} + +describe('reading a stream after a version', () => { + it('gives the events after that version, and the version of the whole stream', async () => { + const store = await aStore(); + await store.append(run, numbered(1, 3), 0); + await store.append(run, numbered(4, 2), 3); + + expect(await store.read(run, 3)).toEqual({ version: 5, events: [{ n: 4 }, { n: 5 }] }); + expect(await store.read(run)).toEqual({ version: 5, events: [1, 2, 3, 4, 5].map((n) => ({ n })) }); + }); + + it('gives no events and the version it was asked after when nothing follows it', async () => { + const store = await aStore(); + await store.append(run, numbered(1, 3), 0); + + expect(await store.read(run, 3)).toEqual({ version: 3, events: [] }); + expect(await store.read('run/nobody-wrote', 0)).toEqual({ version: 0, events: [] }); + }); +}); diff --git a/packages/ledger/src/event-store.ts b/packages/ledger/src/event-store.ts index 0bce9e95e..9ecd7b7b8 100644 --- a/packages/ledger/src/event-store.ts +++ b/packages/ledger/src/event-store.ts @@ -5,13 +5,13 @@ export interface EncodedEvent { readonly data: Schema.JsonObject; } -interface RecordedStream { +export interface RecordedStream { readonly version: number; readonly events: readonly unknown[]; } export interface EventStore { - readonly read: (stream: string) => Promise; + readonly read: (stream: string, after?: number) => Promise; readonly append: (stream: string, events: readonly EncodedEvent[], expectedVersion: number) => Promise; readonly migrate: () => Promise; readonly close: () => Promise; diff --git a/packages/ledger/src/index.ts b/packages/ledger/src/index.ts index d550613cd..5b69aced3 100644 --- a/packages/ledger/src/index.ts +++ b/packages/ledger/src/index.ts @@ -1 +1,5 @@ -export { sqliteLedgerLayer, type SQLiteStoreOptions } from './sqlite-event-store.ts'; +export { eventAppenderOf, type EventAppender } from './event-appender.ts'; +export { eventCodecOf, type EventCodec } from './event-codec.ts'; +export type { EncodedEvent, EventStore, RecordedStream } from './event-store.ts'; +export { sqliteEventStore, sqliteLedgerLayer, type SQLiteStoreOptions } from './sqlite-event-store.ts'; +export { retriedOnVersionConflict, VersionConflict } from './version-conflict.ts'; diff --git a/packages/ledger/src/ledger-service.ts b/packages/ledger/src/ledger-service.ts index 94a64742f..d3e76e9ce 100644 --- a/packages/ledger/src/ledger-service.ts +++ b/packages/ledger/src/ledger-service.ts @@ -1,5 +1,4 @@ import { - Conflict, Ledger, type Decider, type DeclarableReason, @@ -13,11 +12,7 @@ import { eventAppenderOf } from './event-appender.ts'; import type { EventStore } from './event-store.ts'; import { foldEvents } from './fold-events.ts'; import { streamReaderOf } from './stream-reader.ts'; -import type { VersionConflict } from './version-conflict.ts'; - -const retriesOnVersionConflict = 3; - -const changedWhileDeciding = 'The state changed while the command was decided'; +import { retriedOnVersionConflict, type VersionConflict } from './version-conflict.ts'; export function makeLedger(store: EventStore): Ledger['Service'] { const load = streamReaderOf(store); @@ -43,10 +38,6 @@ export function makeLedger(store: EventStore): Ledger['Service'] { return Ledger.of({ load, execute: (stream, decider, command) => - attempt(stream, decider, command).pipe( - Effect.retry({ times: retriesOnVersionConflict }), - Effect.mapError(() => new Conflict({ detail: changedWhileDeciding, kind: 'concurrent_change' })), - Effect.flatMap(Effect.fromResult), - ), + retriedOnVersionConflict(attempt(stream, decider, command)).pipe(Effect.flatMap(Effect.fromResult)), }); } diff --git a/packages/ledger/src/sqlite-event-store.ts b/packages/ledger/src/sqlite-event-store.ts index 2aa66d011..f92c09fe7 100644 --- a/packages/ledger/src/sqlite-event-store.ts +++ b/packages/ledger/src/sqlite-event-store.ts @@ -9,13 +9,13 @@ type AnyDriver = Parameters[0]['driver']; export type SQLiteStoreOptions = Parameters>[0]; -function openSQLiteEventStore(optionsOf: () => SQLiteStoreOptions): EventStore { +export function sqliteEventStore(optionsOf: () => SQLiteStoreOptions): EventStore { const store = getSQLiteEventStore({ ...optionsOf(), schema: { autoMigration: 'None' } }); return { - read: async (stream) => { - const { currentStreamVersion, events } = await store.readStream(stream); + read: async (stream, after = 0) => { + const { currentStreamVersion, events } = await store.readStream(stream, { from: BigInt(after + 1) }); return { - version: Number(currentStreamVersion), + version: Math.max(after, Number(currentStreamVersion)), events: events.map(({ data }: { readonly data: unknown }) => data), }; }, @@ -32,5 +32,5 @@ function openSQLiteEventStore(optionsOf: () => SQLiteS export function sqliteLedgerLayer( optionsOf: () => SQLiteStoreOptions, ): Layer.Layer { - return ledgerLayerOver(() => openSQLiteEventStore(optionsOf)); + return ledgerLayerOver(() => sqliteEventStore(optionsOf)); } diff --git a/packages/ledger/src/version-conflict.ts b/packages/ledger/src/version-conflict.ts index bce10a25b..0d03183ad 100644 --- a/packages/ledger/src/version-conflict.ts +++ b/packages/ledger/src/version-conflict.ts @@ -1,3 +1,15 @@ -import { Data } from 'effect'; +import { Conflict } from '@beonauto/operations'; +import { Data, Effect } from 'effect'; export class VersionConflict extends Data.TaggedError('version_conflict') {} + +const retriesOnVersionConflict = 3; + +const changedWhileDeciding = 'The state changed while the command was decided'; + +export function retriedOnVersionConflict(attempt: Effect.Effect): Effect.Effect { + return attempt.pipe( + Effect.retry({ times: retriesOnVersionConflict }), + Effect.mapError(() => new Conflict({ detail: changedWhileDeciding, kind: 'concurrent_change' })), + ); +} From 7cf7118dc24fcf08210509b82bf7afe2609a54e8 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:03:12 +0100 Subject: [PATCH 04/23] feat(workflow-engine): revise the contract after its review - Values a run holds live once in a value table and frames refer to them by id, so a patch adds an answer once however many places it lands; events are measured in UTF-8 bytes against 1.5 MiB. - Admission names a stale reason, answers not_started for an input to a run that has not started, and dies on another execution or a second start with another document. - Every event and snapshot names its state format; patches apply strictly and the fold decodes the state once, from the snapshot and the tail the run store returns. - Dispatch runs in stream order and the watermark stops before the first event whose output failed; cancels of unseen keys tombstone. - Limits gain 100,000 inputs, 512 MiB of history and the expression work budget; snapshots are due at max(1 MiB, the last snapshot's size) and chunked at 1 MiB of UTF-8. - Events record the receipt with an answer's status or an event's type, and the steps an input ran; `at` is clamped to never go back. - The run store's append fails with the ledger's VersionConflict, and the engine's submit with the ledger's Conflict. - Purity and no-cache scans cover the machine and the run log. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/package.json | 1 + .../src/dispatch/dispatch-watermark.test.ts | 42 +++++- .../src/dispatch/dispatch-watermark.ts | 12 +- .../src/dispatch/run-output.ts | 11 +- .../src/engine/portability.test.ts | 82 ++++++++++-- .../src/engine/workflow-engine.ts | 15 ++- .../src/executor/call-result.ts | 4 + .../workflow-engine/src/executor/executor.ts | 8 +- packages/workflow-engine/src/index.ts | 68 ++++++++-- .../src/machine/admission.test.ts | 81 +++++++++++ .../workflow-engine/src/machine/admission.ts | 70 ++++++++++ .../src/machine/input-receipt.test.ts | 74 +++------- .../src/machine/input-receipt.ts | 55 ++++---- .../workflow-engine/src/machine/instant.ts | 7 + .../workflow-engine/src/machine/limits.ts | 12 ++ .../workflow-engine/src/machine/run-input.ts | 3 +- .../workflow-engine/src/machine/run-state.ts | 126 ++++++++++++------ .../workflow-engine/src/machine/same-json.ts | 25 ++++ .../src/run-log/run-event.test.ts | 49 +++++-- .../workflow-engine/src/run-log/run-event.ts | 23 ++++ .../src/run-log/run-fold.test.ts | 97 ++++++++++++++ .../workflow-engine/src/run-log/run-fold.ts | 67 ++++++++++ .../workflow-engine/src/run-log/run-store.ts | 25 ++-- .../src/run-log/snapshot.test.ts | 60 ++++++--- .../workflow-engine/src/run-log/snapshot.ts | 42 ++++-- .../src/run-log/state-format.ts | 5 + .../src/run-log/state-patch.test.ts | 51 +++++++ .../src/run-log/state-patch.ts | 112 +++++++++++++++- .../src/settlement/record-store.ts | 23 +++- .../src/settlement/run-settlement.ts | 13 -- packages/workflow-engine/src/testing/runs.ts | 69 ++++++---- packages/workflow-engine/src/timers/timers.ts | 9 +- pnpm-lock.yaml | 3 + 33 files changed, 1068 insertions(+), 276 deletions(-) create mode 100644 packages/workflow-engine/src/machine/admission.test.ts create mode 100644 packages/workflow-engine/src/machine/admission.ts create mode 100644 packages/workflow-engine/src/machine/instant.ts create mode 100644 packages/workflow-engine/src/machine/same-json.ts create mode 100644 packages/workflow-engine/src/run-log/run-fold.test.ts create mode 100644 packages/workflow-engine/src/run-log/run-fold.ts create mode 100644 packages/workflow-engine/src/run-log/state-format.ts create mode 100644 packages/workflow-engine/src/run-log/state-patch.test.ts delete mode 100644 packages/workflow-engine/src/settlement/run-settlement.ts diff --git a/packages/workflow-engine/package.json b/packages/workflow-engine/package.json index 27d80d8a7..1eb4cb41f 100644 --- a/packages/workflow-engine/package.json +++ b/packages/workflow-engine/package.json @@ -13,6 +13,7 @@ "test": "vitest run --coverage" }, "dependencies": { + "@beonauto/ledger": "workspace:*", "@beonauto/operations": "workspace:*", "effect": "catalog:" }, diff --git a/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts index b7fb0e4f7..9cd60336c 100644 --- a/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts +++ b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts @@ -1,6 +1,13 @@ import { describe, expect, it } from 'vitest'; -import { DispatchFailed, outputsAbove, type PositionedEvent, type RunOutput } from '../index.ts'; +import { + DispatchFailed, + dispatchedThrough, + outputsAbove, + stateFormat, + type PositionedEvent, + type RunOutput, +} from '../index.ts'; import { at, executionId, openCall } from '../testing/runs.ts'; const arm: RunOutput = { @@ -18,20 +25,45 @@ const settle: RunOutput = { kind: 'settle', executionId, settlement: { status: ' function applied(version: number, outputs: readonly RunOutput[]): PositionedEvent { return { version, - event: { type: 'input_applied', receipt: { kind: 'timer_fired', key: 'k', at }, patch: [], outputs }, + bytes: 100, + event: { + type: 'input_applied', + format: stateFormat, + receipt: { kind: 'timer_fired', key: 'k', at }, + steps: [], + patch: [], + outputs, + }, }; } +const events = [applied(3, [settle]), applied(1, [start]), applied(2, [arm, start]), applied(4, [])]; + describe('the outputs above a watermark', () => { it('are those of the events after it, in the order of the stream', () => { - const events = [applied(3, [settle]), applied(1, [start]), applied(2, [arm, start])]; - expect(outputsAbove(1, events)).toEqual([ { version: 2, output: arm }, { version: 2, output: start }, { version: 3, output: settle }, ]); - expect(outputsAbove(3, events)).toEqual([]); + expect(outputsAbove(4, events)).toEqual([]); + }); +}); + +describe('the watermark after a dispatch', () => { + it('moves to the last event when every output was dispatched, even past events with none', () => { + expect(dispatchedThrough(1, events)).toBe(4); + }); + + it('stops before the event of the first output that failed, so the next wake starts there', () => { + expect(dispatchedThrough(1, events, 3)).toBe(2); + expect(dispatchedThrough(1, events, 2)).toBe(1); + }); + + it('never goes down', () => { + expect([dispatchedThrough(3, events, 2), dispatchedThrough(5, events), dispatchedThrough(2, [])]).toEqual([ + 3, 5, 2, + ]); }); }); diff --git a/packages/workflow-engine/src/dispatch/dispatch-watermark.ts b/packages/workflow-engine/src/dispatch/dispatch-watermark.ts index bf84d850c..71c57d55d 100644 --- a/packages/workflow-engine/src/dispatch/dispatch-watermark.ts +++ b/packages/workflow-engine/src/dispatch/dispatch-watermark.ts @@ -23,9 +23,17 @@ export class DispatchFailed extends Data.TaggedError('dispatch_failed')<{ readonly detail: string; }> {} +function inStreamOrder(events: readonly PositionedEvent[]): readonly PositionedEvent[] { + return events.toSorted((first, second) => first.version - second.version); +} + export function outputsAbove(watermark: number, events: readonly PositionedEvent[]): readonly PositionedOutput[] { - return events + return inStreamOrder(events) .filter(({ version }) => version > watermark) - .toSorted((first, second) => first.version - second.version) .flatMap(({ version, event }) => event.outputs.map((output) => ({ version, output }))); } + +export function dispatchedThrough(watermark: number, events: readonly PositionedEvent[], firstFailed?: number): number { + const last = Math.max(watermark, ...events.map(({ version }) => version)); + return firstFailed === undefined ? last : Math.max(watermark, firstFailed - 1); +} diff --git a/packages/workflow-engine/src/dispatch/run-output.ts b/packages/workflow-engine/src/dispatch/run-output.ts index f9ae948bb..d1cfd0d6c 100644 --- a/packages/workflow-engine/src/dispatch/run-output.ts +++ b/packages/workflow-engine/src/dispatch/run-output.ts @@ -1,18 +1,17 @@ import { Schema } from 'effect'; import { CallKeySchema } from '../executor/call-key.ts'; -import { RunSettlementSchema } from '../settlement/run-settlement.ts'; +import { InstantSchema } from '../machine/instant.ts'; +import { SettlementSchema } from '../settlement/record-store.ts'; import { TimerPurposeSchema } from '../timers/timer-id.ts'; const ExecutionIdSchema = Schema.NonEmptyString; -const MillisecondsSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); - const ArmTimerSchema = Schema.Struct({ kind: Schema.Literal('arm_timer'), executionId: ExecutionIdSchema, timerId: Schema.NonEmptyString, - dueAt: MillisecondsSchema, + dueAt: InstantSchema, purpose: TimerPurposeSchema, }); @@ -27,7 +26,7 @@ const StartCallSchema = Schema.Struct({ key: CallKeySchema, function: Schema.NonEmptyString, arguments: Schema.Json, - longestMs: MillisecondsSchema, + longestMs: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), }); const CancelCallSchema = Schema.Struct({ @@ -38,7 +37,7 @@ const CancelCallSchema = Schema.Struct({ const SettleSchema = Schema.Struct({ kind: Schema.Literal('settle'), executionId: ExecutionIdSchema, - settlement: RunSettlementSchema, + settlement: SettlementSchema, }); export const RunOutputSchema = Schema.Union([ diff --git a/packages/workflow-engine/src/engine/portability.test.ts b/packages/workflow-engine/src/engine/portability.test.ts index 69b876e15..e979841e3 100644 --- a/packages/workflow-engine/src/engine/portability.test.ts +++ b/packages/workflow-engine/src/engine/portability.test.ts @@ -1,4 +1,5 @@ import { readdirSync, readFileSync } from 'node:fs'; +import { builtinModules } from 'node:module'; import { join } from 'node:path'; import { fileURLToPath } from 'node:url'; @@ -11,29 +12,90 @@ interface Forbidden { const source = fileURLToPath(new URL('..', import.meta.url)); +const bareBuiltins = builtinModules.filter((name) => !name.startsWith('_')).join('|'); + const nodeOnly: readonly Forbidden[] = [ { what: 'a node: module', pattern: /from ['"]node:/u }, + { what: 'a Node module by its bare name', pattern: new RegExp(`from ['"](?:${bareBuiltins})(?:/[^'"]*)?['"]`, 'u') }, + { what: 'a dynamic import', pattern: /\bimport\(/u }, { what: 'a Node global', pattern: /\b(?:process|Buffer|require|setImmediate|__dirname)\b/u }, { what: 'code generation', pattern: /\beval\(|new Function\(/u }, { what: 'Temporal', pattern: /@temporalio\//u }, ]; +const impure: readonly Forbidden[] = [ + { what: 'the clock', pattern: /\bDate\.now\(|\bnew Date\b|\bperformance\.now\(/u }, + { what: 'a random source', pattern: /\bMath\.random\(|\bcrypto\.getRandomValues\(|\brandomUUID\(/u }, + { what: 'the host locale or time zone', pattern: /\bIntl\./u }, +]; + +const growingCache: readonly Forbidden[] = [ + { + what: 'a module-level collection', + pattern: /^(?:export )?const \w+(?:: [^=]+)? = new (?:Map|Set|WeakMap|WeakSet)\b/mu, + }, + { what: 'a module-level variable', pattern: /^(?:export )?let /mu }, + { what: 'a module-level list', pattern: /^(?:export )?const \w+(?:: [^=]+)? = \[\];/mu }, +]; + function isProductionSource(file: string): boolean { - return file.endsWith('.ts') && !file.endsWith('.test.ts'); + return file.endsWith('.ts') && !file.endsWith('.test.ts') && !file.startsWith('testing'); +} + +function sourcesUnder(folder: string): readonly string[] { + return readdirSync(join(source, folder), { recursive: true, encoding: 'utf8' }) + .map((file) => join(folder, file)) + .filter((file) => isProductionSource(file)); +} + +function findingsIn(files: readonly string[], forbidden: readonly Forbidden[]): readonly string[] { + return files.flatMap((file) => { + const text = readFileSync(join(source, file), 'utf8'); + return forbidden + .filter(({ pattern }: Forbidden) => pattern.test(text)) + .map(({ what }: Forbidden) => `${file}: ${what}`); + }); } -function findingsIn(file: string): readonly string[] { - const text = readFileSync(join(source, file), 'utf8'); - return nodeOnly - .filter(({ pattern }: Forbidden) => pattern.test(text)) - .map(({ what }: Forbidden) => `${file}: ${what}`); +function caught(text: string, forbidden: readonly Forbidden[]): readonly string[] { + return forbidden.filter(({ pattern }: Forbidden) => pattern.test(text)).map(({ what }: Forbidden) => what); } +const everySource = sourcesUnder('.'); + +const machineAndRunLog = [...sourcesUnder('machine'), ...sourcesUnder('run-log')]; + describe('the engine core', () => { - it('uses no Node-only API, no code generation and no Temporal, so it runs in workerd as it runs in Node', () => { - const files = readdirSync(source, { recursive: true, encoding: 'utf8' }).filter((file) => isProductionSource(file)); + it('uses no Node-only API, no dynamic import, no code generation and no Temporal, so it runs in workerd as in Node', () => { + expect(everySource.length).toBeGreaterThan(20); + expect(findingsIn(everySource, nodeOnly)).toEqual([]); + }); + + it('keeps no module-level cache that could grow with the history of a run', () => { + expect(findingsIn(everySource, growingCache)).toEqual([]); + }); +}); + +describe('the machine and the run log', () => { + it('read no clock, no random source and no locale, so the same state and input decide the same events', () => { + expect(machineAndRunLog.length).toBeGreaterThan(10); + expect(findingsIn(machineAndRunLog, impure)).toEqual([]); + }); +}); - expect(files.length).toBeGreaterThan(10); - expect(files.flatMap((file) => findingsIn(file))).toEqual([]); +describe('the checks of purity', () => { + it('catch what they are there to catch', () => { + expect(caught("import { x } from 'fs';\nconst y = await import('./z.ts');", nodeOnly)).toEqual([ + 'a Node module by its bare name', + 'a dynamic import', + ]); + expect(caught('const t = Date.now();\nconst r = Math.random();\nIntl.DateTimeFormat();', impure)).toEqual([ + 'the clock', + 'a random source', + 'the host locale or time zone', + ]); + expect( + caught('const seen = new Map();\nlet count = 0;\nconst all: string[] = [];', growingCache), + ).toEqual(['a module-level collection', 'a module-level variable', 'a module-level list']); }); }); diff --git a/packages/workflow-engine/src/engine/workflow-engine.ts b/packages/workflow-engine/src/engine/workflow-engine.ts index 4b90f9318..2d574e63e 100644 --- a/packages/workflow-engine/src/engine/workflow-engine.ts +++ b/packages/workflow-engine/src/engine/workflow-engine.ts @@ -1,9 +1,11 @@ +import type { Conflict } from '@beonauto/operations'; import type { Effect } from 'effect'; import type { DispatchWatermark } from '../dispatch/dispatch-watermark.ts'; import type { Executor } from '../executor/executor.ts'; +import type { SubmissionOutcome } from '../machine/admission.ts'; import type { RunInput } from '../machine/run-input.ts'; -import type { RunStore, VersionConflict } from '../run-log/run-store.ts'; +import type { RunStore } from '../run-log/run-store.ts'; import type { RunSerialiser } from '../serialisation/run-serialiser.ts'; import type { RecordStore } from '../settlement/record-store.ts'; import type { Timers } from '../timers/timers.ts'; @@ -18,7 +20,7 @@ export interface EnginePorts { } export interface Submission { - readonly applied: boolean; + readonly outcome: SubmissionOutcome; readonly version: number; } @@ -27,7 +29,14 @@ export interface Wake { readonly dispatchedThrough: number; } +export interface SweepReport { + readonly runs: number; + readonly behind: number; + readonly timersArmedAgain: number; +} + export interface WorkflowEngine { - readonly submit: (input: RunInput) => Effect.Effect; + readonly submit: (input: RunInput) => Effect.Effect; readonly wake: (executionId: string) => Effect.Effect; + readonly sweep: () => Effect.Effect; } diff --git a/packages/workflow-engine/src/executor/call-result.ts b/packages/workflow-engine/src/executor/call-result.ts index 5e174735c..b9f9fbeb3 100644 --- a/packages/workflow-engine/src/executor/call-result.ts +++ b/packages/workflow-engine/src/executor/call-result.ts @@ -1,5 +1,7 @@ import { Schema } from 'effect'; +export const invalidArguments = 'invalid_arguments'; + export const CallResultSchema = Schema.Union([ Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), Schema.Struct({ status: Schema.Literal('rejected'), reason: Schema.String, detail: Schema.String }), @@ -8,3 +10,5 @@ export const CallResultSchema = Schema.Union([ ]); export type CallResult = typeof CallResultSchema.Type; + +export type CallStatus = CallResult['status']; diff --git a/packages/workflow-engine/src/executor/executor.ts b/packages/workflow-engine/src/executor/executor.ts index fc5c54323..7cccf5325 100644 --- a/packages/workflow-engine/src/executor/executor.ts +++ b/packages/workflow-engine/src/executor/executor.ts @@ -3,7 +3,11 @@ import type { Effect } from 'effect'; import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; import type { CancelCall, StartCall } from '../dispatch/run-output.ts'; +export type StartReceipt = 'started' | 'already_started' | 'refused_after_cancel'; + +export type CallCancelReceipt = 'cancelled' | 'already_answered' | 'tombstoned'; + export interface Executor { - readonly start: (call: StartCall, run: RunContext) => Effect.Effect; - readonly cancel: (call: CancelCall, run: RunContext) => Effect.Effect; + readonly start: (call: StartCall, run: RunContext) => Effect.Effect; + readonly cancel: (call: CancelCall, run: RunContext) => Effect.Effect; } diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts index 94439cca3..7e414095d 100644 --- a/packages/workflow-engine/src/index.ts +++ b/packages/workflow-engine/src/index.ts @@ -1,5 +1,6 @@ export { DispatchFailed, + dispatchedThrough, outputsAbove, type DispatchWatermark, type PositionedOutput, @@ -14,10 +15,10 @@ export { type Settle, type StartCall, } from './dispatch/run-output.ts'; -export type { EnginePorts, Submission, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; +export type { EnginePorts, Submission, SweepReport, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; -export { CallResultSchema, type CallResult } from './executor/call-result.ts'; -export type { Executor } from './executor/executor.ts'; +export { CallResultSchema, invalidArguments, type CallResult, type CallStatus } from './executor/call-result.ts'; +export type { CallCancelReceipt, Executor, StartReceipt } from './executor/executor.ts'; export { ReceivedEventSchema, mostEventIdLength, @@ -27,8 +28,27 @@ export { mostWaitingEvents, type ReceivedEvent, } from './inbox/received-event.ts'; -export { InputReceiptSchema, isStale, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; -export { mostEventBytes, mostHeldBytes, mostStepsWithoutWaiting, mostTasksPerInput } from './machine/limits.ts'; +export { + RunMismatch, + outcomeOf, + staleReasonOf, + type StaleReason, + type SubmissionOutcome, +} from './machine/admission.ts'; +export { InputReceiptSchema, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; +export { InstantSchema, clampedAt } from './machine/instant.ts'; +export { + mostCallArgumentsBytes, + mostEventBytes, + mostExpressionWork, + mostHeldBytes, + mostHistoryBytes, + mostInputs, + mostStepsWithoutWaiting, + mostTasksPerInput, + mostWorkPerInput, + taskFrameBytes, +} from './machine/limits.ts'; export type { RunDecider } from './machine/run-decider.ts'; export { RunInputSchema, @@ -46,8 +66,10 @@ export { newRun, type ArmedTimer, type Branch, + type CursorCurrent, type DslError, type FrameBody, + type HeldValue, type InboxState, type ListCursor, type MachineState, @@ -55,15 +77,25 @@ export { type RunState, type TaskFrame, type TryPhase, + type ValueId, type Variables, type WaitingEvent, } from './machine/run-state.ts'; -export { RunEventSchema, type PositionedEvent, type RunEvent } from './run-log/run-event.ts'; -export { VersionConflict, type AppendedEvent, type LoadedRun, type RunStore } from './run-log/run-store.ts'; +export { + RunEventSchema, + StepSchema, + eventBytesOf, + fitsInOneEvent, + type PositionedEvent, + type RunEvent, + type Step, +} from './run-log/run-event.ts'; +export { StreamGap, evolveRun, loadedRunOf, type LoadedRun } from './run-log/run-fold.ts'; +export type { AppendedEvent, RunStore, StoredRun, StoredSnapshot } from './run-log/run-store.ts'; export { SnapshotSchema, isSnapshotDue, - snapshotChunkLength, + mostSnapshotChunkBytes, snapshotChunks, snapshotEveryBytes, snapshotEveryInputs, @@ -71,9 +103,21 @@ export { type SinceSnapshot, type Snapshot, } from './run-log/snapshot.ts'; -export { PatchOperationSchema, type PatchOperation, type StatePatch } from './run-log/state-patch.ts'; +export { StateFormatSchema, stateFormat } from './run-log/state-format.ts'; +export { + PatchFailed, + PatchOperationSchema, + applyStatePatch, + type PatchOperation, + type StatePatch, +} from './run-log/state-patch.ts'; export type { RunSerialiser } from './serialisation/run-serialiser.ts'; -export type { RecordStore, SettleReceipt } from './settlement/record-store.ts'; -export { RunSettlementSchema, type RunSettlement } from './settlement/run-settlement.ts'; +export { + SettlementSchema, + type RecordStore, + type SettleReceipt, + type SettleRequest, + type Settlement, +} from './settlement/record-store.ts'; export { TimerPurposeSchema, timerIdOf, type TimerPurpose } from './timers/timer-id.ts'; -export type { Timers } from './timers/timers.ts'; +export type { ArmReceipt, TimerCancelReceipt, Timers } from './timers/timers.ts'; diff --git a/packages/workflow-engine/src/machine/admission.test.ts b/packages/workflow-engine/src/machine/admission.test.ts new file mode 100644 index 000000000..3bd9b8f42 --- /dev/null +++ b/packages/workflow-engine/src/machine/admission.test.ts @@ -0,0 +1,81 @@ +import { describe, expect, it } from 'vitest'; + +import { newRun, outcomeOf, RunMismatch, staleReasonOf, type RunInput } from '../index.ts'; +import { armedTimer, at, document, executionId, openCall, runningState, started } from '../testing/runs.ts'; + +const fired = (timerId: string): RunInput => ({ kind: 'timer_fired', executionId, at, timerId }); + +const answered = (run: number): RunInput => ({ + kind: 'call_answered', + executionId, + at, + key: { ...openCall, run }, + result: { status: 'succeeded', output: { approved: true } }, +}); + +const received = (id: string): RunInput => ({ + kind: 'event_received', + executionId, + at, + event: { id, type: 'com.acme.approval' }, +}); + +const cancelled: RunInput = { kind: 'cancel_requested', executionId, at }; + +describe('an input that can still change a run', () => { + it('is a start of a new run, the fire of an armed timer, the answer of an open call, a new event or a first cancel', () => { + const reasons = [ + staleReasonOf(newRun, started), + staleReasonOf(runningState, fired(armedTimer)), + staleReasonOf(runningState, answered(1)), + staleReasonOf(runningState, received('event-3')), + staleReasonOf(runningState, cancelled), + ]; + + expect(reasons).toEqual([undefined, undefined, undefined, undefined, undefined]); + expect(outcomeOf()).toBe('applied'); + }); +}); + +describe('an input to a run that has not started', () => { + it('is not stale but too early, so the adapter answers not_found and the caller tries again', () => { + expect( + [fired(armedTimer), answered(1), received('event-3'), cancelled].map((input) => staleReasonOf(newRun, input)), + ).toEqual(['not_started', 'not_started', 'not_started', 'not_started']); + expect(outcomeOf('not_started')).toBe('not_started'); + }); +}); + +describe('a stale input', () => { + it('names why it can no longer change the run', () => { + const reordered = { do: document.do, document: document.document }; + + expect([ + staleReasonOf(runningState, { ...started, document: reordered }), + staleReasonOf({ ...runningState, status: 'ended' }, answered(1)), + staleReasonOf(runningState, fired(`${executionId}/timers/9`)), + staleReasonOf(runningState, answered(2)), + staleReasonOf(runningState, received('event-1')), + staleReasonOf({ ...runningState, cancelRequested: true }, cancelled), + ]).toEqual([ + 'started_before', + 'run_ended', + 'timer_not_armed', + 'call_not_open', + 'event_received_before', + 'cancel_requested_before', + ]); + expect(outcomeOf('timer_not_armed')).toBe('stale'); + }); +}); + +describe('an input that belongs to no run like this one', () => { + it('dies instead of being taken as stale: another execution, or a start with another document', () => { + expect(() => staleReasonOf(runningState, { ...fired(armedTimer), executionId: 'another' })).toThrow(RunMismatch); + expect(() => staleReasonOf(runningState, { ...started, executionId: 'another' })).toThrow(RunMismatch); + expect(() => staleReasonOf(runningState, { ...started, document: { do: [{ other: {} }] } })).toThrow( + new RunMismatch({ detail: `The run of ${executionId} was started again with another document` }), + ); + expect(() => staleReasonOf({ ...runningState, workflow: null }, started)).toThrow(RunMismatch); + }); +}); diff --git a/packages/workflow-engine/src/machine/admission.ts b/packages/workflow-engine/src/machine/admission.ts new file mode 100644 index 000000000..2f4eb401a --- /dev/null +++ b/packages/workflow-engine/src/machine/admission.ts @@ -0,0 +1,70 @@ +import { Data } from 'effect'; + +import { callKeyText } from '../executor/call-key.ts'; +import type { RunInput } from './run-input.ts'; +import type { RunState } from './run-state.ts'; +import { sameJson } from './same-json.ts'; + +export type StaleReason = + | 'not_started' + | 'started_before' + | 'run_ended' + | 'timer_not_armed' + | 'call_not_open' + | 'event_received_before' + | 'cancel_requested_before'; + +export type SubmissionOutcome = 'applied' | 'stale' | 'not_started'; + +export class RunMismatch extends Data.TaggedError('run_mismatch')<{ readonly detail: string }> {} + +function requireSameRun(state: RunState, input: RunInput): void { + if (input.executionId !== state.executionId) { + throw new RunMismatch({ detail: `An input for ${input.executionId} reached the run of ${state.executionId}` }); + } +} + +function startedReason( + state: RunState, + input: Extract, +): StaleReason | undefined { + if (state.status === 'new') { + return undefined; + } + requireSameRun(state, input); + if (!sameJson(state.workflow?.document ?? null, input.document)) { + throw new RunMismatch({ detail: `The run of ${state.executionId} was started again with another document` }); + } + return 'started_before'; +} + +function spentReason(state: RunState, input: Exclude): StaleReason | undefined { + if (input.kind === 'timer_fired') { + return Object.hasOwn(state.timers.armed, input.timerId) ? undefined : 'timer_not_armed'; + } + if (input.kind === 'call_answered') { + return Object.hasOwn(state.calls, callKeyText(input.key)) ? undefined : 'call_not_open'; + } + if (input.kind === 'event_received') { + return state.inbox.receivedIds.includes(input.event.id) ? 'event_received_before' : undefined; + } + return state.cancelRequested ? 'cancel_requested_before' : undefined; +} + +export function staleReasonOf(state: RunState, input: RunInput): StaleReason | undefined { + if (input.kind === 'started') { + return startedReason(state, input); + } + if (state.status === 'new') { + return 'not_started'; + } + requireSameRun(state, input); + return state.status === 'ended' ? 'run_ended' : spentReason(state, input); +} + +export function outcomeOf(reason?: StaleReason): SubmissionOutcome { + if (reason === undefined) { + return 'applied'; + } + return reason === 'not_started' ? 'not_started' : 'stale'; +} diff --git a/packages/workflow-engine/src/machine/input-receipt.test.ts b/packages/workflow-engine/src/machine/input-receipt.test.ts index 44e428581..aace3bc8a 100644 --- a/packages/workflow-engine/src/machine/input-receipt.test.ts +++ b/packages/workflow-engine/src/machine/input-receipt.test.ts @@ -1,71 +1,29 @@ import { describe, expect, it } from 'vitest'; -import { callKeyText, isStale, newRun, receiptOf, type RunInput } from '../index.ts'; -import { armedTimer, at, executionId, openCall, runningState, started } from '../testing/runs.ts'; +import { callKeyText, clampedAt, receiptOf, type RunInput } from '../index.ts'; +import { armedTimer, at, executionId, openCall, started } from '../testing/runs.ts'; -const fired = (timerId: string): RunInput => ({ kind: 'timer_fired', executionId, at, timerId }); - -const answered = (run: number): RunInput => ({ - kind: 'call_answered', - executionId, - at, - key: { ...openCall, run }, - result: { status: 'succeeded', output: { approved: true } }, -}); - -const received = (id: string): RunInput => ({ - kind: 'event_received', - executionId, - at, - event: { id, type: 'com.acme.approval' }, -}); - -const cancelled: RunInput = { kind: 'cancel_requested', executionId, at }; +const inputs: readonly RunInput[] = [ + started, + { kind: 'timer_fired', executionId, at, timerId: armedTimer }, + { kind: 'call_answered', executionId, at, key: openCall, result: { status: 'failed', detail: 'boom' } }, + { kind: 'event_received', executionId, at, event: { id: 'event-9', type: 'com.acme.approval' } }, + { kind: 'cancel_requested', executionId, at }, +]; describe('the receipt of an input', () => { - it('names the input by the key it is deduplicated by', () => { - expect( - [started, fired(armedTimer), answered(1), received('event-9'), cancelled].map((input) => receiptOf(input)), - ).toEqual([ + it('names the input by the key it is deduplicated by, with the status of an answer or the type of an event', () => { + expect(inputs.map((input) => receiptOf(input, 0))).toEqual([ { kind: 'started', key: executionId, at }, { kind: 'timer_fired', key: armedTimer, at }, - { kind: 'call_answered', key: callKeyText(openCall), at }, - { kind: 'event_received', key: 'event-9', at }, + { kind: 'call_answered', key: callKeyText(openCall), at, status: 'failed' }, + { kind: 'event_received', key: 'event-9', at, eventType: 'com.acme.approval' }, { kind: 'cancel_requested', key: executionId, at }, ]); }); -}); - -describe('an input that can still change a run', () => { - it('is a start of a new run, the fire of an armed timer, the answer of an open call, a new event or the first cancel', () => { - expect( - [ - isStale(newRun, started), - isStale(runningState, fired(armedTimer)), - isStale(runningState, answered(1)), - isStale(runningState, received('event-3')), - isStale(runningState, cancelled), - ].every((stale) => !stale), - ).toBe(true); - }); -}); - -describe('a stale input', () => { - it('is a second start, a fire of a timer not armed, an answer of a call not open, an event seen before or a second cancel', () => { - expect([ - isStale(runningState, started), - isStale(runningState, fired(`${executionId}/timers/9`)), - isStale(runningState, answered(2)), - isStale(runningState, received('event-1')), - isStale({ ...runningState, cancelRequested: true }, cancelled), - ]).toEqual([true, true, true, true, true]); - }); - it('is anything but a start for a run that has not started, has ended, or is another run', () => { - expect([ - isStale(newRun, fired(armedTimer)), - isStale({ ...runningState, status: 'ended' }, answered(1)), - isStale({ ...runningState, executionId: 'another' }, received('event-3')), - ]).toEqual([true, true, true]); + it('never goes back in time: an input whose clock is behind the last input takes the time of that input', () => { + expect(receiptOf(started, at + 5000).at).toBe(at + 5000); + expect([clampedAt(at, at - 1), clampedAt(at, at + 1)]).toEqual([at, at + 1]); }); }); diff --git a/packages/workflow-engine/src/machine/input-receipt.ts b/packages/workflow-engine/src/machine/input-receipt.ts index 9a1b1b359..11860055c 100644 --- a/packages/workflow-engine/src/machine/input-receipt.ts +++ b/packages/workflow-engine/src/machine/input-receipt.ts @@ -1,47 +1,38 @@ import { Schema } from 'effect'; import { callKeyText } from '../executor/call-key.ts'; +import { clampedAt, InstantSchema } from './instant.ts'; import type { RunInput } from './run-input.ts'; -import type { RunState } from './run-state.ts'; -export const InputReceiptSchema = Schema.Struct({ - kind: Schema.Literals(['started', 'timer_fired', 'call_answered', 'event_received', 'cancel_requested']), +const keyed = { key: Schema.String, - at: Schema.Int, -}); + at: InstantSchema, +}; -export type InputReceipt = typeof InputReceiptSchema.Type; - -function keyOf(input: RunInput): string { - if (input.kind === 'timer_fired') { - return input.timerId; - } - if (input.kind === 'call_answered') { - return callKeyText(input.key); - } - return input.kind === 'event_received' ? input.event.id : input.executionId; -} +export const InputReceiptSchema = Schema.Union([ + Schema.Struct({ kind: Schema.Literal('started'), ...keyed }), + Schema.Struct({ kind: Schema.Literal('timer_fired'), ...keyed }), + Schema.Struct({ + kind: Schema.Literal('call_answered'), + ...keyed, + status: Schema.Literals(['succeeded', 'rejected', 'failed', 'unreachable']), + }), + Schema.Struct({ kind: Schema.Literal('event_received'), ...keyed, eventType: Schema.String }), + Schema.Struct({ kind: Schema.Literal('cancel_requested'), ...keyed }), +]); -export function receiptOf(input: RunInput): InputReceipt { - return { kind: input.kind, key: keyOf(input), at: input.at }; -} +export type InputReceipt = typeof InputReceiptSchema.Type; -function isSpent(state: RunState, input: RunInput): boolean { - if (input.kind === 'started') { - return state.status !== 'new'; - } +export function receiptOf(input: RunInput, lastInputAt: number): InputReceipt { + const at = clampedAt(lastInputAt, input.at); if (input.kind === 'timer_fired') { - return !Object.hasOwn(state.timers.armed, input.timerId); + return { kind: input.kind, key: input.timerId, at }; } if (input.kind === 'call_answered') { - return !Object.hasOwn(state.calls.open, callKeyText(input.key)); + return { kind: input.kind, key: callKeyText(input.key), at, status: input.result.status }; } - return input.kind === 'event_received' ? state.inbox.receivedIds.includes(input.event.id) : state.cancelRequested; -} - -export function isStale(state: RunState, input: RunInput): boolean { - if (input.kind !== 'started' && (state.status !== 'running' || input.executionId !== state.executionId)) { - return true; + if (input.kind === 'event_received') { + return { kind: input.kind, key: input.event.id, at, eventType: input.event.type }; } - return isSpent(state, input); + return { kind: input.kind, key: input.executionId, at }; } diff --git a/packages/workflow-engine/src/machine/instant.ts b/packages/workflow-engine/src/machine/instant.ts new file mode 100644 index 000000000..eaa887fca --- /dev/null +++ b/packages/workflow-engine/src/machine/instant.ts @@ -0,0 +1,7 @@ +import { Schema } from 'effect'; + +export const InstantSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); + +export function clampedAt(lastInputAt: number, at: number): number { + return Math.max(lastInputAt, at); +} diff --git a/packages/workflow-engine/src/machine/limits.ts b/packages/workflow-engine/src/machine/limits.ts index 64b2f9827..1d22bf1f6 100644 --- a/packages/workflow-engine/src/machine/limits.ts +++ b/packages/workflow-engine/src/machine/limits.ts @@ -1,7 +1,19 @@ export const mostHeldBytes = 4_194_304; +export const taskFrameBytes = 4096; + export const mostEventBytes = 1_572_864; +export const mostCallArgumentsBytes = 270_336; + export const mostTasksPerInput = 100; export const mostStepsWithoutWaiting = 10_000; + +export const mostExpressionWork = 8_000_000; + +export const mostWorkPerInput = 16_000_000; + +export const mostInputs = 100_000; + +export const mostHistoryBytes = 536_870_912; diff --git a/packages/workflow-engine/src/machine/run-input.ts b/packages/workflow-engine/src/machine/run-input.ts index c6a03c571..9583f6b81 100644 --- a/packages/workflow-engine/src/machine/run-input.ts +++ b/packages/workflow-engine/src/machine/run-input.ts @@ -3,11 +3,10 @@ import { Schema } from 'effect'; import { CallKeySchema } from '../executor/call-key.ts'; import { CallResultSchema } from '../executor/call-result.ts'; import { ReceivedEventSchema } from '../inbox/received-event.ts'; +import { InstantSchema } from './instant.ts'; const ExecutionIdSchema = Schema.NonEmptyString; -const InstantSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); - const PositiveMillisecondsSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)); export const RunLimitsSchema = Schema.Struct({ diff --git a/packages/workflow-engine/src/machine/run-state.ts b/packages/workflow-engine/src/machine/run-state.ts index 670de117a..fc6da503a 100644 --- a/packages/workflow-engine/src/machine/run-state.ts +++ b/packages/workflow-engine/src/machine/run-state.ts @@ -3,6 +3,7 @@ import { Schema } from 'effect'; import { CallKeySchema, type CallKey } from '../executor/call-key.ts'; import { ReceivedEventSchema, type ReceivedEvent } from '../inbox/received-event.ts'; import { TimerPurposeSchema, type TimerPurpose } from '../timers/timer-id.ts'; +import { InstantSchema } from './instant.ts'; import { RunLimitsSchema, type RunLimits } from './run-input.ts'; export interface DslError { @@ -13,19 +14,31 @@ export interface DslError { readonly detail?: string; } -export type Variables = Readonly>; +export type ValueId = number; + +export interface HeldValue { + readonly value: Schema.Json; + readonly bytes: number; + readonly holders: number; +} + +export type Variables = Readonly>; + +export type CursorCurrent = + | { readonly kind: 'running'; readonly task: TaskFrame } + | { readonly kind: 'yielding'; readonly timer: string }; export interface ListCursor { readonly pointer: string; readonly position: number; - readonly data: Schema.Json; + readonly data: ValueId; readonly variables: Variables; - readonly current: TaskFrame | null; + readonly current: CursorCurrent | null; } export type Branch = | { readonly state: 'running'; readonly task: TaskFrame } - | { readonly state: 'finished'; readonly output: Schema.Json; readonly flow: string } + | { readonly state: 'finished'; readonly output: ValueId; readonly flow: string } | { readonly state: 'failed'; readonly error: DslError }; export type TryPhase = @@ -37,30 +50,36 @@ export type FrameBody = | { readonly kind: 'list'; readonly list: ListCursor } | { readonly kind: 'for'; - readonly items: readonly Schema.Json[]; + readonly items: ValueId; readonly index: number; - readonly data: Schema.Json; + readonly data: ValueId; readonly list: ListCursor | null; } | { readonly kind: 'fork'; readonly compete: boolean; readonly branches: readonly Branch[] } | { readonly kind: 'try'; readonly attempt: number; readonly startedAt: number; readonly phase: TryPhase } | { readonly kind: 'wait'; readonly timer: string } - | { readonly kind: 'call'; readonly key: CallKey } - | { readonly kind: 'listen'; readonly consumed: readonly Schema.JsonObject[] } - | { readonly kind: 'yield'; readonly timer: string }; + | { + readonly kind: 'call'; + readonly key: CallKey; + readonly primitive: string | null; + readonly name: string | null; + } + | { readonly kind: 'listen'; readonly consumed: readonly ValueId[] }; export interface TaskFrame { readonly reference: string; readonly run: number; - readonly rawInput: Schema.Json; - readonly input: Schema.Json; + readonly rawInput: ValueId; + readonly input: ValueId; readonly variables: Variables; readonly timeout: string | null; readonly body: FrameBody; } export interface MachineState { - readonly context: Schema.Json; + readonly values: Readonly>; + readonly nextValue: ValueId; + readonly context: ValueId; readonly root: TaskFrame | null; } @@ -75,6 +94,7 @@ export type RunOutcome = export interface ArmedTimer { readonly purpose: TimerPurpose; readonly reference: string; + readonly dueAt: number; } export interface WaitingEvent { @@ -94,17 +114,16 @@ export interface InboxState { export interface RunState { readonly executionId: string; readonly status: 'new' | 'running' | 'ended'; - readonly workflow: { readonly document: Schema.JsonObject; readonly input: Schema.Json } | null; + readonly workflow: { readonly document: Schema.JsonObject; readonly input: ValueId } | null; readonly attributes: Schema.JsonObject; readonly limits: RunLimits; readonly startedAt: number; readonly lastInputAt: number; + readonly inputs: number; readonly random: { readonly seed: number; readonly draws: number }; + readonly runs: Readonly>; readonly timers: { readonly next: number; readonly armed: Readonly> }; - readonly calls: { - readonly runs: Readonly>; - readonly open: Readonly>; - }; + readonly calls: Readonly>; readonly inbox: InboxState; readonly heldBytes: number; readonly stepsWithoutWaiting: number; @@ -115,7 +134,9 @@ export interface RunState { const IntSchema = Schema.Int; -const VariablesSchema = Schema.Record(Schema.String, Schema.Json); +const ValueIdSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); + +const VariablesSchema = Schema.Record(Schema.String, ValueIdSchema); const DslErrorSchema = Schema.Struct({ type: Schema.String, @@ -125,20 +146,24 @@ const DslErrorSchema = Schema.Struct({ detail: Schema.optionalKey(Schema.String), }); +const TaskFrameReference = Schema.suspend((): Schema.Codec => TaskFrameSchema); + +const CursorCurrentSchema: Schema.Codec = Schema.Union([ + Schema.Struct({ kind: Schema.Literal('running'), task: TaskFrameReference }), + Schema.Struct({ kind: Schema.Literal('yielding'), timer: Schema.String }), +]); + const ListCursorSchema: Schema.Codec = Schema.Struct({ pointer: Schema.String, position: IntSchema, - data: Schema.Json, + data: ValueIdSchema, variables: VariablesSchema, - current: Schema.NullOr(Schema.suspend((): Schema.Codec => TaskFrameSchema)), + current: Schema.NullOr(CursorCurrentSchema), }); const BranchSchema: Schema.Codec = Schema.Union([ - Schema.Struct({ - state: Schema.Literal('running'), - task: Schema.suspend((): Schema.Codec => TaskFrameSchema), - }), - Schema.Struct({ state: Schema.Literal('finished'), output: Schema.Json, flow: Schema.String }), + Schema.Struct({ state: Schema.Literal('running'), task: TaskFrameReference }), + Schema.Struct({ state: Schema.Literal('finished'), output: ValueIdSchema, flow: Schema.String }), Schema.Struct({ state: Schema.Literal('failed'), error: DslErrorSchema }), ]); @@ -152,24 +177,28 @@ const FrameBodySchema: Schema.Codec = Schema.Union([ Schema.Struct({ kind: Schema.Literal('list'), list: ListCursorSchema }), Schema.Struct({ kind: Schema.Literal('for'), - items: Schema.Array(Schema.Json), + items: ValueIdSchema, index: IntSchema, - data: Schema.Json, + data: ValueIdSchema, list: Schema.NullOr(ListCursorSchema), }), Schema.Struct({ kind: Schema.Literal('fork'), compete: Schema.Boolean, branches: Schema.Array(BranchSchema) }), - Schema.Struct({ kind: Schema.Literal('try'), attempt: IntSchema, startedAt: IntSchema, phase: TryPhaseSchema }), + Schema.Struct({ kind: Schema.Literal('try'), attempt: IntSchema, startedAt: InstantSchema, phase: TryPhaseSchema }), Schema.Struct({ kind: Schema.Literal('wait'), timer: Schema.String }), - Schema.Struct({ kind: Schema.Literal('call'), key: CallKeySchema }), - Schema.Struct({ kind: Schema.Literal('listen'), consumed: Schema.Array(Schema.JsonObject) }), - Schema.Struct({ kind: Schema.Literal('yield'), timer: Schema.String }), + Schema.Struct({ + kind: Schema.Literal('call'), + key: CallKeySchema, + primitive: Schema.NullOr(Schema.String), + name: Schema.NullOr(Schema.String), + }), + Schema.Struct({ kind: Schema.Literal('listen'), consumed: Schema.Array(ValueIdSchema) }), ]); const TaskFrameSchema: Schema.Codec = Schema.Struct({ reference: Schema.String, run: IntSchema, - rawInput: Schema.Json, - input: Schema.Json, + rawInput: ValueIdSchema, + input: ValueIdSchema, variables: VariablesSchema, timeout: Schema.NullOr(Schema.String), body: FrameBodySchema, @@ -187,20 +216,22 @@ const RunOutcomeSchema: Schema.Codec = Schema.Union([ export const RunStateSchema: Schema.Codec = Schema.Struct({ executionId: Schema.String, status: Schema.Literals(['new', 'running', 'ended']), - workflow: Schema.NullOr(Schema.Struct({ document: Schema.JsonObject, input: Schema.Json })), + workflow: Schema.NullOr(Schema.Struct({ document: Schema.JsonObject, input: ValueIdSchema })), attributes: Schema.JsonObject, limits: RunLimitsSchema, - startedAt: IntSchema, - lastInputAt: IntSchema, + startedAt: InstantSchema, + lastInputAt: InstantSchema, + inputs: IntSchema, random: Schema.Struct({ seed: IntSchema, draws: IntSchema }), + runs: Schema.Record(Schema.String, IntSchema), timers: Schema.Struct({ next: IntSchema, - armed: Schema.Record(Schema.String, Schema.Struct({ purpose: TimerPurposeSchema, reference: Schema.String })), - }), - calls: Schema.Struct({ - runs: Schema.Record(Schema.String, IntSchema), - open: Schema.Record(Schema.String, CallKeySchema), + armed: Schema.Record( + Schema.String, + Schema.Struct({ purpose: TimerPurposeSchema, reference: Schema.String, dueAt: InstantSchema }), + ), }), + calls: Schema.Record(Schema.String, CallKeySchema), inbox: Schema.Struct({ waiting: Schema.Array(Schema.Struct({ event: ReceivedEventSchema, bytes: IntSchema })), waitingBytes: IntSchema, @@ -212,7 +243,12 @@ export const RunStateSchema: Schema.Codec = Schema.Struct({ heldBytes: IntSchema, stepsWithoutWaiting: IntSchema, cancelRequested: Schema.Boolean, - machine: Schema.Struct({ context: Schema.Json, root: Schema.NullOr(TaskFrameSchema) }), + machine: Schema.Struct({ + values: Schema.Record(Schema.String, Schema.Struct({ value: Schema.Json, bytes: IntSchema, holders: IntSchema })), + nextValue: ValueIdSchema, + context: ValueIdSchema, + root: Schema.NullOr(TaskFrameSchema), + }), outcome: Schema.NullOr(RunOutcomeSchema), }); @@ -224,13 +260,15 @@ export const newRun: RunState = { limits: { mostDurationMs: 1, longestCallMs: 1 }, startedAt: 0, lastInputAt: 0, + inputs: 0, random: { seed: 0, draws: 0 }, + runs: {}, timers: { next: 1, armed: {} }, - calls: { runs: {}, open: {} }, + calls: {}, inbox: { waiting: [], waitingBytes: 0, receivedIds: [], received: 0, receivedBytes: 0, overflow: null }, heldBytes: 0, stepsWithoutWaiting: 0, cancelRequested: false, - machine: { context: {}, root: null }, + machine: { values: { 0: { value: {}, bytes: 2, holders: 1 } }, nextValue: 1, context: 0, root: null }, outcome: null, }; diff --git a/packages/workflow-engine/src/machine/same-json.ts b/packages/workflow-engine/src/machine/same-json.ts new file mode 100644 index 000000000..75f9d0c7d --- /dev/null +++ b/packages/workflow-engine/src/machine/same-json.ts @@ -0,0 +1,25 @@ +import type { Schema } from 'effect'; + +function isList(value: Schema.Json): value is Schema.JsonArray { + return Array.isArray(value); +} + +function canonical(value: Schema.Json): Schema.Json { + if (typeof value !== 'object' || value === null) { + return value; + } + if (isList(value)) { + return value.map((item: Schema.Json) => canonical(item)); + } + return Object.fromEntries( + Object.entries(value) + .toSorted(([first]: readonly [string, Schema.Json], [second]: readonly [string, Schema.Json]) => + first < second ? -1 : 1, + ) + .map(([key, item]: readonly [string, Schema.Json]) => [key, canonical(item)]), + ); +} + +export function sameJson(first: Schema.Json, second: Schema.Json): boolean { + return JSON.stringify(canonical(first)) === JSON.stringify(canonical(second)); +} diff --git a/packages/workflow-engine/src/run-log/run-event.test.ts b/packages/workflow-engine/src/run-log/run-event.test.ts index 65ebd67e4..85ceb12a8 100644 --- a/packages/workflow-engine/src/run-log/run-event.test.ts +++ b/packages/workflow-engine/src/run-log/run-event.test.ts @@ -1,7 +1,18 @@ import { Result, Schema } from 'effect'; import { describe, expect, it } from 'vitest'; -import { newRun, RunEventSchema, RunInputSchema, RunStateSchema, type RunEvent, type RunInput } from '../index.ts'; +import { + eventBytesOf, + fitsInOneEvent, + mostEventBytes, + newRun, + RunEventSchema, + RunInputSchema, + RunStateSchema, + stateFormat, + type RunEvent, + type RunInput, +} from '../index.ts'; import { at, executionId, openCall, runningState, started } from '../testing/runs.ts'; function asStored>(schema: S, value: S['Type']): unknown { @@ -17,11 +28,16 @@ function readBack>( const event: RunEvent = { type: 'input_applied', - receipt: { kind: 'call_answered', key: 'k', at }, + format: stateFormat, + receipt: { kind: 'call_answered', key: 'k', at, status: 'succeeded' }, + steps: [ + { reference: openCall.reference, run: 1, outcome: 'completed' }, + { reference: '/do/2', run: 1, outcome: 'waiting' }, + ], patch: [ - { op: 'replace', path: '/machine/context', value: { approved: true } }, - { op: 'add', path: '/timers/armed/e~1timers~13', value: { purpose: 'wait', reference: '/do/2' } }, - { op: 'remove', path: '/calls/open/k' }, + { op: 'replace', path: '/machine/context', value: 3 }, + { op: 'add', path: '/timers/armed/e~1timers~13', value: { purpose: 'wait', reference: '/do/2', dueAt: at + 1 } }, + { op: 'remove', path: '/calls/k' }, ], outputs: [ { kind: 'cancel_timer', executionId, timerId: `${executionId}/timers/2` }, @@ -31,6 +47,10 @@ const event: RunEvent = { ], }; +function near(text: string): RunEvent { + return { ...event, patch: [{ op: 'replace', path: '/machine/context', value: text }] }; +} + const inputs: readonly RunInput[] = [ started, { kind: 'timer_fired', executionId, at, timerId: `${executionId}/timers/1` }, @@ -39,21 +59,30 @@ const inputs: readonly RunInput[] = [ executionId, at, key: openCall, - result: { status: 'rejected', reason: 'conflict', detail: 'x' }, + result: { status: 'rejected', reason: 'invalid_arguments', detail: 'x' }, }, { kind: 'event_received', executionId, at, event: { id: 'event-3', type: 'com.acme.approval', data: [1, 2] } }, { kind: 'cancel_requested', executionId, at }, ]; describe('a run event', () => { - it('is stored as JSON and read back as it was decided', () => { + it('is stored as JSON, naming its state format, and read back as it was decided', () => { expect(readBack(RunEventSchema, asStored(RunEventSchema, event))).toEqual(Result.succeed(event)); }); - it('is refused when it holds an output the engine does not dispatch', () => { - const stored = { ...event, outputs: [{ kind: 'send_email', to: 'someone' }] }; + it('is refused when it names another state format, or holds an output the engine does not dispatch', () => { + expect([ + Result.isFailure(readBack(RunEventSchema, { ...event, format: stateFormat + 1 })), + Result.isFailure(readBack(RunEventSchema, { ...event, outputs: [{ kind: 'send_email', to: 'someone' }] })), + ]).toEqual([true, true]); + }); + + it('is measured in bytes of UTF-8 JSON, and fits when it takes at most 1.5 MiB', () => { + const overhead = eventBytesOf(near('')); - expect(Result.isFailure(readBack(RunEventSchema, stored))).toBe(true); + expect(eventBytesOf(near('é'))).toBe(overhead + 2); + expect(fitsInOneEvent(near('x'.repeat(mostEventBytes - overhead)))).toBe(true); + expect(fitsInOneEvent(near('x'.repeat(mostEventBytes - overhead + 1)))).toBe(false); }); }); diff --git a/packages/workflow-engine/src/run-log/run-event.ts b/packages/workflow-engine/src/run-log/run-event.ts index 1fc428c96..fd8a6ac76 100644 --- a/packages/workflow-engine/src/run-log/run-event.ts +++ b/packages/workflow-engine/src/run-log/run-event.ts @@ -2,18 +2,41 @@ import { Schema } from 'effect'; import { RunOutputSchema } from '../dispatch/run-output.ts'; import { InputReceiptSchema } from '../machine/input-receipt.ts'; +import { mostEventBytes } from '../machine/limits.ts'; +import { StateFormatSchema } from './state-format.ts'; import { PatchOperationSchema } from './state-patch.ts'; +export const StepSchema = Schema.Struct({ + reference: Schema.String, + run: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), + outcome: Schema.Literals(['started', 'skipped', 'waiting', 'completed', 'raised', 'timed_out', 'cancelled']), +}); + export const RunEventSchema = Schema.Struct({ type: Schema.Literal('input_applied'), + format: StateFormatSchema, receipt: InputReceiptSchema, + steps: Schema.Array(StepSchema), patch: Schema.Array(PatchOperationSchema), outputs: Schema.Array(RunOutputSchema), }); +export type Step = typeof StepSchema.Type; + export type RunEvent = typeof RunEventSchema.Type; export interface PositionedEvent { readonly version: number; + readonly bytes: number; readonly event: RunEvent; } + +const utf8 = new TextEncoder(); + +export function eventBytesOf(event: RunEvent): number { + return utf8.encode(JSON.stringify(event)).byteLength; +} + +export function fitsInOneEvent(event: RunEvent): boolean { + return eventBytesOf(event) <= mostEventBytes; +} diff --git a/packages/workflow-engine/src/run-log/run-fold.test.ts b/packages/workflow-engine/src/run-log/run-fold.test.ts new file mode 100644 index 000000000..3a915e10b --- /dev/null +++ b/packages/workflow-engine/src/run-log/run-fold.test.ts @@ -0,0 +1,97 @@ +import { describe, expect, it } from 'vitest'; + +import { + evolveRun, + loadedRunOf, + newRun, + stateFormat, + StreamGap, + type PositionedEvent, + type RunEvent, + type Snapshot, + type StatePatch, +} from '../index.ts'; +import { at, document, executionId } from '../testing/runs.ts'; + +const timer = `${executionId}/timers/1`; + +function applied(version: number, bytes: number, patch: StatePatch): PositionedEvent { + const event: RunEvent = { + type: 'input_applied', + format: stateFormat, + receipt: { kind: 'timer_fired', key: timer, at }, + steps: [], + patch, + outputs: [], + }; + return { version, bytes, event }; +} + +const startedRun = applied(1, 900, [ + { op: 'replace', path: '/executionId', value: executionId }, + { op: 'replace', path: '/status', value: 'running' }, + { op: 'replace', path: '/workflow', value: { document, input: 1 } }, + { op: 'add', path: '/machine/values/1', value: { value: { ticket: 7 }, bytes: 12, holders: 1 } }, + { op: 'replace', path: '/inputs', value: 1 }, + { op: 'replace', path: '/startedAt', value: at }, + { op: 'replace', path: '/lastInputAt', value: at }, +]); + +const armed = applied(2, 300, [ + { + op: 'add', + path: `/timers/armed/${timer.replaceAll('/', '~1')}`, + value: { purpose: 'wait', reference: '/do/0', dueAt: at + 1000 }, + }, + { op: 'replace', path: '/timers/next', value: 2 }, + { op: 'replace', path: '/inputs', value: 2 }, +]); + +const fired = applied(3, 200, [ + { op: 'remove', path: `/timers/armed/${timer.replaceAll('/', '~1')}` }, + { op: 'replace', path: '/machine/context', value: 1 }, + { op: 'replace', path: '/inputs', value: 3 }, + { op: 'replace', path: '/lastInputAt', value: at + 1000 }, +]); + +describe('a run loaded from its stream', () => { + it('is the fold of its events from a new run, with its version and the bytes of its history', () => { + const loaded = loadedRunOf({ snapshot: null, tail: [startedRun, armed, fired] }); + + expect(loaded).toMatchObject({ + version: 3, + historyBytes: 1400, + sinceSnapshot: { inputs: 3, bytes: 1400, snapshotBytes: 0 }, + state: { executionId, status: 'running', inputs: 3, lastInputAt: at + 1000, timers: { next: 2, armed: {} } }, + }); + expect(loadedRunOf({ snapshot: null, tail: [] }).state).toEqual(newRun); + }); + + it('is the same from its latest snapshot and the events after it as from its whole stream', () => { + const atTwo = loadedRunOf({ snapshot: null, tail: [startedRun, armed] }); + const snapshot: Snapshot = { + format: stateFormat, + executionId, + version: 2, + historyBytes: atTwo.historyBytes, + state: atTwo.state, + }; + + const fromSnapshot = loadedRunOf({ snapshot: { snapshot, bytes: 5000 }, tail: [fired] }); + + expect(fromSnapshot).toEqual({ + ...loadedRunOf({ snapshot: null, tail: [startedRun, armed, fired] }), + sinceSnapshot: { inputs: 1, bytes: 200, snapshotBytes: 5000 }, + }); + expect(evolveRun(atTwo.state, fired.event)).toEqual(fromSnapshot.state); + }); + + it('dies on a gap in its events, and on a patch that leaves a state this format does not describe', () => { + expect(() => loadedRunOf({ snapshot: null, tail: [startedRun, fired] })).toThrow( + new StreamGap({ expected: 2, found: 3 }), + ); + expect(() => + evolveRun(newRun, applied(1, 10, [{ op: 'replace', path: '/status', value: 'paused' }]).event), + ).toThrow(/Expected "new" \| "running" \| "ended"/u); + }); +}); diff --git a/packages/workflow-engine/src/run-log/run-fold.ts b/packages/workflow-engine/src/run-log/run-fold.ts new file mode 100644 index 000000000..c1cb6efdd --- /dev/null +++ b/packages/workflow-engine/src/run-log/run-fold.ts @@ -0,0 +1,67 @@ +import { Data, Schema } from 'effect'; + +import { newRun, RunStateSchema, type RunState } from '../machine/run-state.ts'; +import type { PositionedEvent, RunEvent } from './run-event.ts'; +import type { StoredRun } from './run-store.ts'; +import type { SinceSnapshot } from './snapshot.ts'; +import { applyStatePatch } from './state-patch.ts'; + +export interface LoadedRun { + readonly state: RunState; + readonly version: number; + readonly historyBytes: number; + readonly sinceSnapshot: SinceSnapshot; +} + +export class StreamGap extends Data.TaggedError('stream_gap')<{ + readonly expected: number; + readonly found: number; +}> {} + +interface FoldStart { + readonly state: RunState; + readonly version: number; + readonly historyBytes: number; + readonly snapshotBytes: number; +} + +const decodeState = Schema.decodeUnknownSync(RunStateSchema); + +export function evolveRun(state: RunState, event: RunEvent): RunState { + return decodeState(applyStatePatch(state, event.patch)); +} + +function startOf({ snapshot }: StoredRun): FoldStart { + return snapshot === null + ? { state: newRun, version: 0, historyBytes: 0, snapshotBytes: 0 } + : { + state: snapshot.snapshot.state, + version: snapshot.snapshot.version, + historyBytes: snapshot.snapshot.historyBytes, + snapshotBytes: snapshot.bytes, + }; +} + +function requireContiguous(tail: readonly PositionedEvent[], after: number): void { + for (const [index, { version }] of tail.entries()) { + if (version !== after + index + 1) { + throw new StreamGap({ expected: after + index + 1, found: version }); + } + } +} + +export function loadedRunOf(stored: StoredRun): LoadedRun { + const start = startOf(stored); + requireContiguous(stored.tail, start.version); + const folded = stored.tail.reduce( + (document: unknown, { event }: PositionedEvent) => applyStatePatch(document, event.patch), + start.state, + ); + const tailBytes = stored.tail.reduce((sum, { bytes }: PositionedEvent) => sum + bytes, 0); + return { + state: decodeState(folded), + version: start.version + stored.tail.length, + historyBytes: start.historyBytes + tailBytes, + sinceSnapshot: { inputs: stored.tail.length, bytes: tailBytes, snapshotBytes: start.snapshotBytes }, + }; +} diff --git a/packages/workflow-engine/src/run-log/run-store.ts b/packages/workflow-engine/src/run-log/run-store.ts index 188c52ad9..c38cdbf14 100644 --- a/packages/workflow-engine/src/run-log/run-store.ts +++ b/packages/workflow-engine/src/run-log/run-store.ts @@ -1,13 +1,17 @@ -import { Data, type Effect } from 'effect'; +import type { VersionConflict } from '@beonauto/ledger'; +import type { Effect } from 'effect'; -import type { RunState } from '../machine/run-state.ts'; import type { PositionedEvent, RunEvent } from './run-event.ts'; -import type { SinceSnapshot, Snapshot } from './snapshot.ts'; +import type { Snapshot } from './snapshot.ts'; -export interface LoadedRun { - readonly state: RunState; - readonly version: number; - readonly sinceSnapshot: SinceSnapshot; +export interface StoredSnapshot { + readonly snapshot: Snapshot; + readonly bytes: number; +} + +export interface StoredRun { + readonly snapshot: StoredSnapshot | null; + readonly tail: readonly PositionedEvent[]; } export interface AppendedEvent { @@ -15,13 +19,8 @@ export interface AppendedEvent { readonly bytes: number; } -export class VersionConflict extends Data.TaggedError('version_conflict')<{ - readonly executionId: string; - readonly expectedVersion: number; -}> {} - export interface RunStore { - readonly load: (executionId: string) => Effect.Effect; + readonly load: (executionId: string) => Effect.Effect; readonly append: ( executionId: string, event: RunEvent, diff --git a/packages/workflow-engine/src/run-log/snapshot.test.ts b/packages/workflow-engine/src/run-log/snapshot.test.ts index ff13befff..bb0eabfd3 100644 --- a/packages/workflow-engine/src/run-log/snapshot.test.ts +++ b/packages/workflow-engine/src/run-log/snapshot.test.ts @@ -1,46 +1,76 @@ import { Result } from 'effect'; import { describe, expect, it } from 'vitest'; -import { isSnapshotDue, snapshotChunkLength, snapshotChunks, snapshotFromChunks, type Snapshot } from '../index.ts'; +import { + isSnapshotDue, + mostSnapshotChunkBytes, + snapshotChunks, + snapshotFromChunks, + stateFormat, + type Snapshot, +} from '../index.ts'; import { executionId, runningState } from '../testing/runs.ts'; const highSurrogate = /[\uD800-\uDBFF]$/u; const lowSurrogate = /^[\uDC00-\uDFFF]/u; +const utf8 = new TextEncoder(); + function snapshotHolding(text: string): Snapshot { - return { format: 1, executionId, version: 4000, state: { ...runningState, machine: { context: text, root: null } } }; + const values = { ...runningState.machine.values, 0: { value: text, bytes: text.length, holders: 1 } }; + return { + format: stateFormat, + executionId, + version: 4000, + historyBytes: 9_000_000, + state: { ...runningState, machine: { ...runningState.machine, values } }, + }; } describe('a snapshot', () => { - it('is due after 1,000 inputs or 1 MiB of history since the last one, whichever comes first', () => { + it('is due after 1,000 inputs, or after as many bytes of events as the last snapshot took, and at least 1 MiB', () => { expect([ - isSnapshotDue({ inputs: 999, bytes: 1_048_575 }), - isSnapshotDue({ inputs: 1000, bytes: 0 }), - isSnapshotDue({ inputs: 1, bytes: 1_048_576 }), - ]).toEqual([false, true, true]); + isSnapshotDue({ inputs: 999, bytes: 1_048_575, snapshotBytes: 0 }), + isSnapshotDue({ inputs: 1000, bytes: 0, snapshotBytes: 0 }), + isSnapshotDue({ inputs: 1, bytes: 1_048_576, snapshotBytes: 900_000 }), + isSnapshotDue({ inputs: 1, bytes: 2_000_000, snapshotBytes: 3_000_000 }), + isSnapshotDue({ inputs: 1, bytes: 3_000_000, snapshotBytes: 3_000_000 }), + ]).toEqual([false, true, true, false, true]); }); - it('round-trips a running state through its chunks', () => { - const snapshot: Snapshot = { format: 1, executionId, version: 1000, state: runningState }; + it('round-trips a running state through its chunks, with the bytes of history it covers', () => { + const snapshot: Snapshot = { + format: stateFormat, + executionId, + version: 1000, + historyBytes: 2_100_000, + state: runningState, + }; expect(snapshotFromChunks(snapshotChunks(snapshot))).toEqual(Result.succeed(snapshot)); }); - it('is stored in chunks of at most 65,536 UTF-16 code units that never split a character', () => { - const snapshot = snapshotHolding(`${'😀'.repeat(70_000)}x${'😀'.repeat(70_000)}é`); + it('is stored in chunks of at most 1 MiB of UTF-8, well under the 2 MB a row takes, that never split a character', () => { + const snapshot = snapshotHolding(`${'😀'.repeat(300_000)}x${'€'.repeat(200_000)}é${'é'.repeat(300_000)}`); const chunks = snapshotChunks(snapshot); expect(chunks.length).toBeGreaterThan(2); - expect(chunks.every((chunk) => chunk.length <= snapshotChunkLength)).toBe(true); + expect(Math.max(...chunks.map((chunk) => utf8.encode(chunk).byteLength))).toBeLessThanOrEqual( + mostSnapshotChunkBytes, + ); expect(chunks.filter((chunk) => highSurrogate.test(chunk))).toEqual([]); expect(chunks.filter((chunk) => lowSurrogate.test(chunk))).toEqual([]); - expect(Math.max(...chunks.map((chunk) => new TextEncoder().encode(chunk).byteLength))).toBeLessThanOrEqual(196_608); expect(snapshotFromChunks(chunks)).toEqual(Result.succeed(snapshot)); }); - it('is refused when its chunks do not make a snapshot of this format', () => { - expect(Result.isFailure(snapshotFromChunks(['{"format":2}']))).toBe(true); + it('is refused when its chunks do not make a snapshot of this state format', () => { + const snapshot: Snapshot = { format: stateFormat, executionId, version: 1, historyBytes: 10, state: runningState }; + const text = snapshotChunks(snapshot) + .join('') + .replace(`"format":${stateFormat}`, `"format":${stateFormat + 1}`); + + expect(Result.isFailure(snapshotFromChunks([text]))).toBe(true); }); }); diff --git a/packages/workflow-engine/src/run-log/snapshot.ts b/packages/workflow-engine/src/run-log/snapshot.ts index 4467d5702..c55932202 100644 --- a/packages/workflow-engine/src/run-log/snapshot.ts +++ b/packages/workflow-engine/src/run-log/snapshot.ts @@ -1,17 +1,19 @@ import { Schema } from 'effect'; import { RunStateSchema } from '../machine/run-state.ts'; +import { StateFormatSchema } from './state-format.ts'; export const snapshotEveryInputs = 1000; export const snapshotEveryBytes = 1_048_576; -export const snapshotChunkLength = 65_536; +export const mostSnapshotChunkBytes = 1_048_576; export const SnapshotSchema = Schema.Struct({ - format: Schema.Literal(1), + format: StateFormatSchema, executionId: Schema.NonEmptyString, version: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), + historyBytes: Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)), state: RunStateSchema, }); @@ -20,32 +22,46 @@ export type Snapshot = typeof SnapshotSchema.Type; export interface SinceSnapshot { readonly inputs: number; readonly bytes: number; + readonly snapshotBytes: number; } const SnapshotTextSchema = Schema.fromJsonString(SnapshotSchema); const encodeSnapshot = Schema.encodeSync(SnapshotTextSchema); -export const decodeSnapshot = Schema.decodeUnknownResult(SnapshotTextSchema); +const decodeSnapshot = Schema.decodeUnknownResult(SnapshotTextSchema); -export function isSnapshotDue({ inputs, bytes }: SinceSnapshot): boolean { - return inputs >= snapshotEveryInputs || bytes >= snapshotEveryBytes; +export function isSnapshotDue({ inputs, bytes, snapshotBytes }: SinceSnapshot): boolean { + return inputs >= snapshotEveryInputs || bytes >= Math.max(snapshotEveryBytes, snapshotBytes); } -const lastSingleUnitCodePoint = 0xff_ff; - -function chunkEnd(text: string, start: number): number { - const end = Math.min(start + snapshotChunkLength, text.length); - const splitsAPair = Number(text.codePointAt(end - 1)) > lastSingleUnitCodePoint; - return splitsAPair ? end - 1 : end; +function utf8LengthOf(codePoint: number): number { + if (codePoint < 0x80) { + return 1; + } + if (codePoint < 0x8_00) { + return 2; + } + return codePoint < 0x1_00_00 ? 3 : 4; } export function snapshotChunks(snapshot: Snapshot): readonly string[] { const text = encodeSnapshot(snapshot); const chunks: string[] = []; - for (let start = 0; start < text.length; start = chunkEnd(text, start)) { - chunks.push(text.slice(start, chunkEnd(text, start))); + let start = 0; + let end = 0; + let bytes = 0; + for (const character of text) { + const length = utf8LengthOf(Number(character.codePointAt(0))); + if (bytes + length > mostSnapshotChunkBytes) { + chunks.push(text.slice(start, end)); + start = end; + bytes = 0; + } + bytes += length; + end += character.length; } + chunks.push(text.slice(start)); return chunks; } diff --git a/packages/workflow-engine/src/run-log/state-format.ts b/packages/workflow-engine/src/run-log/state-format.ts new file mode 100644 index 000000000..6149e1507 --- /dev/null +++ b/packages/workflow-engine/src/run-log/state-format.ts @@ -0,0 +1,5 @@ +import { Schema } from 'effect'; + +export const stateFormat = 1; + +export const StateFormatSchema = Schema.Literal(stateFormat); diff --git a/packages/workflow-engine/src/run-log/state-patch.test.ts b/packages/workflow-engine/src/run-log/state-patch.test.ts new file mode 100644 index 000000000..60d25e4ee --- /dev/null +++ b/packages/workflow-engine/src/run-log/state-patch.test.ts @@ -0,0 +1,51 @@ +import { describe, expect, it } from 'vitest'; + +import { applyStatePatch, PatchFailed } from '../index.ts'; + +const state = { + machine: { context: 0, values: { '0': { value: 1 } } }, + list: ['a', 'b'], + rows: [{ name: 'a' }], + 'a/b': { '~': 1 }, +}; + +describe('a state patch', () => { + it('adds, replaces and removes members of objects and lists, by JSON Pointer, without changing the state it was given', () => { + const patched = applyStatePatch(state, [ + { op: 'add', path: '/machine/values/1', value: { value: 2 } }, + { op: 'replace', path: '/machine/context', value: 1 }, + { op: 'add', path: '/list/1', value: 'between' }, + { op: 'add', path: '/list/-', value: 'last' }, + { op: 'remove', path: '/list/0' }, + { op: 'replace', path: '/list/2', value: 'z' }, + { op: 'replace', path: '/a~1b/~0', value: 2 }, + { op: 'replace', path: '/rows/0/name', value: 'b' }, + { op: 'remove', path: '/machine/values/0' }, + ]); + + expect(patched).toEqual({ + machine: { context: 1, values: { '1': { value: 2 } } }, + list: ['between', 'b', 'z'], + rows: [{ name: 'b' }], + 'a/b': { '~': 2 }, + }); + expect(state.list).toEqual(['a', 'b']); + }); + + it.each([ + [{ op: 'replace', path: '/machine/missing', value: 1 }, 'The state has no member missing'], + [{ op: 'remove', path: '/list/2' }, 'The state has no member 2'], + [{ op: 'add', path: '/machine/context', value: 1 }, 'The member context is there already'], + [{ op: 'add', path: '/list/3', value: 'x' }, 'The list has no position 3'], + [{ op: 'add', path: '/list/one', value: 'x' }, 'The list has no position one'], + [{ op: 'replace', path: '/gone/deeper', value: 1 }, 'The state has no member gone on the way'], + [ + { op: 'replace', path: '/machine/context/deeper', value: 1 }, + 'The path goes through a value that is neither an object nor a list', + ], + [{ op: 'replace', path: '', value: {} }, 'A path names a member under the root, such as /machine/context'], + [{ op: 'remove', path: 'machine' }, 'A path names a member under the root, such as /machine/context'], + ] as const)('fails at once, never guessing, on %j', (operation, detail) => { + expect(() => applyStatePatch(state, [operation])).toThrow(new PatchFailed({ ...operation, detail })); + }); +}); diff --git a/packages/workflow-engine/src/run-log/state-patch.ts b/packages/workflow-engine/src/run-log/state-patch.ts index 31e0e1ca5..f3e416f62 100644 --- a/packages/workflow-engine/src/run-log/state-patch.ts +++ b/packages/workflow-engine/src/run-log/state-patch.ts @@ -1,13 +1,113 @@ -import { Schema } from 'effect'; - -const PointerSchema = Schema.String; +import { Data, Schema } from 'effect'; export const PatchOperationSchema = Schema.Union([ - Schema.Struct({ op: Schema.Literal('add'), path: PointerSchema, value: Schema.Json }), - Schema.Struct({ op: Schema.Literal('replace'), path: PointerSchema, value: Schema.Json }), - Schema.Struct({ op: Schema.Literal('remove'), path: PointerSchema }), + Schema.Struct({ op: Schema.Literal('add'), path: Schema.String, value: Schema.Json }), + Schema.Struct({ op: Schema.Literal('replace'), path: Schema.String, value: Schema.Json }), + Schema.Struct({ op: Schema.Literal('remove'), path: Schema.String }), ]); export type PatchOperation = typeof PatchOperationSchema.Type; export type StatePatch = readonly PatchOperation[]; + +export class PatchFailed extends Data.TaggedError('patch_failed')<{ + readonly op: PatchOperation['op']; + readonly path: string; + readonly detail: string; +}> {} + +type Container = Readonly> | readonly unknown[]; + +type Edit = (container: Container, token: string) => unknown; + +const arrayIndex = /^(?:0|[1-9]\d*)$/u; + +function isList(value: unknown): value is readonly unknown[] { + return Array.isArray(value); +} + +function isRecord(value: unknown): value is Readonly> { + return typeof value === 'object' && value !== null && !Array.isArray(value); +} + +function tokensOf(operation: PatchOperation): readonly string[] { + if (operation.path === '' || !operation.path.startsWith('/')) { + throw new PatchFailed({ ...operation, detail: 'A path names a member under the root, such as /machine/context' }); + } + return operation.path + .slice(1) + .split('/') + .map((token) => token.replaceAll('~1', '/').replaceAll('~0', '~')); +} + +function indexIn(list: readonly unknown[], token: string, operation: PatchOperation, room: number): number { + const index = token === '-' && operation.op === 'add' ? list.length : Number(token); + if (!(arrayIndex.test(token) || token === '-') || index > list.length - 1 + room) { + throw new PatchFailed({ ...operation, detail: `The list has no position ${token}` }); + } + return index; +} + +function hasMember(container: Container, token: string): boolean { + return isList(container) + ? arrayIndex.test(token) && Number(token) < container.length + : Object.hasOwn(container, token); +} + +function childOf(container: Container, token: string): unknown { + return isList(container) ? container[Number(token)] : container[token]; +} + +function withChild(container: Container, token: string, child: unknown): Container { + if (isList(container)) { + return container.map((item, index) => (index === Number(token) ? child : item)); + } + return { ...container, [token]: child }; +} + +function edited(document: unknown, tokens: readonly string[], operation: PatchOperation, edit: Edit): unknown { + const [token = '', ...rest] = tokens; + if (!isRecord(document) && !isList(document)) { + throw new PatchFailed({ + ...operation, + detail: 'The path goes through a value that is neither an object nor a list', + }); + } + if (rest.length === 0) { + return edit(document, token); + } + if (!hasMember(document, token)) { + throw new PatchFailed({ ...operation, detail: `The state has no member ${token} on the way` }); + } + return withChild(document, token, edited(childOf(document, token), rest, operation, edit)); +} + +function editOf(operation: PatchOperation): Edit { + return (container, token) => { + if (operation.op === 'add') { + if (isList(container)) { + return container.toSpliced(indexIn(container, token, operation, 1), 0, operation.value); + } + if (Object.hasOwn(container, token)) { + throw new PatchFailed({ ...operation, detail: `The member ${token} is there already` }); + } + return { ...container, [token]: operation.value }; + } + if (!hasMember(container, token)) { + throw new PatchFailed({ ...operation, detail: `The state has no member ${token}` }); + } + if (operation.op === 'replace') { + return withChild(container, token, operation.value); + } + return isList(container) + ? container.toSpliced(Number(token), 1) + : Object.fromEntries(Object.entries(container).filter(([key]: readonly [string, unknown]) => key !== token)); + }; +} + +export function applyStatePatch(document: unknown, patch: StatePatch): unknown { + return patch.reduce( + (patched: unknown, operation: PatchOperation) => edited(patched, tokensOf(operation), operation, editOf(operation)), + document, + ); +} diff --git a/packages/workflow-engine/src/settlement/record-store.ts b/packages/workflow-engine/src/settlement/record-store.ts index 86838d88a..80fe26962 100644 --- a/packages/workflow-engine/src/settlement/record-store.ts +++ b/packages/workflow-engine/src/settlement/record-store.ts @@ -1,10 +1,27 @@ -import type { Effect } from 'effect'; +import { Schema, type Effect } from 'effect'; import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; -import type { Settle } from '../dispatch/run-output.ts'; + +export const SettlementSchema = Schema.Union([ + Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), + Schema.Struct({ + status: Schema.Literal('rejected'), + reason: Schema.Literals(['invalid_input', 'unavailable']), + detail: Schema.String, + }), + Schema.Struct({ status: Schema.Literal('failed') }), +]); + +export type Settlement = typeof SettlementSchema.Type; export type SettleReceipt = 'recorded' | 'already_recorded' | 'settled_otherwise' | 'unknown_execution'; +export interface SettleRequest { + readonly executionId: string; + readonly settlement: Settlement; +} + export interface RecordStore { - readonly settle: (settle: Settle, run: RunContext) => Effect.Effect; + readonly settle: (request: SettleRequest, run: RunContext) => Effect.Effect; + readonly liveRuns: () => Effect.Effect; } diff --git a/packages/workflow-engine/src/settlement/run-settlement.ts b/packages/workflow-engine/src/settlement/run-settlement.ts deleted file mode 100644 index 8d9305490..000000000 --- a/packages/workflow-engine/src/settlement/run-settlement.ts +++ /dev/null @@ -1,13 +0,0 @@ -import { Schema } from 'effect'; - -export const RunSettlementSchema = Schema.Union([ - Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), - Schema.Struct({ - status: Schema.Literal('rejected'), - reason: Schema.Literals(['invalid_input', 'unavailable']), - detail: Schema.String, - }), - Schema.Struct({ status: Schema.Literal('failed') }), -]); - -export type RunSettlement = typeof RunSettlementSchema.Type; diff --git a/packages/workflow-engine/src/testing/runs.ts b/packages/workflow-engine/src/testing/runs.ts index 8816ef130..9fcc6af72 100644 --- a/packages/workflow-engine/src/testing/runs.ts +++ b/packages/workflow-engine/src/testing/runs.ts @@ -1,18 +1,26 @@ import { callKeyText, type CallKey } from '../executor/call-key.ts'; -import type { RunInput } from '../machine/run-input.ts'; +import type { Started } from '../machine/run-input.ts'; import { newRun, type RunState, type TaskFrame } from '../machine/run-state.ts'; export const executionId = '0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a'; +export const document = { document: { dsl: '1.0.3', namespace: 'acme', name: 'triage', version: '1.0.0' }, do: [] }; + export const openCall: CallKey = { executionId, reference: '/do/1/fork/branches/0/ask', run: 1 }; export const armedTimer = `${executionId}/timers/1`; +export const at = 1_791_100_060_000; + +const ticket = 1; + +const approval = 2; + const waiting: TaskFrame = { reference: '/do/1/fork/branches/1/pause', run: 1, - rawInput: { ticket: 7 }, - input: { ticket: 7 }, + rawInput: ticket, + input: ticket, variables: {}, timeout: null, body: { kind: 'wait', timer: armedTimer }, @@ -21,18 +29,18 @@ const waiting: TaskFrame = { const asking: TaskFrame = { reference: openCall.reference, run: 1, - rawInput: { ticket: 7 }, - input: { ticket: 7 }, - variables: { attempt: 1 }, + rawInput: ticket, + input: ticket, + variables: { attempt: approval }, timeout: `${executionId}/timers/2`, - body: { kind: 'call', key: openCall }, + body: { kind: 'call', key: openCall, primitive: 'inference', name: 'classify' }, }; const forking: TaskFrame = { reference: '/do/1', run: 1, - rawInput: { ticket: 7 }, - input: { ticket: 7 }, + rawInput: ticket, + input: ticket, variables: {}, timeout: null, body: { @@ -50,20 +58,22 @@ export const runningState: RunState = { ...newRun, executionId, status: 'running', - workflow: { document: { document: { dsl: '1.0.3' }, do: [] }, input: { ticket: 7 } }, + workflow: { document, input: ticket }, attributes: { owner: 'tests' }, limits: { mostDurationMs: 2_592_000_000, longestCallMs: 600_000 }, - startedAt: 1_791_100_000_000, - lastInputAt: 1_791_100_000_000, + startedAt: at - 60_000, + lastInputAt: at - 60_000, + inputs: 3, random: { seed: 42, draws: 1 }, + runs: { '/do/1': 1, [openCall.reference]: 1, [waiting.reference]: 1 }, timers: { next: 3, armed: { - [armedTimer]: { purpose: 'wait', reference: waiting.reference }, - [`${executionId}/timers/2`]: { purpose: 'timeout', reference: asking.reference }, + [armedTimer]: { purpose: 'wait', reference: waiting.reference, dueAt: at + 60_000 }, + [`${executionId}/timers/2`]: { purpose: 'timeout', reference: asking.reference, dueAt: at + 600_000 }, }, }, - calls: { runs: { [openCall.reference]: 1 }, open: { [callKeyText(openCall)]: openCall } }, + calls: { [callKeyText(openCall)]: openCall }, inbox: { waiting: [{ event: { id: 'event-2', type: 'com.acme.approval', data: { approved: true } }, bytes: 80 }], waitingBytes: 80, @@ -72,32 +82,41 @@ export const runningState: RunState = { receivedBytes: 160, overflow: null, }, - heldBytes: 9000, - stepsWithoutWaiting: 0, + heldBytes: 16_500, machine: { - context: { seen: 1 }, + values: { + 0: { value: { seen: 1 }, bytes: 10, holders: 1 }, + [ticket]: { value: { ticket: 7 }, bytes: 12, holders: 7 }, + [approval]: { value: 1, bytes: 1, holders: 1 }, + }, + nextValue: 3, + context: 0, root: { reference: '/', run: 1, - rawInput: { ticket: 7 }, - input: { ticket: 7 }, + rawInput: ticket, + input: ticket, variables: {}, timeout: null, body: { kind: 'list', - list: { pointer: '/do', position: 1, data: { ticket: 7 }, variables: {}, current: forking }, + list: { + pointer: '/do', + position: 1, + data: ticket, + variables: {}, + current: { kind: 'running', task: forking }, + }, }, }, }, }; -export const at = 1_791_100_060_000; - -export const started: RunInput = { +export const started: Started = { kind: 'started', executionId, at, - document: { document: { dsl: '1.0.3' }, do: [] }, + document, input: { ticket: 7 }, limits: { mostDurationMs: 2_592_000_000, longestCallMs: 600_000 }, attributes: { owner: 'tests' }, diff --git a/packages/workflow-engine/src/timers/timers.ts b/packages/workflow-engine/src/timers/timers.ts index db4f9cf27..780bf0389 100644 --- a/packages/workflow-engine/src/timers/timers.ts +++ b/packages/workflow-engine/src/timers/timers.ts @@ -3,7 +3,12 @@ import type { Effect } from 'effect'; import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; import type { ArmTimer, CancelTimer } from '../dispatch/run-output.ts'; +export type ArmReceipt = 'armed' | 'already_armed' | 'refused_after_cancel'; + +export type TimerCancelReceipt = 'cancelled' | 'already_fired' | 'tombstoned'; + export interface Timers { - readonly arm: (timer: ArmTimer, run: RunContext) => Effect.Effect; - readonly cancel: (timer: CancelTimer, run: RunContext) => Effect.Effect; + readonly arm: (timer: ArmTimer, run: RunContext) => Effect.Effect; + readonly cancel: (timer: CancelTimer, run: RunContext) => Effect.Effect; + readonly sweep: (run: RunContext, armed: readonly ArmTimer[]) => Effect.Effect; } diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 115498ede..162e6bd4a 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -448,6 +448,9 @@ importers: packages/workflow-engine: dependencies: + '@beonauto/ledger': + specifier: workspace:* + version: link:../ledger '@beonauto/operations': specifier: workspace:* version: link:../operations From 72f9fa0d1c329f59869aa072d43a0d66684893c2 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:03:22 +0100 Subject: [PATCH 05/23] docs(workflow-engine): describe the revised contract and its invariants The README describes the run's log as a state-transition log, the value table and how it bounds a patch, submission outcomes, receipts and tombstones, the dispatch order, the sweep over live runs, state formats, the new bounds and snapshot rules, Node's serialisation, and invariants 22 to 31, each marked as the engine's or an adapter's. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/README.md | 231 ++++++++++++++++++----------- 1 file changed, 142 insertions(+), 89 deletions(-) diff --git a/packages/workflow-engine/README.md b/packages/workflow-engine/README.md index 1903b5b85..57fdad383 100644 --- a/packages/workflow-engine/README.md +++ b/packages/workflow-engine/README.md @@ -1,133 +1,186 @@ # @beonauto/workflow-engine -The core of the workflow engine that runs on the ledger: the contract between the machine that runs a workflow and the adapters that store, time and execute for it. It knows workflows, the inputs a run takes and an executor that performs calls. It does not know brains, prompts, models or specs: whatever an adapter needs to know about the run, such as who started it, it passes as opaque `attributes` and gets back with every output. +The core of the workflow engine that runs on the ledger: the contract between the machine that runs a workflow and the adapters that store, time and execute for it. It knows workflows, the inputs a run takes and an executor that performs calls. It does not know brains, prompts, models or specs: whatever an adapter needs to know about a run, such as who started it, it passes as opaque `attributes` and gets back with every output. -The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. This package holds the contract: the types, the ports, the idempotency keys, the dispatch watermark and the invariants below. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, unchanged. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. +The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. This package holds the contract: the types, the ports, the idempotency keys, the dispatch watermark, the fold of a run's log and the invariants below. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, unchanged. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. ## How a run moves -1. An adapter submits an input for an execution, with `at`, the time on its own clock. The machine never reads a clock. -2. Holding the run's serialisation, the engine loads the run: its latest snapshot and the events after it, folded with `evolve`. -3. It asks the decider. A stale input, or one that changes nothing, appends nothing. Otherwise the decision is one event, appended with the version the engine read as the expected version. This is the load-decide-append loop of `@beonauto/ledger`: a version conflict loads the run again and decides again, up to three more times. -4. When a snapshot is due, the engine saves one. -5. It dispatches the outputs of every event above the run's dispatch watermark, in the order of the stream, and then moves the watermark up to the last event whose outputs were all dispatched. -6. `wake(executionId)` does step 5 again. An adapter wakes a run after a crash, on a restart, and from a sweep that finds runs whose watermark is behind their stream. +1. An adapter submits an input for an execution, with `at`, the time on its own clock. The machine never reads a clock, and it takes `max(at, lastInputAt)` as the time of the input, so time in a run never goes back however the adapters' clocks drift. +2. Holding the run's serialisation, the engine loads the run: `RunStore.load` gives the latest snapshot and the events after it, and `loadedRunOf` folds them. +3. `staleReasonOf(state, input)` says whether the input can still change the run. An input to a run whose `started` has not arrived is `not_started`: the engine answers `{ outcome: 'not_started' }`, which an adapter answers as `not_found` so the caller tries again, as the API does today. Any other reason is `stale`. Both append nothing. +4. Otherwise the machine decides, and the decision is one event, appended with the version the engine read as the expected version: the ledger's load-decide-append loop, which loads and decides again after a version conflict, up to three more times, and then fails with `Conflict`. The engine answers `{ outcome: 'applied' }`. +5. When a snapshot is due, the engine saves one. +6. It dispatches the outputs of every event above the run's dispatch watermark, in the order of the stream, and stops at the first output that fails. The watermark moves to the last event whose outputs were all dispatched. +7. `wake(executionId)` does step 6 again. `sweep()` walks the live runs, wakes each, and has `Timers.sweep` arm again any of the run's armed timers the timer store lost. ## Layers and ports -| Layer | Folder | What it holds | Port | -| ------------- | ------------------- | ----------------------------------------------------- | --------------------------------------- | -| machine | `src/machine` | inputs, state, staleness, limits, the decider's type | none: pure | -| run log | `src/run-log` | the run's events, its state patches, snapshots | `RunStore` (Emmett, one stream per run) | -| timers | `src/timers` | timer ids and what each timer is for | `Timers` | -| inbox | `src/inbox` | the external events a run receives, and their limits | none: events arrive as `event_received` | -| executor | `src/executor` | call keys and call results | `Executor` | -| dispatch | `src/dispatch` | outputs and the dispatch watermark | `DispatchWatermark` | -| serialisation | `src/serialisation` | one input at a time for each run | `RunSerialiser` | -| settlement | `src/settlement` | the settlement of a run and the store that records it | `RecordStore` (the brain's ledger) | -| engine | `src/engine` | the ports together and the engine's own interface | `WorkflowEngine` | +| Layer | Folder | What it holds | Port | +| ------------- | ------------------- | -------------------------------------------------------------- | --------------------------------------- | +| machine | `src/machine` | inputs, state, admission, the clock clamp, limits, the decider | none: pure | +| run log | `src/run-log` | events, state patches, state formats, the fold, snapshots | `RunStore` (one Emmett stream per run) | +| timers | `src/timers` | timer ids and what each timer is for | `Timers` | +| inbox | `src/inbox` | the external events a run receives, and their limits | none: events arrive as `event_received` | +| executor | `src/executor` | call keys and call results | `Executor` | +| dispatch | `src/dispatch` | outputs, the watermark, the order of a dispatch | `DispatchWatermark` | +| serialisation | `src/serialisation` | one input at a time for each run | `RunSerialiser` | +| settlement | `src/settlement` | the record store's settlements and its live runs | `RecordStore` (the brain's ledger) | +| engine | `src/engine` | the ports together and the engine's own interface | `WorkflowEngine` | Every port answers with an Effect. None of them is a clock: time comes in with the inputs. ## Inputs -| Input | Carries | Stale when | -| ------------------ | ------------------------------------------------------------------------------- | ------------------------------------------------------------- | -| `started` | the document, the input, the limits, the attributes and a seed for random draws | the run has started before | -| `timer_fired` | the timer id | the timer is not armed: never armed, already fired, cancelled | -| `call_answered` | the call key and the result: succeeded, rejected, failed or unreachable | the call is not open: never started, answered, cancelled | -| `event_received` | the event, with an `id` of 1 to 256 characters and a `type` | the run has received an event with that id | -| `cancel_requested` | nothing more | a cancel was requested before | +| Input | Carries | Stale when | +| ------------------ | ------------------------------------------------------------------------------- | -------------------------------------------------------- | +| `started` | the document, the input, the limits, the attributes and a seed for random draws | the run has started (`started_before`) | +| `timer_fired` | the timer id | the timer is not armed: never armed, fired, cancelled | +| `call_answered` | the call key and the result: succeeded, rejected, failed or unreachable | the call is not open: never started, answered, cancelled | +| `event_received` | the event, with an `id` of 1 to 256 characters and a `type` | the run has received an event with that id | +| `cancel_requested` | nothing more | a cancel was requested before | -Every input carries the execution id and `at`. Any input but `started` is stale for a run that has not started, has ended, or is another execution. `isStale(state, input)` is this table. +Every input carries the execution id and `at`. Any input but `started` is `not_started` for a run that has not started and `run_ended` for one that has ended. Two inputs are never taken as stale, because they mean an adapter routed wrongly: an input for another execution, and a second `started` with another document. `staleReasonOf` dies on them with `RunMismatch`. -## Events +## The run's log -A run's stream holds one kind of event, `input_applied`: +A run's stream is a state-transition log, not classic event sourcing. Each event records the change an input made to the run's state, as a patch, rather than a domain fact for the fold to interpret. Replaying a run therefore applies patches and evaluates nothing: no expression, no retry arithmetic, no version of the machine's code. A run started under one version of the machine loads under the next. What the events give up in meaning they get back in two fields written for people reading the log. -- `receipt`: the kind of the input, the key it is deduplicated by and its `at`. The input's payload is not stored again: what it changed is in the patch. -- `patch`: the change the input made to the run's state, as JSON Patch operations (RFC 6902: `add`, `replace`, `remove`) addressed by JSON Pointer. +Each event, `input_applied`, holds: + +- `format`: the state format its patch applies to. +- `receipt`: the kind of the input, the key it is deduplicated by and its time; for an answer, the result's status, and for an external event, its type. The input's payload is not stored again: what it changed is in the patch. +- `steps`: each task the input stepped, as `{ reference, run, outcome }`, with outcome `started`, `skipped`, `waiting`, `completed`, `raised`, `timed_out` or `cancelled`. +- `patch`: the change to the state, as JSON Patch operations (RFC 6902 `add`, `replace` and `remove`) addressed by JSON Pointer. - `outputs`: what the engine must do because of this input. -`evolve` applies the patch and nothing else, so replaying a run never evaluates an expression and never depends on the version of the machine that decided it. +An input that is stale or not started appends nothing; logging it is the adapter's job. + +## Held data and the size of a patch + +Values a run holds live once, in the state's value table, `machine.values`: the run's input, the result of a call, the data of an event, the output of an expression. Each value has an id that counts up within the run and is never reused, its size in bytes of UTF-8 JSON, and the number of holders. Frames, list cursors, variables, branches and the context refer to values by id. Passing a value on, as a call's output becomes the cursor's data, the next task's input and the context, adds a holder, not a copy; a value leaves the table when its last holder lets go. + +The data a run holds is the sizes of the values in the table, plus 4 KiB for each open frame, plus the document. That bounds the state, and so every snapshot, at 4 MiB of held data. It also bounds a patch: an input adds each value it made once, as `add /machine/values/`, and otherwise moves ids, so a call that answers with 1 MiB gives a patch of about 1 MiB however many places the answer lands. + +A transformation can still make more data than it was given. The machine measures each event before the append; one over 1.5 MiB ends the run with a `raised` runtime error, status 500, the ending today's limits give, recorded as a small final event: the outcome, the frames gone, timers and calls cancelled and the execution settled. + +Frames never store the document: they name tasks by reference, a JSON Pointer into `workflow.document`. + +## State formats + +Every event and every snapshot names its state format, `stateFormat`, today 1. A change to the state's schema is a new format, and a new format ships only with an upcaster for snapshots and one for patches. `evolve` applies a patch strictly: an `add` to a member that exists, or a `replace` or `remove` of one that does not, dies with `PatchFailed`, and the result must decode as the state, so a skew between a log and the code that reads it is caught when the run loads, never folded into a wrong state. + +## Outputs and receipts -## Outputs +| Output | Carries | Idempotent by | Port | Receipts | +| -------------- | ------------------------------------------------------------- | ------------- | -------------------- | ------------------------------------------------------------------------ | +| `arm_timer` | timer id, due time, purpose | timer id | `Timers.arm` | `armed`, `already_armed`, `refused_after_cancel` | +| `cancel_timer` | timer id | timer id | `Timers.cancel` | `cancelled`, `already_fired`, `tombstoned` | +| `start_call` | call key, the function, its arguments, the longest it may run | call key | `Executor.start` | `started`, `already_started`, `refused_after_cancel` | +| `cancel_call` | call key | call key | `Executor.cancel` | `cancelled`, `already_answered`, `tombstoned` | +| `settle` | execution id and settlement | execution id | `RecordStore.settle` | `recorded`, `already_recorded`, `settled_otherwise`, `unknown_execution` | -| Output | Carries | Idempotent by | Port | -| -------------- | ---------------------------------------------------------------------------------- | ------------- | -------------------- | -| `arm_timer` | timer id, due time, purpose | timer id | `Timers.arm` | -| `cancel_timer` | timer id | timer id | `Timers.cancel` | -| `start_call` | call key, the call as the document names it, its arguments, the longest it may run | call key | `Executor.start` | -| `cancel_call` | call key | call key | `Executor.cancel` | -| `settle` | execution id and settlement: succeeded, rejected or failed | execution id | `RecordStore.settle` | +A cancel for a key the timer store or the executor has never seen is recorded as a tombstone, so a start of that key that arrives later is refused. Answers come back as inputs: a fired timer as `timer_fired`, a finished call as `call_answered`. Every receipt is final; only a failure to answer, `DispatchFailed`, leaves an output to be dispatched again. -Answers come back as inputs: a fired timer as `timer_fired`, a finished call as `call_answered`. `RecordStore.settle` answers `recorded`, `already_recorded`, `settled_otherwise` or `unknown_execution`; each of them is final, and only a failure to answer leaves the output to be dispatched again. +The settlement vocabulary is the record store's: `succeeded` with an output, `rejected` with `invalid_input` or `unavailable` and a detail, or `failed`. + +The machine checks only the size of a call's arguments, at most 264 KiB as JSON. The executor checks what they mean, such as a primitive, a name and an input, and that a workflow does not call another workflow, and answers `rejected` with reason `invalid_arguments` and a detail when they are wrong. The machine maps a rejection's reason to the task's error through the table the interpreter uses today (`primitives/orchestration/src/interpreter/call-task.ts`), with `invalid_arguments` as a `validation` error, status 400, carrying the executor's detail. ## Idempotency keys -| Key | Made of | Where it is deduplicated | -| ------------ | --------------------------------------------------------- | ---------------------------------------------------- | -| execution id | given by the adapter that starts the run | the run's status; `RecordStore` by execution id | -| timer id | `/timers/`, n counting up within the run | `state.timers.armed`; `Timers` by id | -| call key | execution id, the task's reference, the run of that task | `state.calls.open`; `Executor` by `callKeyText(key)` | -| event id | the external event's own `id` | `state.inbox.receivedIds` | -| snapshot | execution id and stream version | the run store keeps only the latest | -| watermark | execution id | the watermark store | +| Key | Made of | Deduplicated in | +| ------------ | --------------------------------------------------------- | ----------------------------------------------- | +| execution id | given by the adapter that starts the run | the run's status; `RecordStore` by execution id | +| timer id | `/timers/`, n counting up within the run | `state.timers.armed`; `Timers` by id | +| call key | execution id, the task's reference, the run of that task | `state.calls`; `Executor` by `callKeyText(key)` | +| event id | the external event's own `id` | `state.inbox.receivedIds` | +| value id | n counting up within the run | `state.machine.values` | + +A run keeps one counter of runs for each task reference, `state.runs`, so the third time a task runs its run is 3, and a call it starts has that run in its key. An adapter that gives a call its own execution, as a call to a spec has, derives that execution's id from the call key, so a call dispatched twice is one execution. -The event store does not deduplicate anything: Emmett appends a message with an id it has seen before as a new message, on SQLite and on D1. Deduplication lives in the run's state. +The event store deduplicates nothing: Emmett appends a message with an id it has seen before as a new message, on SQLite and on D1. Deduplication lives in the run's state. ## The dispatch watermark -The watermark of a run is a stream version. Every output of every event at or below it has been dispatched at least once. Outputs above it are dispatched on the next submit or wake, in the order of the stream, and the watermark then moves up to the last event whose outputs were all dispatched. A crash between the append and the dispatch loses nothing: the outputs are in the event, and the next wake dispatches them. +The watermark of a run is a stream version. Every output of every event at or below it has been dispatched at least once. A dispatch takes the outputs above it in the order of the stream and stops at the first that fails; the watermark then moves to the last event whose outputs were all dispatched, `dispatchedThrough(watermark, events, firstFailed)`, and never goes down. A crash between the append and the dispatch loses nothing: the outputs are in the event, and the next wake dispatches them. + +## Serialisation + +One input at a time for each run is the adapter's job. On Cloudflare it is the run's Durable Object, whose single thread takes one request at a time. On Node it is one process for each SQLite file, holding a lock per run in memory; a second process on the same file is outside the contract, and nothing claims or leases a run. A PostgreSQL adapter, later, will take a lease per run. Whatever slips past, the expected version of the append catches. + +On Node the timers table goes through the ledger's own SQLite driver or lives in a separate file: written through a second SQLite library to the ledger's file, committed cancels were lost (`spikes/node/results/lost-write-repeat.json` on branch `spike/engine-node`). + +## Live runs and the sweep + +The live runs are the executions the record store has recorded `started` with `finishesLater` and not yet settled, what `awaitsSettlement` in `@beonauto/specs` answers: `RecordStore.liveRuns()`. On Node that is one query over the one file. `sweep()` walks them; for each it calls `wake`, and `Timers.sweep` with the run's armed timers, which arms again any the timer store lost, such as an alarm of an evicted Durable Object that gave up after its retries. ## State and snapshots -A run's state is plain JSON: no `Map`, `Set`, `Date`, `undefined`, class or function, so `JSON.parse(JSON.stringify(state))` is the state. It holds the run's definition and attributes, the machine's frames and context, the armed timers and open calls, the inbox, the random draws made so far, the cancel request and, once the run ends, its outcome. +A run's state is plain JSON: no `Map`, `Set`, `Date`, `undefined`, class or function, so `JSON.parse(JSON.stringify(state))` is the state. -A snapshot is `{ format, executionId, version, state }`, the state folded from the events up to `version`. It is due after 1,000 inputs or 1 MiB of events since the last one, whichever comes first. It is stored in chunks of at most 65,536 UTF-16 code units, at most 192 KiB in UTF-8 each, cut between characters, never inside one; the run store keeps only the latest snapshot. +A snapshot is `{ format, executionId, version, historyBytes, state }`, the state folded from events 1 to `version`, written only once event `version` is durable; the run store keeps only the latest. A snapshot is due after 1,000 inputs, or once the events since the last snapshot take as many bytes as that snapshot did, and at least 1 MiB, so writing snapshots never costs more bytes than the history it covers. + +A snapshot holds at most about 5.3 MiB: the held data (4 MiB: the values, the document and the frames), the events waiting in the inbox (1 MiB), the ids of the events received (1,024 of at most 256 characters), and the timers, calls and run counters, a few dozen bytes for each frame. It is stored in chunks of at most 1 MiB of UTF-8, cut between characters, never inside one. D1 and Durable Object SQLite both take rows of at most 2 MB; a 1 MiB chunk leaves room for the row's other columns, and an event, at most 1.5 MiB, fits in a row too. ## Limits -| Limit | Value | Why | -| --------------------------- | ------------ | -------------------------------------------------------------------------------- | -| data a run holds | 4 MiB | measured by size, not by object identity; it bounds the state, so every snapshot | -| one event | 1.5 MiB | a call may answer with 1 MiB; D1 and Durable Object rows take up to 2 MB | -| tasks in one input | 100 | then the machine arms a timer due at once and goes on when it fires | -| tasks without waiting | 10,000 | as the interpreter has it | -| events waiting in the inbox | 64, 1 MiB | as the interpreter has it | -| events a run receives | 1,024, 4 MiB | as the interpreter has it; it also bounds the event ids kept for deduplication | +| Limit | Value | When it is reached | +| --------------------------- | ------------ | ------------------------------------------------------------------- | +| data a run holds | 4 MiB | the run ends, raised, as today | +| one event | 1.5 MiB | the run ends, raised, in a small final event | +| arguments of a call | 264 KiB | the task raises a validation error | +| tasks in one input | 100 | the machine arms a timer due at once and goes on when it fires | +| expression work | 8,000,000 | for one expression, as today | +| work in one input | 16,000,000 | as above: a timer due at once, then the rest | +| tasks without waiting | 10,000 | the run ends, raised, as today | +| events waiting in the inbox | 64, 1 MiB | the run ends, raised, as today | +| events a run receives | 1,024, 4 MiB | the run ends, raised, as today; it also bounds the ids kept | +| inputs a run takes | 100,000 | checked in `decide` from `state.inputs`: the run ends, raised | +| history | 512 MiB | checked by the engine from the bytes appended: the run ends, raised | + +## On the ledger's loop -The 16 MiB a run could hold on Temporal comes down to 4 MiB, and snapshots are chunked: a snapshot is then at most about 5 MiB in chunks of at most 192 KiB. Bringing the limit under 1 MiB instead would break what callers rely on today: a call answers with up to 1 MiB, and a workflow's output may take 1 MiB. Temporal's limits on history, 40,000 events and 8 MiB, go: a run is loaded from its snapshot, so the length of its history no longer costs time. +The engine decides and appends the way the ledger does, with the ledger's own pieces from `@beonauto/ledger`. The run store's append fails with the ledger's `VersionConflict`, and the engine retries it with `retriedOnVersionConflict`, which fails with `Conflict` after three more attempts. The SQLite run store, in the adapters' step, reads a run's tail with `EventStore.read(stream, after)` and appends with `eventAppenderOf`. ## Invariants -Each sentence is something a reviewer can check against the code or a test. - -1. A run has exactly one stream, and the stream's version is the number of inputs the run applied. -2. `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source and no storage, and the same state and input give the same events. -3. Time in a run is only ever an input's `at`; random draws come from the seed in `started` and the number of draws in the state. -4. `evolve(state, event)` applies the event's patch and does nothing else. -5. An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version. -6. A stale input appends nothing, and neither does an input that would change nothing. -7. A late answer, a duplicate answer, a second delivery of an event and the fire of a cancelled timer are stale inputs. -8. The deduplication state is bounded: armed timers and open calls are what is outstanding, and a run keeps at most 1,024 event ids. -9. Timer ids are never reused within a run, and a call key's run counts up for each reference. -10. The outputs of a run are exactly the `outputs` of its events, and an output is dispatched only after the event that holds it is appended. -11. Every output is idempotent by its key, so dispatching it twice has the effect of dispatching it once. -12. The watermark never goes down, and every output of every event at or below it has been dispatched at least once. -13. A run's outcome is in its stream before the record store is asked to record it, and the `settle` output is dispatched again until the record store answers. -14. The run log and the record store may be different stores: the contract assumes two writes, each idempotent by execution id, the second retried. An adapter whose two stores are one database may make them one transaction. -15. Loading a run from its latest snapshot and the events after it gives the same state as folding its whole stream. -16. Only the latest snapshot of a run is kept, and no stored chunk of it is longer than 65,536 UTF-16 code units. -17. The state of a run is plain JSON, and the data it holds, measured by size, stays at or under 4 MiB. -18. No event is larger than 1.5 MiB as JSON, and no input runs more than 100 tasks. -19. The machine never assumes it is the only writer: it relies only on the expected version of each append. Serialising the inputs of one run is the adapter's job. -20. Nothing in this package uses a Node-only API, generates code, or imports Temporal (`src/engine/portability.test.ts`). -21. No module of the engine keeps a cache that grows with the history of a run: what an input costs in memory is bounded by the input and the state. +Each sentence is something a reviewer can check against the code or a test. **[engine]** marks what this package and the machine guarantee; **[adapter]** marks an obligation of every adapter. + +1. **[engine]** A run has exactly one stream, and the stream's version is the number of inputs the run applied. +2. **[engine]** `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source, no locale and no storage, and the same state and input give the same events (`src/engine/portability.test.ts`). +3. **[engine]** Time in a run is only ever an input's clamped `at`; random draws come from the seed in `started` and the number of draws in the state. +4. **[engine]** `evolve(state, event)` applies the event's patch strictly and does nothing else (`src/run-log/run-fold.test.ts`, `src/run-log/state-patch.test.ts`). +5. **[engine]** An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version. +6. **[engine]** A stale input appends nothing, and neither does an input to a run that has not started, nor one that would change nothing. +7. **[engine]** A late answer, a duplicate answer, a second delivery of an event and the fire of a cancelled timer are stale inputs (`src/machine/admission.test.ts`). +8. **[engine]** The deduplication state is bounded: armed timers and open calls are what is outstanding, and a run keeps at most 1,024 event ids. +9. **[engine]** Timer ids and value ids are never reused within a run, and the run counter of a task reference only counts up. +10. **[engine]** The outputs of a run are exactly the `outputs` of its events, and an output is dispatched only after the event that holds it is appended. +11. **[adapter]** Every output is idempotent by its key, so dispatching it twice has the effect of dispatching it once, and a cancel of a key never seen leaves a tombstone that refuses a later start. +12. **[engine]** The watermark never goes down, and every output of every event at or below it has been dispatched at least once (`src/dispatch/dispatch-watermark.test.ts`). +13. **[engine]** A run's outcome is in its stream before the record store is asked to record it, and the `settle` output is dispatched again until the record store answers. +14. **[adapter]** The run log and the record store may be different stores: two writes, each idempotent by execution id, the second retried. An adapter whose two stores are one database may make them one transaction. +15. **[engine]** Loading a run from its latest snapshot and the events after it gives the same state as folding its whole stream (`src/run-log/run-fold.test.ts`). +16. **[adapter]** Only the latest snapshot of a run is kept, in chunks of at most 1 MiB of UTF-8 (`src/run-log/snapshot.test.ts` for the chunks). +17. **[engine]** The state of a run is plain JSON, and the data it holds, measured by size, stays at or under 4 MiB. +18. **[engine]** No event is larger than 1.5 MiB as JSON: the machine measures each event before the append and ends the run instead (`src/run-log/run-event.test.ts`); and no input runs more than 100 tasks. +19. **[adapter]** Inputs of one run are applied one at a time; the machine relies only on the expected version of each append. +20. **[engine]** Nothing in this package uses a Node-only API, a dynamic import or code generation, or imports Temporal (`src/engine/portability.test.ts`). +21. **[engine]** No module of the engine keeps a cache that grows with the history of a run: what an input costs in memory is bounded by the input and the state (`src/engine/portability.test.ts`). +22. **[engine]** No two events of a stream have the same receipt kind and key. +23. **[engine]** `lastInputAt` never decreases (`src/machine/input-receipt.test.ts`). +24. **[engine]** A snapshot at version v is the fold of events 1 to v, and is written only after event v is durable; **[adapter]** the run store writes it only then. +25. **[engine]** A dispatch takes outputs in the order of the stream and stops at the first that fails (`src/dispatch/dispatch-watermark.test.ts`). +26. **[engine]** Every armed timer and every open call has exactly one `arm_timer` or `start_call` and at most one cancel in the stream. +27. **[engine]** An ended run has no armed timers and no open calls. +28. **[engine]** `settle` appears once in a stream, in its last event. +29. **[engine]** A run takes at most 100,000 inputs and 512 MiB of history; the input that would go past either ends the run. +30. **[engine]** Every event and every snapshot names its state format, and one of another format is refused when it is read (`src/run-log/run-event.test.ts`, `src/run-log/snapshot.test.ts`). +31. **[engine]** Every `arm_timer` is due at or after the time of the input that armed it. ## Open design points -- Whether an event may reach a run before its `started` input: the contract calls it stale, so an adapter must deliver events only to a started run. -- Whether the arguments of a `call` are checked by the machine, as the interpreter checks a spec call today, or by the executor, which would answer `rejected`. The engine itself does not know specs. -- The 1 MiB snapshot interval with a 4 MiB state can write up to four bytes of snapshot for each byte of history. -- What replaces Temporal's limits on history, if anything, now that a run's history no longer costs time to load. +- `evolve` decodes the whole state after each patch, so a load costs one decode of the state whatever the tail, and an applied input one more. The machine's step measures that against the 9 ms a fold from a snapshot every 1,000 events took in workerd (`spikes/cloudflare/results/fold.json` on branch `spike/engine-cloudflare`); checking only the patched paths is the fallback. +- A workflow that calls a workflow raises a `configuration` error today and a `validation` error once the executor rejects it with `invalid_arguments`; keeping `configuration` needs a reason of its own. +- The machine needs the DSL, expressions and policy that live in the orchestration primitive today; where the machine itself lives, beside them or with this package, is decided when it is written. +- What replaces Temporal's limits on a run's history is decided here as 100,000 inputs and 512 MiB, both well above what a workflow could reach on Temporal; real use may move them. From c20ee5b022883181e0ba3e8890bba0aa5488fc69 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:03:31 +0100 Subject: [PATCH 06/23] docs(global): weigh alternatives and costs in the workflow engine decision The record now names the run's stream a state-transition log and says why, weighs Cloudflare Workflows, Restate and Inngest, names the interpreter rewrite as the main cost, and states the snapshot rules, bounds, Node serialisation and timers, PostgreSQL assumptions, migration, retention and what replaces Temporal's UI. The 30 s CPU figure is Cloudflare's documentation: local workerd ran 35 s of CPU without a limit (spikes/cloudflare/results/fold.json, cpuLimits). Co-Authored-By: Claude Opus 5.5 --- .../0001-workflow-engine-on-the-ledger.md | 32 +++++++++++++------ 1 file changed, 22 insertions(+), 10 deletions(-) diff --git a/docs/decisions/0001-workflow-engine-on-the-ledger.md b/docs/decisions/0001-workflow-engine-on-the-ledger.md index 373a2cee7..0a9dc9a0a 100644 --- a/docs/decisions/0001-workflow-engine-on-the-ledger.md +++ b/docs/decisions/0001-workflow-engine-on-the-ledger.md @@ -12,30 +12,42 @@ Workflow specs run on Temporal today, the wrong place for them: - Tenant data is stored twice: Temporal's history holds each workflow's document, input, the outputs of its calls and its events, unencrypted, beside the brain's ledger. - We already have the store: the ledger keeps every brain's events with Emmett on SQLite, on the sqlite3, D1 and Durable Object drivers. +We weighed three other engines. Cloudflare Workflows runs only on Cloudflare, so self-hosted servers would need a second engine. Restate and Inngest are services of their own: a self-hosted server would run one beside it, as with Temporal, and each keeps step results in its own store, so tenant data would still be stored twice. + ## Decision -A run is a decider in Emmett's workflow shape, on the load-decide-append loop the ledger already uses: +A run is a decider in Emmett's workflow shape, on the load-decide-append loop the ledger already uses, with the ledger's own conflict retry: -- `decide(input, state)` says what happened and `evolve(state, event)` folds it. The inputs are a start, a timer fired, a call answered, an event received and a cancel request, each with the time it arrived. -- One stream per run. An applied input appends one event, with the version it was decided on as the expected version. The event holds the change to the state as a JSON Patch, so replay evaluates nothing, and the outputs: arm or cancel a timer, start or cancel a call, settle. -- Outputs are dispatched after the append, behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. +- `decide(input, state)` says what an input changes and `evolve(state, event)` applies it. The inputs are a start, a timer fired, a call answered, an event received and a cancel request, each with the time it arrived, never earlier than the input before it. +- One stream per run. An applied input appends one event, with the version it was decided on as the expected version. +- The stream is a state-transition log, not classic event sourcing. Each event holds the change to the state as a JSON Patch, a receipt naming the input, the steps it ran and the outputs: arm or cancel a timer, start or cancel a call, settle. Replay applies patches and evaluates nothing, so a run started under one version of the interpreter loads under the next, and the receipt and steps keep the log readable. Every event and snapshot names its state format; a new format ships with upcasters. +- Outputs are dispatched after the append, in stream order behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. A cancel that arrives before its start leaves a tombstone. - Deduplication lives in the run's state. Emmett is the store, never the engine: we use neither its workflow handler, which folds the whole stream for every input, nor its processors. -- A snapshot follows every 1,000 inputs or 1 MiB of events, whichever comes first; only the latest is kept, in chunks of at most 192 KiB. A run may hold 4 MiB instead of 16. +- A snapshot follows 1,000 inputs, or as many bytes of events as the last snapshot took and at least 1 MiB; only the latest is kept, in chunks of at most 1 MiB under the 2 MB row limit of D1 and Durable Object SQLite. A run may hold 4 MiB instead of 16, take 100,000 inputs and write 512 MiB of history. - Four adapters sit behind small ports: run store, timers, executor and record store, with the watermark and per-run serialisation beside them. - On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes runs that fell behind. -- Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process. +- Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process; a second process on that file is unsupported. Timers go through the ledger's own SQLite driver or a separate file. +- A PostgreSQL adapter, later, assumes no 2 MB row limit and takes a lease per run for serialisation. ## Consequences +The main cost is rewriting the interpreter as a machine that steps from state to state instead of an async function Temporal replays. Its DSL, expressions and policy stay. + We give up Temporal's durable timers, deduplicated delivery, replay, web UI and operator tools. We must build and keep correct: -- Timers that fire at least once, on a table in Node and on the object's alarm on Cloudflare; a fire of a timer no longer armed changes nothing. +- Timers that fire at least once; a fire of a timer no longer armed changes nothing. - Deduplication in state, every key bounded, or snapshots grow with the run. -- The outbox watermark: a crash between append and dispatch loses nothing. -- The sweep on Cloudflare, for alarms that fire late after eviction or give up after their retries. +- The watermark: a crash between append and dispatch loses nothing. +- The sweep, for alarms that fire late after eviction or give up after their retries. - Serialisation per run: an in-process lock in Node, the object's thread on Cloudflare. - Settlement in two stores, the run's stream and the brain's record, each idempotent by execution id, the second retried. +Reads over the ledger and Studio replace Temporal's UI. + +Nothing running on Temporal is migrated: nothing is in production, and the switch happens before a release. + +Ended streams are kept; a deletion policy is a later decision. + Tenant data is stored once, and a workflow needs no service beyond the server. ## Evidence @@ -51,7 +63,7 @@ Branch `spike/engine-node`: Branch `spike/engine-cloudflare`: -- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. +- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. Cloudflare documents 30 s of CPU per request by default, which local workerd does not enforce: 35 s of CPU finished under the default and 2 s under `cpu_ms = 50`. - `spikes/cloudflare/results/timers.json`: alarms fire 5 ms late at p99; an evicted object's alarm fired 15.6 s late; a sweep re-armed one that had given up. - `spikes/cloudflare/results/settlement.json`: the record was written exactly once, or given up as intended, under every injected D1 fault and crash; D1 refuses eleven events in one append. - `spikes/cloudflare/results/portability.json`: the interpreter, the DSL policy, jq and both Cloudflare ledger drivers run in workerd; the workflow SDK's validators run once precompiled. From 6e4ad3dd11d45154d16979f331558ecce60e0d3b Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:34:01 +0100 Subject: [PATCH 07/23] feat(ledger): share the load-decide-append loop as decisionLoop decisionLoop(load, append, decider) is the loop Ledger.execute runs, now exported so a caller with its own load, such as the workflow engine's load from a snapshot and its tail, runs the same loop and the same retry instead of a copy. It answers with what the load gave, the decided events and the folded state. Co-Authored-By: Claude Opus 5.5 --- packages/ledger/README.md | 1 + packages/ledger/src/decision-loop.test.ts | 84 +++++++++++++++++++++++ packages/ledger/src/index.ts | 1 + packages/ledger/src/ledger-service.ts | 64 +++++++++++++---- 4 files changed, 137 insertions(+), 13 deletions(-) create mode 100644 packages/ledger/src/decision-loop.test.ts diff --git a/packages/ledger/README.md b/packages/ledger/README.md index 033db9b41..ecc1d7823 100644 --- a/packages/ledger/README.md +++ b/packages/ledger/README.md @@ -28,6 +28,7 @@ Code that keeps its own streams, such as `@beonauto/workflow-engine`, uses the s - `EventStore.read(stream, after)` gives the events after version `after` and the version of the whole stream, so a reader that holds a snapshot at version `after` reads only the tail. Emmett answers a read past the end of a stream with version 0; `read` answers with `after` instead. - `eventAppenderOf(store)` encodes and appends events with an expected version, at most eight in one append, and fails with `VersionConflict` when another writer appended first. - `retriedOnVersionConflict(attempt)` runs a load-decide-append attempt again after a version conflict, up to three more times, and then fails with `Conflict`. +- `decisionLoop(load, append, decider)` is the load-decide-append loop itself, the one `Ledger.execute` runs: it loads, decides, appends the decided events with the loaded version expected, retries with `retriedOnVersionConflict`, and answers with what the load gave, the events and the folded state. The ledger's load folds the whole stream; a caller with snapshots passes a load that folds a snapshot and its tail. ## Creating the layer diff --git a/packages/ledger/src/decision-loop.test.ts b/packages/ledger/src/decision-loop.test.ts new file mode 100644 index 000000000..a36ea2254 --- /dev/null +++ b/packages/ledger/src/decision-loop.test.ts @@ -0,0 +1,84 @@ +import { Conflict } from '@beonauto/operations'; +import { Effect, Result } from 'effect'; +import { describe, expect, it } from 'vitest'; + +import { decisionLoop, VersionConflict, type DecisionLoop } from './index.ts'; +import { outcomeOf } from './testing/open-ledger.ts'; +import { tally, type Amounts } from './testing/tally.ts'; + +interface Loaded { + readonly state: number; + readonly version: number; + readonly loadedFrom: string; +} + +interface Appended { + readonly stream: string; + readonly events: readonly unknown[]; + readonly expectedVersion: number; +} + +interface Counted { + readonly type: 'counted'; + readonly by: number; +} + +interface Loop { + readonly loop: DecisionLoop; + readonly appended: readonly Appended[]; +} + +function loopOver(loaded: Loaded, conflicts: number): Loop { + const appended: Appended[] = []; + let remaining = conflicts; + const loop = decisionLoop( + () => Effect.succeed(loaded), + (stream, events, expectedVersion) => + Effect.suspend(() => { + remaining -= 1; + appended.push({ stream, events, expectedVersion }); + return remaining >= 0 ? Effect.fail(new VersionConflict()) : Effect.void; + }), + tally, + ); + return { loop, appended }; +} + +const snapshotted: Loaded = { state: 40, version: 7, loadedFrom: 'a snapshot and two events' }; + +describe('the decision loop over any store', () => { + it('decides on what its load gave, appends with that version expected, and gives the load back with the events', async () => { + const { loop, appended } = loopOver(snapshotted, 0); + + expect(await Effect.runPromise(loop('run/1', [1, 1]))).toEqual({ + loaded: snapshotted, + events: [ + { type: 'counted', by: 1 }, + { type: 'counted', by: 1 }, + ], + state: 42, + version: 9, + }); + expect(appended).toEqual([ + { + stream: 'run/1', + events: [ + { type: 'counted', by: 1 }, + { type: 'counted', by: 1 }, + ], + expectedVersion: 7, + }, + ]); + }); + + it('loads and decides again after a version conflict, and fails with Conflict after three more attempts', async () => { + const { loop, appended } = loopOver(snapshotted, 4); + + expect(await outcomeOf(loop('run/1', [1]))).toEqual( + Result.fail( + new Conflict({ detail: 'The state changed while the command was decided', kind: 'concurrent_change' }), + ), + ); + expect(appended).toHaveLength(4); + }); +}); diff --git a/packages/ledger/src/index.ts b/packages/ledger/src/index.ts index 5b69aced3..72c83a1a0 100644 --- a/packages/ledger/src/index.ts +++ b/packages/ledger/src/index.ts @@ -1,5 +1,6 @@ export { eventAppenderOf, type EventAppender } from './event-appender.ts'; export { eventCodecOf, type EventCodec } from './event-codec.ts'; export type { EncodedEvent, EventStore, RecordedStream } from './event-store.ts'; +export { decisionLoop, type Decided, type DecisionLoop, type StreamAppend, type StreamLoad } from './ledger-service.ts'; export { sqliteEventStore, sqliteLedgerLayer, type SQLiteStoreOptions } from './sqlite-event-store.ts'; export { retriedOnVersionConflict, VersionConflict } from './version-conflict.ts'; diff --git a/packages/ledger/src/ledger-service.ts b/packages/ledger/src/ledger-service.ts index d3e76e9ce..0c0f83142 100644 --- a/packages/ledger/src/ledger-service.ts +++ b/packages/ledger/src/ledger-service.ts @@ -1,5 +1,6 @@ import { Ledger, + type Conflict, type Decider, type DeclarableReason, type Rejection, @@ -14,30 +15,67 @@ import { foldEvents } from './fold-events.ts'; import { streamReaderOf } from './stream-reader.ts'; import { retriedOnVersionConflict, type VersionConflict } from './version-conflict.ts'; -export function makeLedger(store: EventStore): Ledger['Service'] { - const load = streamReaderOf(store); - const append = eventAppenderOf(store); +export type StreamLoad = (stream: string) => Effect.Effect; + +export type StreamAppend = ( + stream: string, + events: readonly Event[], + expectedVersion: number, +) => Effect.Effect; + +export interface Decided extends StreamState { + readonly loaded: Loaded; + readonly events: readonly Event[]; +} - const attempt = ( +export type DecisionLoop = ( + stream: string, + command: Command, +) => Effect.Effect, Rejection | Conflict>; + +export function decisionLoop< + Loaded extends StreamState, + State, + Command, + Event extends TypedEvent, + R extends DeclarableReason, +>( + load: StreamLoad, + append: StreamAppend, + decider: Decider, +): DecisionLoop { + const attempt = ( stream: string, - decider: Decider, command: Command, - ): Effect.Effect, Rejection>, VersionConflict> => + ): Effect.Effect, Rejection>, VersionConflict> => Effect.gen(function* () { - const { state, version } = yield* load(stream, decider); - const decided = decider.decide(command, state); + const loaded = yield* load(stream); + const decided = decider.decide(command, loaded.state); if (Result.isSuccess(decided) && decided.success.length > 0) { - yield* append(stream, decider.eventSchema, decided.success, version); + yield* append(stream, decided.success, loaded.version); } - return Result.map(decided, (events) => ({ - state: foldEvents(decider.evolve, state, events), - version: version + events.length, + return Result.map(decided, (events: readonly Event[]) => ({ + loaded, + events, + state: foldEvents(decider.evolve, loaded.state, events), + version: loaded.version + events.length, })); }); + return (stream, command) => + retriedOnVersionConflict(attempt(stream, command)).pipe(Effect.flatMap(Effect.fromResult)); +} + +export function makeLedger(store: EventStore): Ledger['Service'] { + const load = streamReaderOf(store); + const append = eventAppenderOf(store); return Ledger.of({ load, execute: (stream, decider, command) => - retriedOnVersionConflict(attempt(stream, decider, command)).pipe(Effect.flatMap(Effect.fromResult)), + decisionLoop( + (named: string) => load(named, decider), + (named, events, expectedVersion) => append(named, decider.eventSchema, events, expectedVersion), + decider, + )(stream, command).pipe(Effect.map(({ state, version }) => ({ state, version }))), }); } From 49c0c191078973c0f2d4ba8b682aed2f85d03af6 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:35:43 +0100 Subject: [PATCH 08/23] feat(operations): keep the settlement and call-result vocabulary once SettlementSchema and CallResultSchema, with invalidArguments, now live in @beonauto/operations beside Outcome. The workflow engine's run outputs, inputs and record store use them instead of copies; the record store's settlement in @beonauto/specs and the interpreter's RunSettlement are derived from them. The interpreter's call result and the Temporal worker's spec result keep their own shapes until Temporal goes. Co-Authored-By: Claude Opus 5.5 --- packages/operations/src/index.ts | 2 ++ .../src/outcome}/call-result.ts | 0 packages/operations/src/outcome/settlement.ts | 13 +++++++++++++ packages/specs/src/execution/execution-settler.ts | 4 ++-- .../workflow-engine/src/dispatch/run-output.ts | 2 +- packages/workflow-engine/src/index.ts | 9 +-------- packages/workflow-engine/src/machine/run-input.ts | 2 +- .../src/settlement/record-store.ts | 15 ++------------- primitives/orchestration/src/interpreter/host.ts | 7 ++----- 9 files changed, 24 insertions(+), 30 deletions(-) rename packages/{workflow-engine/src/executor => operations/src/outcome}/call-result.ts (100%) create mode 100644 packages/operations/src/outcome/settlement.ts diff --git a/packages/operations/src/index.ts b/packages/operations/src/index.ts index 0e0d1b475..78c52fdfc 100644 --- a/packages/operations/src/index.ts +++ b/packages/operations/src/index.ts @@ -5,6 +5,7 @@ export { BrainReader } from './ledger/brain-reader.ts'; export { BrainContext, type BrainAddress } from './caller/brain-context.ts'; export { BrainWriter } from './ledger/brain-writer.ts'; export { Caller, CallerIdentitySchema, type CallerIdentity } from './caller/caller.ts'; +export { CallResultSchema, invalidArguments, type CallResult, type CallStatus } from './outcome/call-result.ts'; export { makeCatalog, type Catalog } from './catalog/catalog.ts'; export { Conflict, type ConflictKind } from './outcome/conflict.ts'; export type { Decider, StreamState, TypedEvent } from './ledger/decider.ts'; @@ -55,6 +56,7 @@ export type { BrainRequest, OrgRequest } from './dispatch/request.ts'; export type { Method, Route } from './definition/route.ts'; export type { OperationKind, OperationScope } from './caller/operation-scope.ts'; export { settle } from './dispatch/settle.ts'; +export { SettlementSchema, type Settlement } from './outcome/settlement.ts'; export type { StreamReader, StreamWriter } from './ledger/stream-ports.ts'; export { Unavailable, UnavailableKindSchema, type UnavailableKind } from './outcome/unavailable.ts'; export { randomUUIDv7 } from './uuid/uuid-v7.ts'; diff --git a/packages/workflow-engine/src/executor/call-result.ts b/packages/operations/src/outcome/call-result.ts similarity index 100% rename from packages/workflow-engine/src/executor/call-result.ts rename to packages/operations/src/outcome/call-result.ts diff --git a/packages/operations/src/outcome/settlement.ts b/packages/operations/src/outcome/settlement.ts new file mode 100644 index 000000000..211b9740b --- /dev/null +++ b/packages/operations/src/outcome/settlement.ts @@ -0,0 +1,13 @@ +import { Schema } from 'effect'; + +export const SettlementSchema = Schema.Union([ + Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), + Schema.Struct({ + status: Schema.Literal('rejected'), + reason: Schema.Literals(['invalid_input', 'unavailable']), + detail: Schema.String, + }), + Schema.Struct({ status: Schema.Literal('failed') }), +]); + +export type Settlement = typeof SettlementSchema.Type; diff --git a/packages/specs/src/execution/execution-settler.ts b/packages/specs/src/execution/execution-settler.ts index 33d9e2281..b9187f0cc 100644 --- a/packages/specs/src/execution/execution-settler.ts +++ b/packages/specs/src/execution/execution-settler.ts @@ -4,6 +4,7 @@ import { OrgIdSchema, streamPrefixOfBrain, type Conflict, + type Settlement as RunSettlement, type StreamWriter, } from '@beonauto/operations'; import { DateTime, Effect, Schema } from 'effect'; @@ -22,8 +23,7 @@ export interface ExecutionAddress { export type Settlement = | { readonly status: 'succeeded'; readonly output: Schema.Json; readonly record: Schema.JsonObject } - | { readonly status: 'rejected'; readonly reason: 'invalid_input' | 'unavailable'; readonly detail: string } - | { readonly status: 'failed' }; + | Exclude; export type SettleExecution = ( execution: ExecutionAddress, diff --git a/packages/workflow-engine/src/dispatch/run-output.ts b/packages/workflow-engine/src/dispatch/run-output.ts index d1cfd0d6c..861161ecd 100644 --- a/packages/workflow-engine/src/dispatch/run-output.ts +++ b/packages/workflow-engine/src/dispatch/run-output.ts @@ -1,8 +1,8 @@ +import { SettlementSchema } from '@beonauto/operations'; import { Schema } from 'effect'; import { CallKeySchema } from '../executor/call-key.ts'; import { InstantSchema } from '../machine/instant.ts'; -import { SettlementSchema } from '../settlement/record-store.ts'; import { TimerPurposeSchema } from '../timers/timer-id.ts'; const ExecutionIdSchema = Schema.NonEmptyString; diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts index 7e414095d..edb5212af 100644 --- a/packages/workflow-engine/src/index.ts +++ b/packages/workflow-engine/src/index.ts @@ -17,7 +17,6 @@ export { } from './dispatch/run-output.ts'; export type { EnginePorts, Submission, SweepReport, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; -export { CallResultSchema, invalidArguments, type CallResult, type CallStatus } from './executor/call-result.ts'; export type { CallCancelReceipt, Executor, StartReceipt } from './executor/executor.ts'; export { ReceivedEventSchema, @@ -112,12 +111,6 @@ export { type StatePatch, } from './run-log/state-patch.ts'; export type { RunSerialiser } from './serialisation/run-serialiser.ts'; -export { - SettlementSchema, - type RecordStore, - type SettleReceipt, - type SettleRequest, - type Settlement, -} from './settlement/record-store.ts'; +export type { RecordStore, SettleReceipt, SettleRequest } from './settlement/record-store.ts'; export { TimerPurposeSchema, timerIdOf, type TimerPurpose } from './timers/timer-id.ts'; export type { ArmReceipt, TimerCancelReceipt, Timers } from './timers/timers.ts'; diff --git a/packages/workflow-engine/src/machine/run-input.ts b/packages/workflow-engine/src/machine/run-input.ts index 9583f6b81..bc97fea50 100644 --- a/packages/workflow-engine/src/machine/run-input.ts +++ b/packages/workflow-engine/src/machine/run-input.ts @@ -1,7 +1,7 @@ +import { CallResultSchema } from '@beonauto/operations'; import { Schema } from 'effect'; import { CallKeySchema } from '../executor/call-key.ts'; -import { CallResultSchema } from '../executor/call-result.ts'; import { ReceivedEventSchema } from '../inbox/received-event.ts'; import { InstantSchema } from './instant.ts'; diff --git a/packages/workflow-engine/src/settlement/record-store.ts b/packages/workflow-engine/src/settlement/record-store.ts index 80fe26962..2d35f35f8 100644 --- a/packages/workflow-engine/src/settlement/record-store.ts +++ b/packages/workflow-engine/src/settlement/record-store.ts @@ -1,19 +1,8 @@ -import { Schema, type Effect } from 'effect'; +import type { Settlement } from '@beonauto/operations'; +import type { Effect } from 'effect'; import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; -export const SettlementSchema = Schema.Union([ - Schema.Struct({ status: Schema.Literal('succeeded'), output: Schema.Json }), - Schema.Struct({ - status: Schema.Literal('rejected'), - reason: Schema.Literals(['invalid_input', 'unavailable']), - detail: Schema.String, - }), - Schema.Struct({ status: Schema.Literal('failed') }), -]); - -export type Settlement = typeof SettlementSchema.Type; - export type SettleReceipt = 'recorded' | 'already_recorded' | 'settled_otherwise' | 'unknown_execution'; export interface SettleRequest { diff --git a/primitives/orchestration/src/interpreter/host.ts b/primitives/orchestration/src/interpreter/host.ts index 30f4ecd1a..87cfb78d2 100644 --- a/primitives/orchestration/src/interpreter/host.ts +++ b/primitives/orchestration/src/interpreter/host.ts @@ -1,4 +1,4 @@ -import type { CallerIdentity } from '@beonauto/operations'; +import type { CallerIdentity, Settlement } from '@beonauto/operations'; import type { Json } from '../dsl/json.ts'; @@ -18,10 +18,7 @@ export type SpecCallResult = | { readonly status: 'rejected'; readonly reason: string; readonly detail: string } | { readonly status: 'failed'; readonly detail: string }; -export type RunSettlement = - | { readonly status: 'succeeded'; readonly output: Json } - | { readonly status: 'rejected'; readonly reason: 'invalid_input' | 'unavailable'; readonly detail: string } - | { readonly status: 'failed' }; +export type RunSettlement = Settlement; export interface SettleRequest { readonly org: string; From cd9641dd9e3351588eed4c3b82c88143827891ce Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:43:12 +0100 Subject: [PATCH 09/23] feat(orchestration): bound the cache of compiled expressions The module-level map of compiled expressions grew with every expression a process ever compiled. Under Temporal each workflow had an isolate of its own; on the workflow engine one process serves every tenant. The cache now holds at most 262,144 characters of expression source and lets go of the expression used longest ago. A compiled expression measured 22 to 34 bytes of heap per source character, so the cache holds at most about 9 MiB, and compiling one again took 10 to 150 us (1,000 expressions of 1 to 100 terms, parsed and validated). Co-Authored-By: Claude Opus 5.5 --- primitives/orchestration/README.md | 1 + .../src/dsl/bounded-cache.test.ts | 32 ++++++++++++++ .../orchestration/src/dsl/bounded-cache.ts | 43 +++++++++++++++++++ .../orchestration/src/dsl/expressions.test.ts | 25 ++++++++++- .../orchestration/src/dsl/expressions.ts | 5 ++- 5 files changed, 104 insertions(+), 2 deletions(-) create mode 100644 primitives/orchestration/src/dsl/bounded-cache.test.ts create mode 100644 primitives/orchestration/src/dsl/bounded-cache.ts diff --git a/primitives/orchestration/README.md b/primitives/orchestration/README.md index 816be58ff..6ee3bbd65 100644 --- a/primitives/orchestration/README.md +++ b/primitives/orchestration/README.md @@ -202,6 +202,7 @@ What a server spends on workflows is bounded by limits a tenant cannot raise, so - **Events** (`send_execution_event`): each at most 256 KiB as JSON, its `type` and `id` at most 256 characters and its `source` and `subject` at most 1024. A workflow holds at most 64 events it has not consumed, and 1 MiB of them by the same estimate; it takes at most 1024 events, or 4 MiB of them as JSON, over its life, counting events it ignores as repeated. One more fails the workflow with a `runtime` error at once, whatever it is doing; its execution settles `rejected`, and later events are rejected as `not_found`. - **A workflow held in the worker's cache** therefore takes at most about 18 MiB: the 16 MiB above, 1 MiB of events, the ids of up to 1024 events, and about 0.45 MiB of its own (300 idle cached workflows grew the process by 132 MiB). The worker caches at most 16 workflows, about 290 MiB in all. - **Replaying a history**, which the worker does when a workflow that is not cached has something to do, grew the process by about 7.5 times the history's bytes (66 MiB for a history of 8.8 MiB, 189 MiB for 26.4 MiB, with the bundle built ahead). The worker runs at most 2 workflow tasks at once. A workflow stops before its history passes 8 MiB, checked before each task that waits or calls a spec; tasks already running when it was checked can add their outputs (each nested output at most 1 MiB) and events can still arrive, up to the limits above, so a history ends a little past 8 MiB, about 60 to 75 MiB to replay. Temporal itself ends a workflow whose history reaches its limit, 50 MiB by default (`limit.historySize.error`): replaying one of those takes up to about 375 MiB. +- **Compiled expressions** are kept in one cache for the process, holding at most 262,144 characters of expression source and letting go of the expression used longest ago. A compiled expression measured 22 to 34 bytes of heap for each character of its source, so the cache holds at most about 9 MiB; compiling one again took 10 to 150 µs. Under Temporal each workflow had an isolate, and its cache, of its own; without a bound, one process serving every tenant would keep every expression it ever compiled. - **The server itself**, with workflows offered and the bundle built ahead, takes about 390 MiB idle. With the defaults, workflows can make the process hold at most about 390 + 290 + 2 x 75 = 830 MiB in the expected worst case, and 390 + 290 + 2 x 375 = 1430 MiB if two workflows near Temporal's own history limit replay at once. Requests to the API add what they carry (each body at most 1 MiB), and nested executions add what their primitives use, at most `ORCHESTRATION_NESTED_EXECUTIONS` (32) of them at once. An operator who wants the second figure lower sets Temporal's history limit for the namespace lower. diff --git a/primitives/orchestration/src/dsl/bounded-cache.test.ts b/primitives/orchestration/src/dsl/bounded-cache.test.ts new file mode 100644 index 000000000..2d527b2d8 --- /dev/null +++ b/primitives/orchestration/src/dsl/bounded-cache.test.ts @@ -0,0 +1,32 @@ +import { describe, expect, it } from 'vitest'; + +import { boundedCacheOf } from './bounded-cache.ts'; + +describe('a bounded cache', () => { + it('keeps entries while their keys fit in the characters it may hold', () => { + const cache = boundedCacheOf(10); + cache.set('aaaa', 1); + cache.set('bbbb', 2); + + expect([cache.get('aaaa'), cache.get('bbbb'), cache.characters()]).toEqual([1, 2, 8]); + }); + + it('forgets the entry used longest ago to make room, so a read keeps an entry', () => { + const cache = boundedCacheOf(10); + cache.set('aaaa', 1); + cache.set('bbbb', 2); + cache.get('aaaa'); + cache.set('cccc', 3); + + expect([cache.get('aaaa'), cache.get('bbbb'), cache.get('cccc'), cache.characters()]).toEqual([1, undefined, 3, 8]); + }); + + it('counts a key set again once, and keeps nothing whose key alone is larger than all it may hold', () => { + const cache = boundedCacheOf(10); + cache.set('aaaa', 1); + cache.set('aaaa', 2); + cache.set('x'.repeat(11), 3); + + expect([cache.get('aaaa'), cache.get('x'.repeat(11)), cache.characters()]).toEqual([2, undefined, 4]); + }); +}); diff --git a/primitives/orchestration/src/dsl/bounded-cache.ts b/primitives/orchestration/src/dsl/bounded-cache.ts new file mode 100644 index 000000000..5cd5b6c54 --- /dev/null +++ b/primitives/orchestration/src/dsl/bounded-cache.ts @@ -0,0 +1,43 @@ +export interface BoundedCache { + readonly get: (key: string) => V | undefined; + readonly set: (key: string, value: V) => void; + readonly characters: () => number; +} + +export function boundedCacheOf(mostCharacters: number): BoundedCache { + const entries = new Map(); + const held = { characters: 0 }; + const forget = (key: string): void => { + if (entries.delete(key)) { + held.characters -= key.length; + } + }; + const forgetOldestUntilFits = (): void => { + for (const key of entries.keys()) { + if (held.characters <= mostCharacters) { + return; + } + forget(key); + } + }; + return { + get: (key) => { + const value = entries.get(key); + if (value !== undefined) { + entries.delete(key); + entries.set(key, value); + } + return value; + }, + set: (key, value) => { + forget(key); + if (key.length > mostCharacters) { + return; + } + entries.set(key, value); + held.characters += key.length; + forgetOldestUntilFits(); + }, + characters: () => held.characters, + }; +} diff --git a/primitives/orchestration/src/dsl/expressions.test.ts b/primitives/orchestration/src/dsl/expressions.test.ts index cb3436212..fa6a8d70e 100644 --- a/primitives/orchestration/src/dsl/expressions.test.ts +++ b/primitives/orchestration/src/dsl/expressions.test.ts @@ -1,11 +1,21 @@ import { describe, expect, it } from 'vitest'; -import { checkExpression, enclosedBody, expressionSource, runExpression } from './expressions.ts'; +import { + checkExpression, + enclosedBody, + expressionSource, + mostCompiledCharacters, + runExpression, +} from './expressions.ts'; const now = Date.parse('2026-10-01T09:00:00.000Z'); const budget = { now, mostWork: 8_000_000 }; +function padded(index: number): string { + return `("${'x'.repeat(1000)}" | length) + .a + ${index}`; +} + describe('an expression', () => { it('is the body of a string enclosed in ${ }', () => { expect(enclosedBody('${ .name }')).toBe(' .name '); @@ -22,6 +32,19 @@ describe('an expression', () => { expect(runExpression('empty', null, {}, budget)).toMatchObject({ value: null }); }); + it('is compiled once and kept in a cache of bounded size, and compiled again once the cache let it go', () => { + const filling = Math.ceil(mostCompiledCharacters / padded(0).length) + 1; + + const first = runExpression(padded(0), { a: 1 }, {}, budget); + const others = Array.from({ length: filling }, (_, index) => + runExpression(padded(index + 1), { a: 0 }, {}, budget), + ); + + expect(first).toMatchObject({ value: 1001 }); + expect(others.at(-1)).toMatchObject({ value: 1000 + filling }); + expect(runExpression(padded(0), { a: 1 }, {}, budget)).toEqual(first); + }); + it('runs again on the same variables', () => { const variables = { context: { list: [1, 2, 3] } }; diff --git a/primitives/orchestration/src/dsl/expressions.ts b/primitives/orchestration/src/dsl/expressions.ts index 54525d565..1469e56cd 100644 --- a/primitives/orchestration/src/dsl/expressions.ts +++ b/primitives/orchestration/src/dsl/expressions.ts @@ -1,5 +1,6 @@ import { parse, runAst, validate, type Value } from '@gabrielbryk/jq-ts'; +import { boundedCacheOf } from './bounded-cache.ts'; import { isJson, isList, isObject, measureOf, mostValueDepth, type Json, type JsonEntry } from './json.ts'; export type Variables = Readonly>; @@ -23,7 +24,9 @@ const limits = { maxSteps: 200_000, maxDepth: 200, maxOutputs: 10_000 }; const longestProblem = 1000; -const compiled = new Map(); +export const mostCompiledCharacters = 262_144; + +const compiled = boundedCacheOf(mostCompiledCharacters); const converted = new WeakMap(); From 55a9d77313432d810a5703325e70c50dd8ec3c5c Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:43:53 +0100 Subject: [PATCH 10/23] test(workflow-engine): scan for host timers and tell caches from constants The portability scan also refuses setTimeout and setInterval and the JavaScript Temporal global. The cache scan now flags a Map or Set built empty at module level, a cache, and lets through a constant collection of literals and a WeakMap, whose entries go with the values they describe, ahead of the DSL joining the package. Co-Authored-By: Claude Opus 5.5 --- .../src/engine/portability.test.ts | 19 +++++++++++++------ 1 file changed, 13 insertions(+), 6 deletions(-) diff --git a/packages/workflow-engine/src/engine/portability.test.ts b/packages/workflow-engine/src/engine/portability.test.ts index e979841e3..61e620407 100644 --- a/packages/workflow-engine/src/engine/portability.test.ts +++ b/packages/workflow-engine/src/engine/portability.test.ts @@ -20,7 +20,9 @@ const nodeOnly: readonly Forbidden[] = [ { what: 'a dynamic import', pattern: /\bimport\(/u }, { what: 'a Node global', pattern: /\b(?:process|Buffer|require|setImmediate|__dirname)\b/u }, { what: 'code generation', pattern: /\beval\(|new Function\(/u }, + { what: 'a host timer', pattern: /\bset(?:Timeout|Interval)\(/u }, { what: 'Temporal', pattern: /@temporalio\//u }, + { what: 'the Temporal global', pattern: /\bTemporal\./u }, ]; const impure: readonly Forbidden[] = [ @@ -32,7 +34,7 @@ const impure: readonly Forbidden[] = [ const growingCache: readonly Forbidden[] = [ { what: 'a module-level collection', - pattern: /^(?:export )?const \w+(?:: [^=]+)? = new (?:Map|Set|WeakMap|WeakSet)\b/mu, + pattern: /^(?:export )?const \w+(?:: [^=]+)? = new (?:Map|Set)(?:<[^>]*>)?\(\);/mu, }, { what: 'a module-level variable', pattern: /^(?:export )?let /mu }, { what: 'a module-level list', pattern: /^(?:export )?const \w+(?:: [^=]+)? = \[\];/mu }, @@ -71,7 +73,7 @@ describe('the engine core', () => { expect(findingsIn(everySource, nodeOnly)).toEqual([]); }); - it('keeps no module-level cache that could grow with the history of a run', () => { + it('keeps no module-level cache that could grow with the history of a run: a collection built empty at module level is one; a constant collection of literals is not, nor a WeakMap, whose entries go with the values they describe', () => { expect(findingsIn(everySource, growingCache)).toEqual([]); }); }); @@ -85,10 +87,12 @@ describe('the machine and the run log', () => { describe('the checks of purity', () => { it('catch what they are there to catch', () => { - expect(caught("import { x } from 'fs';\nconst y = await import('./z.ts');", nodeOnly)).toEqual([ - 'a Node module by its bare name', - 'a dynamic import', - ]); + expect( + caught( + "import { x } from 'fs';\nconst y = await import('./z.ts');\nsetTimeout(f, 1);\nTemporal.Now.instant();", + nodeOnly, + ), + ).toEqual(['a Node module by its bare name', 'a dynamic import', 'a host timer', 'the Temporal global']); expect(caught('const t = Date.now();\nconst r = Math.random();\nIntl.DateTimeFormat();', impure)).toEqual([ 'the clock', 'a random source', @@ -97,5 +101,8 @@ describe('the checks of purity', () => { expect( caught('const seen = new Map();\nlet count = 0;\nconst all: string[] = [];', growingCache), ).toEqual(['a module-level collection', 'a module-level variable', 'a module-level list']); + expect( + caught("const kinds = new Set(['a', 'b']);\nconst sizes = new WeakMap();", growingCache), + ).toEqual([]); }); }); From b5187ca7d305de123730efb5d459b689579d6bfc Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:46:16 +0100 Subject: [PATCH 11/23] refactor(workflow-engine): move the DSL from the orchestration primitive The DSL the machine needs moves to packages/workflow-engine/src/dsl with its tests, as renames: JSON helpers, expressions with the jq work budget and the bounded cache, durations and their limits, tasks, and the policy with its checks and nesting rules. Eighteen of the files are unchanged. retained-size, the interpreter's estimate of held data under Temporal, moves beside the interpreter. The orchestration primitive imports the DSL through the engine's new dsl/* entry, so the bundled workflow code takes neither Effect nor the ledger from the engine's main entry; the replay of the recorded histories against a freshly built bundle passes. The jq dependency and its pnpm patch, registered by package and version, follow the code. The two nesting and fork tests that start a workflow through the interpreter stay with it, in refused-documents.test.ts. The workflow bundle's inputs, and the image's bundle and final stages, now include packages/workflow-engine. The pre-commit lint refuses a commit whose moved files do not resolve, so the renames and the import changes are one commit. Co-Authored-By: Claude Opus 5.5 --- packages/server/Dockerfile | 5 +++ packages/workflow-engine/package.json | 7 +++-- .../src/dsl/bounded-cache.test.ts | 0 .../workflow-engine}/src/dsl/bounded-cache.ts | 0 .../src/dsl/duration-limits.ts | 0 .../src/dsl/durations.test.ts | 0 .../workflow-engine}/src/dsl/durations.ts | 0 .../src/dsl/expression-work.test.ts | 0 .../src/dsl/expressions.test.ts | 0 .../workflow-engine}/src/dsl/expressions.ts | 0 .../workflow-engine}/src/dsl/json.test.ts | 0 .../workflow-engine}/src/dsl/json.ts | 0 .../workflow-engine}/src/dsl/nesting.test.ts | 11 +------ .../workflow-engine}/src/dsl/nesting.ts | 10 +++--- .../workflow-engine}/src/dsl/policy-checks.ts | 0 .../workflow-engine}/src/dsl/policy.test.ts | 0 .../workflow-engine}/src/dsl/policy.ts | 0 .../src/dsl/task-checks.test.ts | 0 .../src/dsl/task-policy.test.ts | 0 .../workflow-engine}/src/dsl/task-policy.ts | 0 .../workflow-engine}/src/dsl/tasks.test.ts | 0 .../workflow-engine}/src/dsl/tasks.ts | 0 .../src/dsl/workflow-bounds.test.ts | 7 ++--- .../workflow-engine/src/testing/workflows.ts | 16 ++++++++++ pnpm-lock.yaml | 12 +++++-- primitives/orchestration/README.md | 6 ++-- primitives/orchestration/package.json | 2 +- .../src/document/dsl-validation.test.ts | 2 +- .../src/document/dsl-validation.ts | 5 ++- .../src/document/workflow-document.ts | 10 +++--- .../src/document/workflow-summary.ts | 3 +- .../src/events/send-execution-event.ts | 2 +- .../src/interpreter/call-task.ts | 5 +-- .../src/interpreter/error-tasks.ts | 13 ++++++-- .../src/interpreter/evaluation.ts | 20 ++++++++++-- .../src/interpreter/flow-tasks.ts | 7 +++-- .../src/interpreter/held-data.test.ts | 2 +- .../orchestration/src/interpreter/holding.ts | 5 +-- .../orchestration/src/interpreter/host.ts | 3 +- .../orchestration/src/interpreter/inbox.ts | 5 +-- .../src/interpreter/interpreter.ts | 5 +-- .../src/interpreter/invocation.ts | 7 +++-- .../src/interpreter/listen-task.ts | 9 +++--- .../src/interpreter/raised-error.ts | 2 +- .../src/interpreter/refused-documents.test.ts | 31 +++++++++++++++++++ .../retained-size.test.ts | 0 .../src/{dsl => interpreter}/retained-size.ts | 0 .../src/interpreter/retry-policy.ts | 5 +-- .../src/interpreter/run-state.ts | 5 +-- .../src/interpreter/settlement.ts | 3 +- .../src/interpreter/task-bodies.ts | 3 +- .../src/interpreter/task-runner.ts | 7 +++-- .../orchestration/src/interpreter/timeouts.ts | 5 +-- .../src/interpreter/workflow-run.ts | 3 +- .../src/primitive/execution-input.test.ts | 2 +- .../src/primitive/orchestration-client.ts | 2 +- .../src/primitive/orchestration-primitive.ts | 2 +- .../orchestration/src/testing/workflows.ts | 2 +- .../orchestration/src/worker/activities.ts | 2 +- .../src/worker/command-paths.test.ts | 2 +- .../src/worker/orchestration-worker.test.ts | 2 +- .../orchestration/src/worker/spec-results.ts | 2 +- .../src/worker/temporal-settings.ts | 3 +- .../src/worker/workflow-bundle.test.ts | 3 +- .../src/worker/workflow-bundle.ts | 15 ++++----- .../src/workflow/interpreter-workflow.ts | 3 +- 66 files changed, 183 insertions(+), 100 deletions(-) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/bounded-cache.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/bounded-cache.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/duration-limits.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/durations.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/durations.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/expression-work.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/expressions.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/expressions.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/json.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/json.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/nesting.test.ts (83%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/nesting.ts (82%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/policy-checks.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/policy.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/policy.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/task-checks.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/task-policy.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/task-policy.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/tasks.test.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/tasks.ts (100%) rename {primitives/orchestration => packages/workflow-engine}/src/dsl/workflow-bounds.test.ts (85%) create mode 100644 packages/workflow-engine/src/testing/workflows.ts create mode 100644 primitives/orchestration/src/interpreter/refused-documents.test.ts rename primitives/orchestration/src/{dsl => interpreter}/retained-size.test.ts (100%) rename primitives/orchestration/src/{dsl => interpreter}/retained-size.ts (100%) diff --git a/packages/server/Dockerfile b/packages/server/Dockerfile index a360e5573..f05a0d3d2 100644 --- a/packages/server/Dockerfile +++ b/packages/server/Dockerfile @@ -15,6 +15,7 @@ COPY packages/identity/package.json packages/identity/ COPY packages/ledger/package.json packages/ledger/ COPY packages/operations/package.json packages/operations/ COPY packages/specs/package.json packages/specs/ +COPY packages/workflow-engine/package.json packages/workflow-engine/ COPY primitives/inference/package.json primitives/inference/ COPY primitives/orchestration/package.json primitives/orchestration/ RUN pnpm install --frozen-lockfile --prod --no-optional --ignore-scripts --filter @beonauto/server... @@ -40,11 +41,14 @@ RUN npm install --global --no-fund --no-audit pnpm@12.8.1 COPY package.json pnpm-lock.yaml pnpm-workspace.yaml ./ COPY patches patches/ COPY packages/config/package.json packages/config/ +COPY packages/ledger/package.json packages/ledger/ COPY packages/operations/package.json packages/operations/ COPY packages/specs/package.json packages/specs/ +COPY packages/workflow-engine/package.json packages/workflow-engine/ COPY primitives/orchestration/package.json primitives/orchestration/ RUN pnpm install --frozen-lockfile --prod --ignore-scripts --filter @beonauto/orchestration... COPY primitives/orchestration/build-workflow-bundle.ts primitives/orchestration/ +COPY packages/workflow-engine/src packages/workflow-engine/src COPY primitives/orchestration/src primitives/orchestration/src RUN node primitives/orchestration/build-workflow-bundle.ts /app/workflow-bundle @@ -63,6 +67,7 @@ COPY packages/ledger/src packages/ledger/src COPY packages/operations/src packages/operations/src COPY packages/server/src packages/server/src COPY packages/specs/src packages/specs/src +COPY packages/workflow-engine/src packages/workflow-engine/src COPY primitives/inference/src primitives/inference/src COPY primitives/orchestration/src primitives/orchestration/src COPY primitives/orchestration/check-workflow-bundle.ts primitives/orchestration/ diff --git a/packages/workflow-engine/package.json b/packages/workflow-engine/package.json index 1eb4cb41f..6c960bfde 100644 --- a/packages/workflow-engine/package.json +++ b/packages/workflow-engine/package.json @@ -5,7 +5,8 @@ "license": "Elastic-2.0", "type": "module", "exports": { - ".": "./src/index.ts" + ".": "./src/index.ts", + "./dsl/*": "./src/dsl/*.ts" }, "scripts": { "lint": "oxlint -c ../../.oxlintrc.json", @@ -15,10 +16,12 @@ "dependencies": { "@beonauto/ledger": "workspace:*", "@beonauto/operations": "workspace:*", + "@gabrielbryk/jq-ts": "1.7.0", "effect": "catalog:" }, "devDependencies": { "@vitest/coverage-v8": "catalog:", - "vitest": "catalog:" + "vitest": "catalog:", + "yaml": "2.9.1" } } diff --git a/primitives/orchestration/src/dsl/bounded-cache.test.ts b/packages/workflow-engine/src/dsl/bounded-cache.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/bounded-cache.test.ts rename to packages/workflow-engine/src/dsl/bounded-cache.test.ts diff --git a/primitives/orchestration/src/dsl/bounded-cache.ts b/packages/workflow-engine/src/dsl/bounded-cache.ts similarity index 100% rename from primitives/orchestration/src/dsl/bounded-cache.ts rename to packages/workflow-engine/src/dsl/bounded-cache.ts diff --git a/primitives/orchestration/src/dsl/duration-limits.ts b/packages/workflow-engine/src/dsl/duration-limits.ts similarity index 100% rename from primitives/orchestration/src/dsl/duration-limits.ts rename to packages/workflow-engine/src/dsl/duration-limits.ts diff --git a/primitives/orchestration/src/dsl/durations.test.ts b/packages/workflow-engine/src/dsl/durations.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/durations.test.ts rename to packages/workflow-engine/src/dsl/durations.test.ts diff --git a/primitives/orchestration/src/dsl/durations.ts b/packages/workflow-engine/src/dsl/durations.ts similarity index 100% rename from primitives/orchestration/src/dsl/durations.ts rename to packages/workflow-engine/src/dsl/durations.ts diff --git a/primitives/orchestration/src/dsl/expression-work.test.ts b/packages/workflow-engine/src/dsl/expression-work.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/expression-work.test.ts rename to packages/workflow-engine/src/dsl/expression-work.test.ts diff --git a/primitives/orchestration/src/dsl/expressions.test.ts b/packages/workflow-engine/src/dsl/expressions.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/expressions.test.ts rename to packages/workflow-engine/src/dsl/expressions.test.ts diff --git a/primitives/orchestration/src/dsl/expressions.ts b/packages/workflow-engine/src/dsl/expressions.ts similarity index 100% rename from primitives/orchestration/src/dsl/expressions.ts rename to packages/workflow-engine/src/dsl/expressions.ts diff --git a/primitives/orchestration/src/dsl/json.test.ts b/packages/workflow-engine/src/dsl/json.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/json.test.ts rename to packages/workflow-engine/src/dsl/json.test.ts diff --git a/primitives/orchestration/src/dsl/json.ts b/packages/workflow-engine/src/dsl/json.ts similarity index 100% rename from primitives/orchestration/src/dsl/json.ts rename to packages/workflow-engine/src/dsl/json.ts diff --git a/primitives/orchestration/src/dsl/nesting.test.ts b/packages/workflow-engine/src/dsl/nesting.test.ts similarity index 83% rename from primitives/orchestration/src/dsl/nesting.test.ts rename to packages/workflow-engine/src/dsl/nesting.test.ts index 0431c5327..bddf522c6 100644 --- a/primitives/orchestration/src/dsl/nesting.test.ts +++ b/packages/workflow-engine/src/dsl/nesting.test.ts @@ -1,6 +1,6 @@ import { describe, expect, it } from 'vitest'; -import { interpret, workflow } from '../testing/workflows.ts'; +import { workflow } from '../testing/workflows.ts'; import type { Json } from './json.ts'; import { rejectionsOf } from './policy.ts'; @@ -51,15 +51,6 @@ describe('the nesting of task lists', () => { }); }); -describe('a workflow that starts with tasks nested too deeply', () => { - it('is rejected before it runs any of them, however it was stored', async () => { - const { settlement, commands } = await interpret(workflow(`do: ${nestedDo(65)}`)); - - expect(settlement).toMatchObject({ status: 'rejected', reason: 'invalid_input' }); - expect(commands.map(({ kind }) => kind)).toStrictEqual(['deadline', 'settle']); - }); -}); - describe('the nesting of values in a document', () => { it('may reach 512 levels, and is rejected past that, at the value that is too deep', () => { expect(rejectionsOf(documentNesting(508))).toStrictEqual([]); diff --git a/primitives/orchestration/src/dsl/nesting.ts b/packages/workflow-engine/src/dsl/nesting.ts similarity index 82% rename from primitives/orchestration/src/dsl/nesting.ts rename to packages/workflow-engine/src/dsl/nesting.ts index b45873b9c..290d754d1 100644 --- a/primitives/orchestration/src/dsl/nesting.ts +++ b/packages/workflow-engine/src/dsl/nesting.ts @@ -1,5 +1,5 @@ import { entriesOf, field, isList, isObject, mostValueDepth, type Json, type JsonObject } from './json.ts'; -import { forbidden, type Located, type Rejection } from './policy-checks.ts'; +import { forbidden, type Rejection } from './policy-checks.ts'; import { nestedTaskLists, pointerTo, taskEntries, type TaskEntry } from './tasks.ts'; const mostTaskNesting = 64; @@ -22,10 +22,10 @@ function deepValueIn(value: Json, pointer: string, room: number): string | undef if (room === 0) { return pointer; } - const children: readonly Located[] = isList(value) - ? value.map((item, index): Located => [item, pointerTo(pointer, index)]) - : entriesOf(value).map(([key, item]): Located => [item, pointerTo(pointer, key)]); - return firstFound(children, ([child, at]) => deepValueIn(child ?? null, at, room - 1)); + const children: readonly (readonly [Json, string])[] = isList(value) + ? value.map((item, index): readonly [Json, string] => [item, pointerTo(pointer, index)]) + : entriesOf(value).map(([key, item]): readonly [Json, string] => [item, pointerTo(pointer, key)]); + return firstFound(children, ([child, at]) => deepValueIn(child, at, room - 1)); } function deepTaskListIn(list: Json | undefined, pointer: string, room: number): string | undefined { diff --git a/primitives/orchestration/src/dsl/policy-checks.ts b/packages/workflow-engine/src/dsl/policy-checks.ts similarity index 100% rename from primitives/orchestration/src/dsl/policy-checks.ts rename to packages/workflow-engine/src/dsl/policy-checks.ts diff --git a/primitives/orchestration/src/dsl/policy.test.ts b/packages/workflow-engine/src/dsl/policy.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/policy.test.ts rename to packages/workflow-engine/src/dsl/policy.test.ts diff --git a/primitives/orchestration/src/dsl/policy.ts b/packages/workflow-engine/src/dsl/policy.ts similarity index 100% rename from primitives/orchestration/src/dsl/policy.ts rename to packages/workflow-engine/src/dsl/policy.ts diff --git a/primitives/orchestration/src/dsl/task-checks.test.ts b/packages/workflow-engine/src/dsl/task-checks.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/task-checks.test.ts rename to packages/workflow-engine/src/dsl/task-checks.test.ts diff --git a/primitives/orchestration/src/dsl/task-policy.test.ts b/packages/workflow-engine/src/dsl/task-policy.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/task-policy.test.ts rename to packages/workflow-engine/src/dsl/task-policy.test.ts diff --git a/primitives/orchestration/src/dsl/task-policy.ts b/packages/workflow-engine/src/dsl/task-policy.ts similarity index 100% rename from primitives/orchestration/src/dsl/task-policy.ts rename to packages/workflow-engine/src/dsl/task-policy.ts diff --git a/primitives/orchestration/src/dsl/tasks.test.ts b/packages/workflow-engine/src/dsl/tasks.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/tasks.test.ts rename to packages/workflow-engine/src/dsl/tasks.test.ts diff --git a/primitives/orchestration/src/dsl/tasks.ts b/packages/workflow-engine/src/dsl/tasks.ts similarity index 100% rename from primitives/orchestration/src/dsl/tasks.ts rename to packages/workflow-engine/src/dsl/tasks.ts diff --git a/primitives/orchestration/src/dsl/workflow-bounds.test.ts b/packages/workflow-engine/src/dsl/workflow-bounds.test.ts similarity index 85% rename from primitives/orchestration/src/dsl/workflow-bounds.test.ts rename to packages/workflow-engine/src/dsl/workflow-bounds.test.ts index 63f3e19cf..8b0181e14 100644 --- a/primitives/orchestration/src/dsl/workflow-bounds.test.ts +++ b/packages/workflow-engine/src/dsl/workflow-bounds.test.ts @@ -1,6 +1,6 @@ import { describe, expect, it } from 'vitest'; -import { interpret, workflow } from '../testing/workflows.ts'; +import { workflow } from '../testing/workflows.ts'; import { durationLimitRejections } from './duration-limits.ts'; import { rejectionsOf } from './policy.ts'; @@ -50,14 +50,11 @@ describe('a fork', () => { expect(rejectionsOf(workflow(`do:\n - wide: { fork: { branches: [${branches(32)}] } }`))).toStrictEqual([]); }); - it('is refused with more, when stored and when a workflow starts with it', async () => { + it('is refused with more', () => { const document = workflow(`do:\n - wide: { fork: { branches: [${branches(33)}] } }`); - const { settlement, commands } = await interpret(document); expect(rejectionsOf(document)).toStrictEqual([ { pointer: '/do/0/wide/fork/branches', detail: 'A fork may have at most 32 branches, not 33', forbidden: true }, ]); - expect(settlement).toMatchObject({ status: 'rejected', reason: 'invalid_input' }); - expect(commands.map(({ kind }) => kind)).toStrictEqual(['deadline', 'settle']); }); }); diff --git a/packages/workflow-engine/src/testing/workflows.ts b/packages/workflow-engine/src/testing/workflows.ts new file mode 100644 index 000000000..3b882730e --- /dev/null +++ b/packages/workflow-engine/src/testing/workflows.ts @@ -0,0 +1,16 @@ +import { Schema } from 'effect'; +import { parse } from 'yaml'; + +import type { JsonObject } from '../dsl/json.ts'; + +export const header = { dsl: '1.0.3', namespace: 'acme', name: 'test', version: '1.0.0' }; + +const decodeObject = Schema.decodeUnknownSync(Schema.JsonObject); + +export function yamlObject(source: string): JsonObject { + return decodeObject(parse(source)); +} + +export function workflow(source: string): JsonObject { + return { document: header, ...yamlObject(source) }; +} diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 162e6bd4a..1d1868c6c 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -454,6 +454,9 @@ importers: '@beonauto/operations': specifier: workspace:* version: link:../operations + '@gabrielbryk/jq-ts': + specifier: 1.7.0 + version: 1.7.0(patch_hash=6b0cd1e2b6d395b64b799bf623225c6e0852f20d90393219ac87ca8797a2ce56)(typescript@7.0.2) effect: specifier: 'catalog:' version: 4.0.0 @@ -464,6 +467,9 @@ importers: vitest: specifier: 'catalog:' version: 5.0.2(@opentelemetry/api@1.9.1)(@types/node@26.6.3)(@vitest/coverage-v8@5.0.2)(vite@8.3.1(@types/node@26.6.3)(esbuild@0.28.1)(jiti@2.7.0)(terser@5.51.2)(yaml@2.9.1)) + yaml: + specifier: 2.9.1 + version: 2.9.1 primitives/inference: dependencies: @@ -541,9 +547,9 @@ importers: '@beonauto/specs': specifier: workspace:* version: link:../../packages/specs - '@gabrielbryk/jq-ts': - specifier: 1.7.0 - version: 1.7.0(patch_hash=6b0cd1e2b6d395b64b799bf623225c6e0852f20d90393219ac87ca8797a2ce56)(typescript@7.0.2) + '@beonauto/workflow-engine': + specifier: workspace:* + version: link:../../packages/workflow-engine '@openworkflowspec/sdk': specifier: 1.0.3-alpha8 version: 1.0.3-alpha8 diff --git a/primitives/orchestration/README.md b/primitives/orchestration/README.md index 6ee3bbd65..3625ca8d5 100644 --- a/primitives/orchestration/README.md +++ b/primitives/orchestration/README.md @@ -111,7 +111,7 @@ Workflows of every org share the threads of the worker, so no document may make Measured on an Apple M-series machine, the expressions that do the most work per unit, run up to the budget, take at most about 30 ms and allocate at most about 25 MB (`@base64` of a 400000-character string); before the budget, `"x" * 40000000 | length` took 0.5 s and 0.8 GB, and `"a" * 200000 | indices("a" * 100000)` 2.5 s. -The count comes from a patch of `@gabrielbryk/jq-ts` 1.7.0 (`patches/@gabrielbryk__jq-ts@1.7.0.patch` at the root of the repository), which the library offers no hook for. It adds the `maxWork` limit and the `usage` it reports, charges the operations listed above, and replaces the library's string search, which in the engine's worst case takes time in proportion to the product of the two lengths, with the Knuth-Morris-Pratt search, linear in their sum. A new version of the library needs the patch ported, and `src/dsl/expression-work.test.ts` fails for any charge that goes missing. +The count comes from a patch of `@gabrielbryk/jq-ts` 1.7.0 (`patches/@gabrielbryk__jq-ts@1.7.0.patch` at the root of the repository), which the library offers no hook for. It adds the `maxWork` limit and the `usage` it reports, charges the operations listed above, and replaces the library's string search, which in the engine's worst case takes time in proportion to the product of the two lengths, with the Knuth-Morris-Pratt search, linear in their sum. A new version of the library needs the patch ported, and `packages/workflow-engine/src/dsl/expression-work.test.ts` fails for any charge that goes missing. ### Executing a spec @@ -252,7 +252,7 @@ A worker given only the workflow code bundles it with webpack and swc when it st ## Not in this version -Starting workflows from schedules or events, `run`, `emit`, outbound calls, catalogs and functions beyond `execute_spec`, executing a workflow from a workflow (or any spec that finishes later), checking inputs and outputs against their schemas, listening `until` a condition, correlating events, search attributes, and showing the progress of a run. A function is added by adding an activity to the activity contract (`src/workflow/activity-contract.ts`), a body to the `call` task (`src/interpreter/call-task.ts`), and its name to the policy (`src/dsl/task-policy.ts`). +Starting workflows from schedules or events, `run`, `emit`, outbound calls, catalogs and functions beyond `execute_spec`, executing a workflow from a workflow (or any spec that finishes later), checking inputs and outputs against their schemas, listening `until` a condition, correlating events, search attributes, and showing the progress of a run. A function is added by adding an activity to the activity contract (`src/workflow/activity-contract.ts`), a body to the `call` task (`src/interpreter/call-task.ts`), and its name to the policy (`packages/workflow-engine/src/dsl/task-policy.ts`). ## Testing @@ -260,4 +260,4 @@ The integration tests run against a Temporal dev server that `@temporalio/testin ## Source -`src/index.ts` is the entry point, and `src/settings.ts` the entry for reading the settings without loading Temporal. `src/dsl` holds what reads a document and is safe in the workflow sandbox: JSON, durations, jq expressions, tasks and the policy. `src/document` parses a spec document: YAML, the DSL schema and graph, the issues and the summary. `src/interpreter` runs a workflow over the `WorkflowHost` interface. `src/workflow` is the Temporal workflow: the entry module that is bundled, and the host built on Temporal's workflow API. `src/worker` holds the worker, its activities, its settings, Temporal's runtime, the workflow bundle and its failure converter, which removes stack traces from recorded failures. `src/primitive` holds the primitive and its Temporal client, and `src/events` the event operation. +`src/index.ts` is the entry point, and `src/settings.ts` the entry for reading the settings without loading Temporal. What reads a document and is safe in the workflow sandbox, JSON, durations, jq expressions, tasks and the policy, is the DSL of `@beonauto/workflow-engine` (`packages/workflow-engine/src/dsl`), imported through the package's `dsl/*` entry, so the bundled workflow code takes neither Effect nor the ledger from the engine's main entry. `src/document` parses a spec document: YAML, the DSL schema and graph, the issues and the summary. `src/interpreter` runs a workflow over the `WorkflowHost` interface. `src/workflow` is the Temporal workflow: the entry module that is bundled, and the host built on Temporal's workflow API. `src/worker` holds the worker, its activities, its settings, Temporal's runtime, the workflow bundle and its failure converter, which removes stack traces from recorded failures. `src/primitive` holds the primitive and its Temporal client, and `src/events` the event operation. diff --git a/primitives/orchestration/package.json b/primitives/orchestration/package.json index 65e27dae1..514c59261 100644 --- a/primitives/orchestration/package.json +++ b/primitives/orchestration/package.json @@ -23,7 +23,7 @@ "@beonauto/config": "workspace:*", "@beonauto/operations": "workspace:*", "@beonauto/specs": "workspace:*", - "@gabrielbryk/jq-ts": "1.7.0", + "@beonauto/workflow-engine": "workspace:*", "@openworkflowspec/sdk": "1.0.3-alpha8", "@temporalio/activity": "1.24.0", "@temporalio/client": "1.24.0", diff --git a/primitives/orchestration/src/document/dsl-validation.test.ts b/primitives/orchestration/src/document/dsl-validation.test.ts index 9dc9d0701..c24f6769a 100644 --- a/primitives/orchestration/src/document/dsl-validation.test.ts +++ b/primitives/orchestration/src/document/dsl-validation.test.ts @@ -1,6 +1,6 @@ +import type { Json, JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { describe, expect, it } from 'vitest'; -import type { Json, JsonObject } from '../dsl/json.ts'; import { header, workflow } from '../testing/workflows.ts'; import { dslProblems, schemaProblems, type Problem } from './dsl-validation.ts'; diff --git a/primitives/orchestration/src/document/dsl-validation.ts b/primitives/orchestration/src/document/dsl-validation.ts index 23d837a34..d7ac1320b 100644 --- a/primitives/orchestration/src/document/dsl-validation.ts +++ b/primitives/orchestration/src/document/dsl-validation.ts @@ -1,8 +1,7 @@ +import type { JsonObject } from '@beonauto/workflow-engine/dsl/json'; +import { taskKinds } from '@beonauto/workflow-engine/dsl/tasks'; import { buildGraph, Classes, SchemaValidationError, WorkflowValidationError } from '@openworkflowspec/sdk'; -import type { JsonObject } from '../dsl/json.ts'; -import { taskKinds } from '../dsl/tasks.ts'; - export interface Problem { readonly pointer: string; readonly detail: string; diff --git a/primitives/orchestration/src/document/workflow-document.ts b/primitives/orchestration/src/document/workflow-document.ts index de8ae93d1..dd9b8057c 100644 --- a/primitives/orchestration/src/document/workflow-document.ts +++ b/primitives/orchestration/src/document/workflow-document.ts @@ -1,12 +1,12 @@ import { readYaml, type LocatedProblem, type Position, type YamlKind } from '@beonauto/config'; import { InvalidInput, type Issue } from '@beonauto/operations'; +import { durationLimitRejections } from '@beonauto/workflow-engine/dsl/duration-limits'; +import { mostValueDepth, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; +import { nestingRejections } from '@beonauto/workflow-engine/dsl/nesting'; +import { rejectionsOf } from '@beonauto/workflow-engine/dsl/policy'; +import type { Rejection } from '@beonauto/workflow-engine/dsl/policy-checks'; import { Effect } from 'effect'; -import { durationLimitRejections } from '../dsl/duration-limits.ts'; -import { mostValueDepth, type JsonObject } from '../dsl/json.ts'; -import { nestingRejections } from '../dsl/nesting.ts'; -import type { Rejection } from '../dsl/policy-checks.ts'; -import { rejectionsOf } from '../dsl/policy.ts'; import { defaultMostDuration } from '../interpreter/workflow-run.ts'; import { dslProblems, type Problem } from './dsl-validation.ts'; diff --git a/primitives/orchestration/src/document/workflow-summary.ts b/primitives/orchestration/src/document/workflow-summary.ts index 38f7e1f10..d57f53526 100644 --- a/primitives/orchestration/src/document/workflow-summary.ts +++ b/primitives/orchestration/src/document/workflow-summary.ts @@ -1,6 +1,5 @@ import type { SpecSummary } from '@beonauto/specs'; - -import { objectField, textField, type JsonObject } from '../dsl/json.ts'; +import { objectField, textField, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; export function summaryOf(document: JsonObject): SpecSummary { const header = objectField(document, 'document') ?? {}; diff --git a/primitives/orchestration/src/events/send-execution-event.ts b/primitives/orchestration/src/events/send-execution-event.ts index 0974d64ec..abaa68381 100644 --- a/primitives/orchestration/src/events/send-execution-event.ts +++ b/primitives/orchestration/src/events/send-execution-event.ts @@ -2,9 +2,9 @@ import { randomUUID } from 'node:crypto'; import { BrainContext, defineCommand, NotFound, quoted } from '@beonauto/operations'; import { getExecution } from '@beonauto/specs'; +import { jsonBytesOf } from '@beonauto/workflow-engine/dsl/json'; import { DateTime, Effect, Schema, SchemaTransformation } from 'effect'; -import { jsonBytesOf } from '../dsl/json.ts'; import { mostReceivedEventBytes, mostReceivedEvents, diff --git a/primitives/orchestration/src/interpreter/call-task.ts b/primitives/orchestration/src/interpreter/call-task.ts index a4c360568..ea18caace 100644 --- a/primitives/orchestration/src/interpreter/call-task.ts +++ b/primitives/orchestration/src/interpreter/call-task.ts @@ -1,5 +1,6 @@ -import { field, isObject, jsonBytesOf, textField, type Json } from '../dsl/json.ts'; -import { executeSpecFunction } from '../dsl/task-policy.ts'; +import { field, isObject, jsonBytesOf, textField, type Json } from '@beonauto/workflow-engine/dsl/json'; +import { executeSpecFunction } from '@beonauto/workflow-engine/dsl/task-policy'; + import { evaluateTemplate, placeOf } from './evaluation.ts'; import type { SpecCall, SpecCallResult } from './host.ts'; import type { Body, Invocation } from './invocation.ts'; diff --git a/primitives/orchestration/src/interpreter/error-tasks.ts b/primitives/orchestration/src/interpreter/error-tasks.ts index 0014b3b59..8946f52a0 100644 --- a/primitives/orchestration/src/interpreter/error-tasks.ts +++ b/primitives/orchestration/src/interpreter/error-tasks.ts @@ -1,5 +1,14 @@ -import type { Variables } from '../dsl/expressions.ts'; -import { entriesOf, field, isObject, objectField, textField, type JsonEntry, type JsonObject } from '../dsl/json.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; +import { + entriesOf, + field, + isObject, + objectField, + textField, + type JsonEntry, + type JsonObject, +} from '@beonauto/workflow-engine/dsl/json'; + import { evaluateTemplate, holds, placeOf } from './evaluation.ts'; import { bodyOf, type Body, type Invocation } from './invocation.ts'; import { RaisedError, errorAsJson, errorFromJson, raised, type DslError } from './raised-error.ts'; diff --git a/primitives/orchestration/src/interpreter/evaluation.ts b/primitives/orchestration/src/interpreter/evaluation.ts index 0dbe76a49..632c89534 100644 --- a/primitives/orchestration/src/interpreter/evaluation.ts +++ b/primitives/orchestration/src/interpreter/evaluation.ts @@ -1,6 +1,20 @@ -import { readDuration } from '../dsl/durations.ts'; -import { enclosedBody, expressionSource, runExpression, type Variables } from '../dsl/expressions.ts'; -import { entriesOf, isList, isObject, isTruthy, type Json, type JsonEntry, type JsonObject } from '../dsl/json.ts'; +import { readDuration } from '@beonauto/workflow-engine/dsl/durations'; +import { + enclosedBody, + expressionSource, + runExpression, + type Variables, +} from '@beonauto/workflow-engine/dsl/expressions'; +import { + entriesOf, + isList, + isObject, + isTruthy, + type Json, + type JsonEntry, + type JsonObject, +} from '@beonauto/workflow-engine/dsl/json'; + import type { Invocation, Place } from './invocation.ts'; import { RaisedError, errorType, raised } from './raised-error.ts'; import { mostActivationWork, mostExpressionWork, type RunState } from './run-state.ts'; diff --git a/primitives/orchestration/src/interpreter/flow-tasks.ts b/primitives/orchestration/src/interpreter/flow-tasks.ts index cda43086b..219f83a26 100644 --- a/primitives/orchestration/src/interpreter/flow-tasks.ts +++ b/primitives/orchestration/src/interpreter/flow-tasks.ts @@ -1,4 +1,4 @@ -import type { Variables } from '../dsl/expressions.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; import { entriesOf, field, @@ -10,8 +10,9 @@ import { type Json, type JsonArray, type JsonObject, -} from '../dsl/json.ts'; -import { taskEntries } from '../dsl/tasks.ts'; +} from '@beonauto/workflow-engine/dsl/json'; +import { taskEntries } from '@beonauto/workflow-engine/dsl/tasks'; + import { evaluateExpression, holds, placeOf } from './evaluation.ts'; import type { Release } from './holding.ts'; import { bodyOf, type Body, type Invocation, type TaskOutcome } from './invocation.ts'; diff --git a/primitives/orchestration/src/interpreter/held-data.test.ts b/primitives/orchestration/src/interpreter/held-data.test.ts index b06a676ac..ef8b1ab19 100644 --- a/primitives/orchestration/src/interpreter/held-data.test.ts +++ b/primitives/orchestration/src/interpreter/held-data.test.ts @@ -1,11 +1,11 @@ import { describe, expect, it } from 'vitest'; -import { retainedBytesOf } from '../dsl/retained-size.ts'; import { fakeHost } from '../testing/fake-host.ts'; import { interpret, runOf, workflow } from '../testing/workflows.ts'; import { mostHeldBytes, taskFrameBytes } from './holding.ts'; import type { RunSettlement } from './host.ts'; import { RaisedError } from './raised-error.ts'; +import { retainedBytesOf } from './retained-size.ts'; import { makeRunState } from './run-state.ts'; function holderOf() { diff --git a/primitives/orchestration/src/interpreter/holding.ts b/primitives/orchestration/src/interpreter/holding.ts index 4e2362410..ead611575 100644 --- a/primitives/orchestration/src/interpreter/holding.ts +++ b/primitives/orchestration/src/interpreter/holding.ts @@ -1,6 +1,7 @@ -import type { Json } from '../dsl/json.ts'; -import { retainedBytesOf } from '../dsl/retained-size.ts'; +import type { Json } from '@beonauto/workflow-engine/dsl/json'; + import { raised } from './raised-error.ts'; +import { retainedBytesOf } from './retained-size.ts'; export type Release = () => void; diff --git a/primitives/orchestration/src/interpreter/host.ts b/primitives/orchestration/src/interpreter/host.ts index 87cfb78d2..17f401b1f 100644 --- a/primitives/orchestration/src/interpreter/host.ts +++ b/primitives/orchestration/src/interpreter/host.ts @@ -1,6 +1,5 @@ import type { CallerIdentity, Settlement } from '@beonauto/operations'; - -import type { Json } from '../dsl/json.ts'; +import type { Json } from '@beonauto/workflow-engine/dsl/json'; export interface SpecCall { readonly org: string; diff --git a/primitives/orchestration/src/interpreter/inbox.ts b/primitives/orchestration/src/interpreter/inbox.ts index e921dc40f..b0f0925c7 100644 --- a/primitives/orchestration/src/interpreter/inbox.ts +++ b/primitives/orchestration/src/interpreter/inbox.ts @@ -1,6 +1,7 @@ -import { isJson, isObject, jsonBytesOf, type JsonObject } from '../dsl/json.ts'; -import { retainedBytesOf } from '../dsl/retained-size.ts'; +import { isJson, isObject, jsonBytesOf, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; + import { raised, type RaisedError } from './raised-error.ts'; +import { retainedBytesOf } from './retained-size.ts'; export type EventFilter = (event: JsonObject) => boolean; diff --git a/primitives/orchestration/src/interpreter/interpreter.ts b/primitives/orchestration/src/interpreter/interpreter.ts index 28459dba1..f0b9e8285 100644 --- a/primitives/orchestration/src/interpreter/interpreter.ts +++ b/primitives/orchestration/src/interpreter/interpreter.ts @@ -1,5 +1,6 @@ -import { field, jsonBytesOf, objectField, type Json } from '../dsl/json.ts'; -import { rejectionsOf } from '../dsl/policy.ts'; +import { field, jsonBytesOf, objectField, type Json } from '@beonauto/workflow-engine/dsl/json'; +import { rejectionsOf } from '@beonauto/workflow-engine/dsl/policy'; + import { placeIn, transform } from './evaluation.ts'; import type { WorkflowHost } from './host.ts'; import { RaisedError, errorType } from './raised-error.ts'; diff --git a/primitives/orchestration/src/interpreter/invocation.ts b/primitives/orchestration/src/interpreter/invocation.ts index 72abc4128..fb0d0123c 100644 --- a/primitives/orchestration/src/interpreter/invocation.ts +++ b/primitives/orchestration/src/interpreter/invocation.ts @@ -1,6 +1,7 @@ -import type { Variables } from '../dsl/expressions.ts'; -import type { Json } from '../dsl/json.ts'; -import type { TaskEntry } from '../dsl/tasks.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; +import type { Json } from '@beonauto/workflow-engine/dsl/json'; +import type { TaskEntry } from '@beonauto/workflow-engine/dsl/tasks'; + import type { Meter, RunState } from './run-state.ts'; export interface Scope { diff --git a/primitives/orchestration/src/interpreter/listen-task.ts b/primitives/orchestration/src/interpreter/listen-task.ts index d0be5cd6e..421b663a0 100644 --- a/primitives/orchestration/src/interpreter/listen-task.ts +++ b/primitives/orchestration/src/interpreter/listen-task.ts @@ -1,4 +1,4 @@ -import { enclosedBody } from '../dsl/expressions.ts'; +import { enclosedBody } from '@beonauto/workflow-engine/dsl/expressions'; import { entriesOf, field, @@ -9,9 +9,10 @@ import { type Json, type JsonEntry, type JsonObject, -} from '../dsl/json.ts'; -import type { Located } from '../dsl/policy-checks.ts'; -import { eventFiltersOf } from '../dsl/task-policy.ts'; +} from '@beonauto/workflow-engine/dsl/json'; +import type { Located } from '@beonauto/workflow-engine/dsl/policy-checks'; +import { eventFiltersOf } from '@beonauto/workflow-engine/dsl/task-policy'; + import { evaluate, placeOf } from './evaluation.ts'; import type { EventFilter } from './inbox.ts'; import type { Body, Invocation } from './invocation.ts'; diff --git a/primitives/orchestration/src/interpreter/raised-error.ts b/primitives/orchestration/src/interpreter/raised-error.ts index 50a6a3abe..a823e321c 100644 --- a/primitives/orchestration/src/interpreter/raised-error.ts +++ b/primitives/orchestration/src/interpreter/raised-error.ts @@ -1,4 +1,4 @@ -import { field, textField, type JsonObject } from '../dsl/json.ts'; +import { field, textField, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; export type ErrorKind = | 'configuration' diff --git a/primitives/orchestration/src/interpreter/refused-documents.test.ts b/primitives/orchestration/src/interpreter/refused-documents.test.ts new file mode 100644 index 000000000..14d81fed2 --- /dev/null +++ b/primitives/orchestration/src/interpreter/refused-documents.test.ts @@ -0,0 +1,31 @@ +import { describe, expect, it } from 'vitest'; + +import { interpret, workflow } from '../testing/workflows.ts'; + +function nestedDo(levels: number): string { + return levels === 1 ? '[{ leaf: { set: { done: true } } }]' : `[{ inner: { do: ${nestedDo(levels - 1)} } }]`; +} + +function branches(count: number): string { + return Array.from({ length: count }, (_, index) => `{ b${index}: { set: { n: ${index} } } }`).join(', '); +} + +describe('a workflow that starts with tasks nested too deeply', () => { + it('is rejected before it runs any of them, however it was stored', async () => { + const { settlement, commands } = await interpret(workflow(`do: ${nestedDo(65)}`)); + + expect(settlement).toMatchObject({ status: 'rejected', reason: 'invalid_input' }); + expect(commands.map(({ kind }) => kind)).toStrictEqual(['deadline', 'settle']); + }); +}); + +describe('a workflow that starts with a fork of more than 32 branches', () => { + it('is rejected before it runs any of them', async () => { + const { settlement, commands } = await interpret( + workflow(`do:\n - wide: { fork: { branches: [${branches(33)}] } }`), + ); + + expect(settlement).toMatchObject({ status: 'rejected', reason: 'invalid_input' }); + expect(commands.map(({ kind }) => kind)).toStrictEqual(['deadline', 'settle']); + }); +}); diff --git a/primitives/orchestration/src/dsl/retained-size.test.ts b/primitives/orchestration/src/interpreter/retained-size.test.ts similarity index 100% rename from primitives/orchestration/src/dsl/retained-size.test.ts rename to primitives/orchestration/src/interpreter/retained-size.test.ts diff --git a/primitives/orchestration/src/dsl/retained-size.ts b/primitives/orchestration/src/interpreter/retained-size.ts similarity index 100% rename from primitives/orchestration/src/dsl/retained-size.ts rename to primitives/orchestration/src/interpreter/retained-size.ts diff --git a/primitives/orchestration/src/interpreter/retry-policy.ts b/primitives/orchestration/src/interpreter/retry-policy.ts index 20f46e6ed..0efe48231 100644 --- a/primitives/orchestration/src/interpreter/retry-policy.ts +++ b/primitives/orchestration/src/interpreter/retry-policy.ts @@ -1,5 +1,6 @@ -import type { Variables } from '../dsl/expressions.ts'; -import { field, isObject, objectField, type Json, type JsonObject } from '../dsl/json.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; +import { field, isObject, objectField, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; + import { holds, millisecondsOf } from './evaluation.ts'; import type { Place } from './invocation.ts'; import { raised } from './raised-error.ts'; diff --git a/primitives/orchestration/src/interpreter/run-state.ts b/primitives/orchestration/src/interpreter/run-state.ts index d06e2cdee..68f77c32e 100644 --- a/primitives/orchestration/src/interpreter/run-state.ts +++ b/primitives/orchestration/src/interpreter/run-state.ts @@ -1,5 +1,6 @@ -import { measureOf, mostValueDepth, objectField, type Json, type JsonObject } from '../dsl/json.ts'; -import type { Components } from '../dsl/policy-checks.ts'; +import { measureOf, mostValueDepth, objectField, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; +import type { Components } from '@beonauto/workflow-engine/dsl/policy-checks'; + import { makeHolding, type Hold } from './holding.ts'; import type { WorkflowHost } from './host.ts'; import { makeInbox, type Inbox } from './inbox.ts'; diff --git a/primitives/orchestration/src/interpreter/settlement.ts b/primitives/orchestration/src/interpreter/settlement.ts index b2be6b9da..04844074e 100644 --- a/primitives/orchestration/src/interpreter/settlement.ts +++ b/primitives/orchestration/src/interpreter/settlement.ts @@ -1,4 +1,5 @@ -import type { Json } from '../dsl/json.ts'; +import type { Json } from '@beonauto/workflow-engine/dsl/json'; + import type { RunSettlement } from './host.ts'; import { describeError, type DslError } from './raised-error.ts'; diff --git a/primitives/orchestration/src/interpreter/task-bodies.ts b/primitives/orchestration/src/interpreter/task-bodies.ts index b1c52ed22..a9f98d909 100644 --- a/primitives/orchestration/src/interpreter/task-bodies.ts +++ b/primitives/orchestration/src/interpreter/task-bodies.ts @@ -1,4 +1,5 @@ -import type { TaskKind } from '../dsl/tasks.ts'; +import type { TaskKind } from '@beonauto/workflow-engine/dsl/tasks'; + import { callTask } from './call-task.ts'; import { raiseTask, tryTask } from './error-tasks.ts'; import { doTask, forkTask, forTask, switchTask } from './flow-tasks.ts'; diff --git a/primitives/orchestration/src/interpreter/task-runner.ts b/primitives/orchestration/src/interpreter/task-runner.ts index 9779c6af2..7d21437ea 100644 --- a/primitives/orchestration/src/interpreter/task-runner.ts +++ b/primitives/orchestration/src/interpreter/task-runner.ts @@ -1,6 +1,7 @@ -import type { Variables } from '../dsl/expressions.ts'; -import { field, objectField, textField, type Json, type JsonObject } from '../dsl/json.ts'; -import { taskEntries, typeOf, type TaskEntry } from '../dsl/tasks.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; +import { field, objectField, textField, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; +import { taskEntries, typeOf, type TaskEntry } from '@beonauto/workflow-engine/dsl/tasks'; + import { holds, placeIn, transform } from './evaluation.ts'; import type { ListResult, Place, Runner, Scope, TaskOutcome } from './invocation.ts'; import { raised } from './raised-error.ts'; diff --git a/primitives/orchestration/src/interpreter/timeouts.ts b/primitives/orchestration/src/interpreter/timeouts.ts index 8b6aa790f..5be138cba 100644 --- a/primitives/orchestration/src/interpreter/timeouts.ts +++ b/primitives/orchestration/src/interpreter/timeouts.ts @@ -1,5 +1,6 @@ -import type { Variables } from '../dsl/expressions.ts'; -import { field, isObject, objectField, type Json, type JsonObject } from '../dsl/json.ts'; +import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; +import { field, isObject, objectField, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; + import { millisecondsOf, placeIn } from './evaluation.ts'; import { raised } from './raised-error.ts'; import type { RunState } from './run-state.ts'; diff --git a/primitives/orchestration/src/interpreter/workflow-run.ts b/primitives/orchestration/src/interpreter/workflow-run.ts index 494223897..efe752a8d 100644 --- a/primitives/orchestration/src/interpreter/workflow-run.ts +++ b/primitives/orchestration/src/interpreter/workflow-run.ts @@ -1,6 +1,5 @@ import type { CallerIdentity } from '@beonauto/operations'; - -import { isJson, isList, isObject, type Json, type JsonObject } from '../dsl/json.ts'; +import { isJson, isList, isObject, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; interface RunExecution { readonly id: string; diff --git a/primitives/orchestration/src/primitive/execution-input.test.ts b/primitives/orchestration/src/primitive/execution-input.test.ts index 8e6b7b316..caac5b7b7 100644 --- a/primitives/orchestration/src/primitive/execution-input.test.ts +++ b/primitives/orchestration/src/primitive/execution-input.test.ts @@ -1,7 +1,7 @@ +import type { Json } from '@beonauto/workflow-engine/dsl/json'; import { Effect } from 'effect'; import { describe, expect, it } from 'vitest'; -import type { Json } from '../dsl/json.ts'; import { brainWith } from '../testing/brain.ts'; import type { OrchestrationClient } from './orchestration-client.ts'; import { makeOrchestration } from './orchestration-primitive.ts'; diff --git a/primitives/orchestration/src/primitive/orchestration-client.ts b/primitives/orchestration/src/primitive/orchestration-client.ts index 06e5685ea..88f995e2a 100644 --- a/primitives/orchestration/src/primitive/orchestration-client.ts +++ b/primitives/orchestration/src/primitive/orchestration-client.ts @@ -1,10 +1,10 @@ import { setTimeout } from 'node:timers/promises'; import { NotFound, Unavailable } from '@beonauto/operations'; +import type { JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { Client, Connection, WorkflowNotFoundError } from '@temporalio/client'; import { Clock, Effect, type Scope } from 'effect'; -import type { JsonObject } from '../dsl/json.ts'; import { defaultLongestNestedExecutionMs, type StartingRun, type WorkflowRun } from '../interpreter/workflow-run.ts'; import { makeThrottle } from '../notices/throttle.ts'; import { connectionOptionsOf, type TemporalSettings } from '../worker/temporal-settings.ts'; diff --git a/primitives/orchestration/src/primitive/orchestration-primitive.ts b/primitives/orchestration/src/primitive/orchestration-primitive.ts index 674d2ec4e..f6bb84ebe 100644 --- a/primitives/orchestration/src/primitive/orchestration-primitive.ts +++ b/primitives/orchestration/src/primitive/orchestration-primitive.ts @@ -1,10 +1,10 @@ import { asSentence, InvalidInput } from '@beonauto/operations'; import { definePrimitive, inWords, type FinishesLater, type Primitive } from '@beonauto/specs'; +import { measureOf, mostValueDepth, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { Effect, type Schema } from 'effect'; import { parseWorkflowDocument } from '../document/workflow-document.ts'; import { summaryOf } from '../document/workflow-summary.ts'; -import { measureOf, mostValueDepth, type JsonObject } from '../dsl/json.ts'; import type { OrchestrationClient } from './orchestration-client.ts'; import { orchestrationDescription } from './orchestration-description.ts'; diff --git a/primitives/orchestration/src/testing/workflows.ts b/primitives/orchestration/src/testing/workflows.ts index 137b76b83..cdb2ca8e5 100644 --- a/primitives/orchestration/src/testing/workflows.ts +++ b/primitives/orchestration/src/testing/workflows.ts @@ -1,7 +1,7 @@ import type { CallerIdentity } from '@beonauto/operations'; +import { isJson, isObject, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { parse } from 'yaml'; -import { isJson, isObject, type Json, type JsonObject } from '../dsl/json.ts'; import type { RunSettlement, SpecCall, SpecCallResult, WorkflowHost } from '../interpreter/host.ts'; import { startWorkflow, type WorkflowStart } from '../interpreter/interpreter.ts'; import type { WorkflowEnding } from '../interpreter/settlement.ts'; diff --git a/primitives/orchestration/src/worker/activities.ts b/primitives/orchestration/src/worker/activities.ts index d532eed5a..69aa9846d 100644 --- a/primitives/orchestration/src/worker/activities.ts +++ b/primitives/orchestration/src/worker/activities.ts @@ -1,10 +1,10 @@ import { Conflict, NotFound } from '@beonauto/operations'; import type { SettleExecution, Settlement } from '@beonauto/specs'; +import { jsonBytesOf } from '@beonauto/workflow-engine/dsl/json'; import { ApplicationFailure } from '@temporalio/activity'; import { ApplicationFailureCategory } from '@temporalio/common'; import { Cause, Effect, Exit } from 'effect'; -import { jsonBytesOf } from '../dsl/json.ts'; import type { RunSettlement, SettleRequest, SpecCall, SpecCallResult } from '../interpreter/host.ts'; import { executionConflict, diff --git a/primitives/orchestration/src/worker/command-paths.test.ts b/primitives/orchestration/src/worker/command-paths.test.ts index 1f024b051..07f6a7c47 100644 --- a/primitives/orchestration/src/worker/command-paths.test.ts +++ b/primitives/orchestration/src/worker/command-paths.test.ts @@ -1,8 +1,8 @@ +import type { JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { Effect } from 'effect'; import { afterAll, beforeAll, describe, expect, it } from 'vitest'; import { recordHistory } from '../../replay-corpus.ts'; -import type { JsonObject } from '../dsl/json.ts'; import { temporalHarness, type TemporalHarness } from '../testing/temporal.ts'; import { runFor, workflow } from '../testing/workflows.ts'; diff --git a/primitives/orchestration/src/worker/orchestration-worker.test.ts b/primitives/orchestration/src/worker/orchestration-worker.test.ts index 24cf0b4fb..a8b34e665 100644 --- a/primitives/orchestration/src/worker/orchestration-worker.test.ts +++ b/primitives/orchestration/src/worker/orchestration-worker.test.ts @@ -1,8 +1,8 @@ +import type { JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { Effect } from 'effect'; import { afterAll, beforeAll, describe, expect, it } from 'vitest'; import { recordHistory } from '../../replay-corpus.ts'; -import type { JsonObject } from '../dsl/json.ts'; import { temporalHarness, type TemporalHarness } from '../testing/temporal.ts'; import { runFor, workflow } from '../testing/workflows.ts'; import type { SpecExecution, SpecExecutionResult } from './dependencies.ts'; diff --git a/primitives/orchestration/src/worker/spec-results.ts b/primitives/orchestration/src/worker/spec-results.ts index e206bfebc..d6979102a 100644 --- a/primitives/orchestration/src/worker/spec-results.ts +++ b/primitives/orchestration/src/worker/spec-results.ts @@ -1,6 +1,6 @@ import type { Outcome } from '@beonauto/operations'; +import { field, textField } from '@beonauto/workflow-engine/dsl/json'; -import { field, textField } from '../dsl/json.ts'; import type { SpecExecutionResult } from './dependencies.ts'; export function specExecutionResultOf(outcome: Outcome): SpecExecutionResult { diff --git a/primitives/orchestration/src/worker/temporal-settings.ts b/primitives/orchestration/src/worker/temporal-settings.ts index 6156c0174..eaab1058c 100644 --- a/primitives/orchestration/src/worker/temporal-settings.ts +++ b/primitives/orchestration/src/worker/temporal-settings.ts @@ -1,7 +1,6 @@ +import { readDuration } from '@beonauto/workflow-engine/dsl/durations'; import { Config, ConfigProvider, Data, Effect, Option } from 'effect'; -import { readDuration } from '../dsl/durations.ts'; - export interface TemporalSettings { readonly address: string; readonly namespace: string; diff --git a/primitives/orchestration/src/worker/workflow-bundle.test.ts b/primitives/orchestration/src/worker/workflow-bundle.test.ts index 28ff994cb..47b170541 100644 --- a/primitives/orchestration/src/worker/workflow-bundle.test.ts +++ b/primitives/orchestration/src/worker/workflow-bundle.test.ts @@ -43,6 +43,7 @@ describe('a workflow bundle built ahead of time', () => { expect(inputs).toContain('pnpm-lock.yaml'); expect(inputs).toContain('patches/@gabrielbryk__jq-ts@1.7.0.patch'); expect(inputs).toContain('primitives/orchestration/src/workflow/workflows.ts'); + expect(inputs).toContain('packages/workflow-engine/src/dsl/expressions.ts'); expect(inputs.filter((file) => /\.test\.ts$|\/testing\//u.test(file))).toEqual([]); expect(Object.values(manifest.inputs).filter((digest) => !/^[0-9a-f]{64}$/u.test(digest))).toEqual([]); }); @@ -86,7 +87,7 @@ describe('a workflow bundle that does not match the server', () => { it('names at most five of the files that differ', async () => { const copy = await copiedBundle('many'); await changedFile(join(copy, 'workflow-bundle.json'), (text) => - text.replaceAll(/("primitives\/orchestration\/src\/dsl\/[^"]+": ")[0-9a-f]{64}/gu, '$1changed'), + text.replaceAll(/("packages\/workflow-engine\/src\/dsl\/[^"]+": ")[0-9a-f]{64}/gu, '$1changed'), ); await expect(verifiedWorkflowBundle(copy)).rejects.toThrow(/this server runs: [^,]+(?:, [^,]+){4}$/u); diff --git a/primitives/orchestration/src/worker/workflow-bundle.ts b/primitives/orchestration/src/worker/workflow-bundle.ts index 6a3c18ca0..ae4a09a10 100644 --- a/primitives/orchestration/src/worker/workflow-bundle.ts +++ b/primitives/orchestration/src/worker/workflow-bundle.ts @@ -18,7 +18,7 @@ const manifestFile = 'workflow-bundle.json'; const mostChangesNamed = 5; -const sourceDirectory = 'primitives/orchestration/src'; +const sourceDirectories = ['packages/workflow-engine/src', 'primitives/orchestration/src']; const decodeManifest = Schema.decodeUnknownOption( Schema.fromJsonString(Schema.Struct({ code: Schema.String, inputs: Schema.Record(Schema.String, Schema.String) })), @@ -64,12 +64,13 @@ export function requireWorkflowBundler(): void { async function bundleInputs(): Promise { const patches = await readdir(join(workspaceRoot, 'patches')); - const sources = await readdir(join(workspaceRoot, sourceDirectory), { recursive: true }); - return [ - 'pnpm-lock.yaml', - ...patches.map((patch) => `patches/${patch}`), - ...sources.filter((file) => isWorkflowSource(file)).map((file) => `${sourceDirectory}/${file}`), - ].toSorted(); + const sources = await Promise.all(sourceDirectories.map((directory) => workflowSourcesIn(directory))); + return ['pnpm-lock.yaml', ...patches.map((patch) => `patches/${patch}`), ...sources.flat()].toSorted(); +} + +async function workflowSourcesIn(directory: string): Promise { + const files = await readdir(join(workspaceRoot, directory), { recursive: true }); + return files.filter((file) => isWorkflowSource(file)).map((file) => `${directory}/${file}`); } function isWorkflowSource(file: string): boolean { diff --git a/primitives/orchestration/src/workflow/interpreter-workflow.ts b/primitives/orchestration/src/workflow/interpreter-workflow.ts index 7fcd376d4..473cc5d09 100644 --- a/primitives/orchestration/src/workflow/interpreter-workflow.ts +++ b/primitives/orchestration/src/workflow/interpreter-workflow.ts @@ -1,4 +1,5 @@ -import type { Json } from '../dsl/json.ts'; +import type { Json } from '@beonauto/workflow-engine/dsl/json'; + import { startWorkflow } from '../interpreter/interpreter.ts'; import { eventSignalName } from './activity-contract.ts'; import { temporalHost, type WorkflowApi } from './temporal-host.ts'; From 01afc204c563ecb8e8a6986d6af59b4b8cf22985 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:49:43 +0100 Subject: [PATCH 12/23] refactor(workflow-engine): take the functions a workflow calls from the caller The DSL policy knew execute_spec, its arguments, the orchestration primitive and the specs of a brain. policyOf(functions) now takes the functions a workflow may call, each with the checks of its arguments, and the words that explain where a workflow reaches the world and how it starts. The orchestration primitive gives execute_spec and its rules in src/document/workflow-functions.ts, with their tests, so the engine package knows no brains, specs or primitives. The messages a document gets are the same as before. Co-Authored-By: Claude Opus 5.5 --- .../workflow-engine/src/dsl/call-functions.ts | 20 ++++++ .../workflow-engine/src/dsl/nesting.test.ts | 6 +- .../workflow-engine/src/dsl/policy.test.ts | 6 +- packages/workflow-engine/src/dsl/policy.ts | 64 ++++++++++------- .../src/dsl/task-checks.test.ts | 6 +- .../src/dsl/task-policy.test.ts | 52 +++++--------- .../workflow-engine/src/dsl/task-policy.ts | 49 +++++-------- .../src/dsl/workflow-bounds.test.ts | 6 +- .../workflow-engine/src/testing/workflows.ts | 15 +++- primitives/orchestration/README.md | 2 +- .../src/document/workflow-document.ts | 4 +- .../src/document/workflow-functions.test.ts | 71 +++++++++++++++++++ .../src/document/workflow-functions.ts | 34 +++++++++ .../src/interpreter/call-task.ts | 2 +- .../src/interpreter/interpreter.ts | 4 +- 15 files changed, 235 insertions(+), 106 deletions(-) create mode 100644 packages/workflow-engine/src/dsl/call-functions.ts create mode 100644 primitives/orchestration/src/document/workflow-functions.test.ts create mode 100644 primitives/orchestration/src/document/workflow-functions.ts diff --git a/packages/workflow-engine/src/dsl/call-functions.ts b/packages/workflow-engine/src/dsl/call-functions.ts new file mode 100644 index 000000000..7cea70fd6 --- /dev/null +++ b/packages/workflow-engine/src/dsl/call-functions.ts @@ -0,0 +1,20 @@ +import type { Json } from './json.ts'; +import type { Rejection } from './policy-checks.ts'; + +export type ArgumentRejections = (arguments_: Json | undefined, pointer: string) => readonly Rejection[]; + +export interface CallFunctions { + readonly argumentChecks: Readonly>; + readonly howAWorkflowReachesTheWorld: string; + readonly howAWorkflowStarts: string; +} + +export function namesOf({ argumentChecks }: CallFunctions): string { + return Object.keys(argumentChecks).join(', '); +} + +export function theFunctions(functions: CallFunctions): string { + return Object.keys(functions.argumentChecks).length === 1 + ? `the one function is ${namesOf(functions)}` + : `the functions are ${namesOf(functions)}`; +} diff --git a/packages/workflow-engine/src/dsl/nesting.test.ts b/packages/workflow-engine/src/dsl/nesting.test.ts index bddf522c6..5d73f762d 100644 --- a/packages/workflow-engine/src/dsl/nesting.test.ts +++ b/packages/workflow-engine/src/dsl/nesting.test.ts @@ -1,8 +1,10 @@ import { describe, expect, it } from 'vitest'; -import { workflow } from '../testing/workflows.ts'; +import { testFunctions, workflow } from '../testing/workflows.ts'; import type { Json } from './json.ts'; -import { rejectionsOf } from './policy.ts'; +import { policyOf } from './policy.ts'; + +const rejectionsOf = policyOf(testFunctions); function nestedDo(levels: number): string { return levels === 1 ? '[{ leaf: { set: { done: true } } }]' : `[{ inner: { do: ${nestedDo(levels - 1)} } }]`; diff --git a/packages/workflow-engine/src/dsl/policy.test.ts b/packages/workflow-engine/src/dsl/policy.test.ts index 931304a63..39a723aaa 100644 --- a/packages/workflow-engine/src/dsl/policy.test.ts +++ b/packages/workflow-engine/src/dsl/policy.test.ts @@ -1,7 +1,9 @@ import { describe, expect, it } from 'vitest'; -import { header, workflow, yamlObject } from '../testing/workflows.ts'; -import { rejectionsOf } from './policy.ts'; +import { header, testFunctions, workflow, yamlObject } from '../testing/workflows.ts'; +import { policyOf } from './policy.ts'; + +const rejectionsOf = policyOf(testFunctions); function pointersRejectedIn(source: string): readonly string[] { return rejectionsOf(workflow(source)).map(({ pointer }) => pointer); diff --git a/packages/workflow-engine/src/dsl/policy.ts b/packages/workflow-engine/src/dsl/policy.ts index e5c4e9ed2..c8fb54a4d 100644 --- a/packages/workflow-engine/src/dsl/policy.ts +++ b/packages/workflow-engine/src/dsl/policy.ts @@ -1,3 +1,4 @@ +import { namesOf, type CallFunctions } from './call-functions.ts'; import { entriesOf, field, @@ -24,22 +25,28 @@ import { import { commonRejections, ownRejections } from './task-policy.ts'; import { kindOf, pointerTo, taskEntries, type TaskEntry } from './tasks.ts'; -const rejectedComponents: Readonly> = { - authentications: 'authentications are not supported in this version: a workflow reaches nothing that needs them', - secrets: 'secrets are not supported in this version: a workflow reaches nothing that needs them', - catalogs: 'catalogs are not supported in this version: a workflow calls only execute_spec', - extensions: 'extensions are not supported in this version', - functions: 'reusable functions are not supported in this version: call execute_spec directly', -}; +type Policy = (document: JsonObject) => readonly Rejection[]; + +function rejectedComponents(functions: CallFunctions): Readonly> { + return { + authentications: 'authentications are not supported in this version: a workflow reaches nothing that needs them', + secrets: 'secrets are not supported in this version: a workflow reaches nothing that needs them', + catalogs: `catalogs are not supported in this version: a workflow calls only ${namesOf(functions)}`, + extensions: 'extensions are not supported in this version', + functions: `reusable functions are not supported in this version: call ${namesOf(functions)} directly`, + }; +} const flowDirectives = new Set(['continue', 'exit', 'end']); -export function rejectionsOf(document: JsonObject): readonly Rejection[] { - const nesting = nestingRejections(document); - return nesting.length > 0 ? nesting : walkedRejectionsOf(document); +export function policyOf(functions: CallFunctions): Policy { + return (document) => { + const nesting = nestingRejections(document); + return nesting.length > 0 ? nesting : walkedRejectionsOf(document, functions); + }; } -function walkedRejectionsOf(document: JsonObject): readonly Rejection[] { +function walkedRejectionsOf(document: JsonObject, functions: CallFunctions): readonly Rejection[] { const use = objectField(document, 'use') ?? {}; const components = { errors: objectField(use, 'errors') ?? {}, @@ -49,13 +56,13 @@ function walkedRejectionsOf(document: JsonObject): readonly Rejection[] { const schedule = field(document, 'schedule') === undefined ? [] - : [forbidden('/schedule', 'schedules are not supported in this version: execute the spec to run it')]; + : [forbidden('/schedule', `schedules are not supported in this version: ${functions.howAWorkflowStarts}`)]; return versionRejections(document).concat( - componentRejections(use, components), + componentRejections(use, components, functions), schedule, workflowDataRejections(document), timeoutRejections(field(document, 'timeout'), '/timeout', components), - taskListRejections(field(document, 'do'), '/do', components), + taskListRejections(field(document, 'do'), '/do', components, functions), ); } @@ -66,8 +73,8 @@ function versionRejections(document: JsonObject): readonly Rejection[] { : [forbidden('/document/dsl', `This runtime runs documents of DSL 1.0.x, not ${dsl ?? 'an unnamed version'}`)]; } -function componentRejections(use: JsonObject, components: Components): readonly Rejection[] { - const rejected = Object.entries(rejectedComponents) +function componentRejections(use: JsonObject, components: Components, functions: CallFunctions): readonly Rejection[] { + const rejected = Object.entries(rejectedComponents(functions)) .filter(([name]: readonly [string, string]) => field(use, name) !== undefined) .map(([name, detail]: readonly [string, string]) => forbidden(`/use/${name}`, detail)); const retries = entriesOf(components.retries).flatMap(([name, policy]: JsonEntry) => @@ -105,20 +112,25 @@ function schemaRejections(schema: JsonObject | undefined, pointer: string): read : [forbidden(`${pointer}/format`, `Schemas are JSON Schema in this version, not ${format}`)]; } -function taskListRejections(list: Json | undefined, pointer: string, components: Components): readonly Rejection[] { +function taskListRejections( + list: Json | undefined, + pointer: string, + components: Components, + functions: CallFunctions, +): readonly Rejection[] { const entries = taskEntries(list, pointer); const names = new Set(entries.map(({ name }) => name)); return entries.flatMap((entry) => - taskRejections(entry, components).concat( + taskRejections(entry, components, functions).concat( jumpRejections(field(entry.task, 'then'), `${entry.reference}/then`, names), ), ); } -function taskRejections(entry: TaskEntry, components: Components): readonly Rejection[] { - return ownRejections(entry, components).concat( +function taskRejections(entry: TaskEntry, components: Components, functions: CallFunctions): readonly Rejection[] { + return ownRejections(entry, components, functions).concat( commonRejections(entry, components), - nestedRejections(entry, components), + nestedRejections(entry, components, functions), ); } @@ -128,13 +140,17 @@ function jumpRejections(then: Json | undefined, pointer: string, siblings: Reado : [rejection(pointer, `then: ${then} names no task in the same list`)]; } -function nestedRejections({ task, reference }: TaskEntry, components: Components): readonly Rejection[] { +function nestedRejections( + { task, reference }: TaskEntry, + components: Components, + functions: CallFunctions, +): readonly Rejection[] { const kind = kindOf(task); if (kind === 'fork') { const branches = field(objectField(task, 'fork') ?? {}, 'branches'); const pointer = `${reference}/fork/branches`; return taskEntries(branches, pointer).flatMap((branch) => - taskRejections(branch, components).concat( + taskRejections(branch, components, functions).concat( branchJumpRejections(field(branch.task, 'then'), `${branch.reference}/then`), ), ); @@ -144,7 +160,7 @@ function nestedRejections({ task, reference }: TaskEntry, components: Components [kind === 'try' ? field(task, 'try') : undefined, `${reference}/try`], [kind === 'try' ? field(objectField(task, 'catch') ?? {}, 'do') : undefined, `${reference}/catch/do`], ]; - return lists.flatMap(([list, pointer]: Located) => taskListRejections(list, pointer, components)); + return lists.flatMap(([list, pointer]: Located) => taskListRejections(list, pointer, components, functions)); } function branchJumpRejections(then: Json | undefined, pointer: string): readonly Rejection[] { diff --git a/packages/workflow-engine/src/dsl/task-checks.test.ts b/packages/workflow-engine/src/dsl/task-checks.test.ts index 242b98df3..b058be7dd 100644 --- a/packages/workflow-engine/src/dsl/task-checks.test.ts +++ b/packages/workflow-engine/src/dsl/task-checks.test.ts @@ -1,7 +1,9 @@ import { describe, expect, it } from 'vitest'; -import { workflow } from '../testing/workflows.ts'; -import { rejectionsOf } from './policy.ts'; +import { testFunctions, workflow } from '../testing/workflows.ts'; +import { policyOf } from './policy.ts'; + +const rejectionsOf = policyOf(testFunctions); function pointersRejectedIn(tasks: string): readonly string[] { return rejectionsOf(workflow(`do:\n${tasks}`)).map(({ pointer }) => pointer); diff --git a/packages/workflow-engine/src/dsl/task-policy.test.ts b/packages/workflow-engine/src/dsl/task-policy.test.ts index 2fba88b42..4e1f70063 100644 --- a/packages/workflow-engine/src/dsl/task-policy.test.ts +++ b/packages/workflow-engine/src/dsl/task-policy.test.ts @@ -1,7 +1,9 @@ import { describe, expect, it } from 'vitest'; -import { workflow } from '../testing/workflows.ts'; -import { rejectionsOf } from './policy.ts'; +import { testFunctions, workflow } from '../testing/workflows.ts'; +import { policyOf } from './policy.ts'; + +const rejectionsOf = policyOf(testFunctions); function rejectedIn(tasks: string): readonly string[] { return rejectionsOf(workflow(`do:\n${tasks}`)).map(({ pointer, detail }) => `${pointer}: ${detail}`); @@ -24,51 +26,29 @@ describe('the policy of the tasks of a document', () => { it.each(['http', 'grpc', 'openapi', 'asyncapi', 'a2a', 'mcp'])('rejects call: %s', (name) => { expect(rejectedIn(` - fetch: { call: ${name}, with: {} }`)).toEqual([ - `/do/0/fetch/call: call: ${name} is not allowed: a workflow reaches the world only through the specs of its brain; call execute_spec`, + `/do/0/fetch/call: call: ${name} is not allowed: a workflow reaches the world only through the functions it is given; call notify`, ]); }); - it('rejects a call of a function other than execute_spec', () => { + it('rejects a call of a function other than those its caller gives', () => { + const twoFunctions = policyOf({ ...testFunctions, argumentChecks: { notify: () => [], page: () => [] } }); + expect(rejectedIn(' - log: { call: log, with: {} }')).toEqual([ - '/do/0/log/call: call: log names no function; the one function is execute_spec', + '/do/0/log/call: call: log names no function; the one function is notify', ]); - }); -}); - -describe('the policy of execute_spec', () => { - it('takes a primitive, a name and an input, as expressions or literals', () => { - expect( - rejectedIn(` - - summarize: - call: execute_spec - with: { primitive: inference, name: '\${ .spec }', input: { text: '\${ .text }' } } -`), - ).toEqual([]); - }); - - it('rejects arguments it does not take and arguments it lacks', () => { - expect( - rejectedIn(` - - nothing: { call: execute_spec } - - partial: { call: execute_spec, with: { name: 3, model: big } } -`), - ).toEqual([ - '/do/0/nothing/with: execute_spec takes with: { primitive, name, input }', - '/do/1/partial/with/model: execute_spec takes no argument model', - '/do/1/partial/with/primitive: execute_spec needs a string primitive', - '/do/1/partial/with/name: execute_spec needs a string name', + expect(twoFunctions(workflow('do:\n - log: { call: log }')).map(({ detail }) => detail)).toEqual([ + 'call: log names no function; the functions are notify, page', ]); }); - it('rejects executing another workflow, and broken expressions in its arguments', () => { + it('leaves the arguments of a call to the checks of its function', () => { expect( rejectedIn(` - - nested: { call: execute_spec, with: { primitive: orchestration, name: other, input: ['\${ .a + }'] } } + - fine: { call: notify, with: { to: '\${ .who }' } } + - bare: { call: notify } + - broken: { call: notify, with: { to: '\${ .a + }' } } `), - ).toEqual([ - '/do/0/nested/with/primitive: A workflow cannot execute another workflow in this version', - expect.stringMatching(/^\/do\/0\/nested\/with\/input\/0: /u), - ]); + ).toEqual(['/do/1/bare/with: notify takes with: { to }', expect.stringMatching(/^\/do\/2\/broken\/with\/to: /u)]); }); }); diff --git a/packages/workflow-engine/src/dsl/task-policy.ts b/packages/workflow-engine/src/dsl/task-policy.ts index 9ea9f35ff..e2c8a0cd3 100644 --- a/packages/workflow-engine/src/dsl/task-policy.ts +++ b/packages/workflow-engine/src/dsl/task-policy.ts @@ -1,3 +1,4 @@ +import { namesOf, theFunctions, type CallFunctions } from './call-functions.ts'; import { entriesOf, field, @@ -27,20 +28,15 @@ import { kindOf, pointerTo, type TaskEntry, type TaskKind } from './tasks.ts'; type OwnRejections = (task: JsonObject, reference: string, components: Components) => readonly Rejection[]; -export const executeSpecFunction = 'execute_spec'; - const mostForkBranches = 32; const outboundCalls = new Set(['http', 'grpc', 'openapi', 'asyncapi', 'a2a', 'mcp']); -const executeSpecArguments = new Set(['primitive', 'name', 'input']); - -const rejectionsByKind: Readonly> = { +const rejectionsByKind: Readonly, OwnRejections>> = { run: (_task, reference) => [ forbidden(`${reference}/run`, 'run tasks (shell, script, container, workflow) are not allowed'), ], emit: (_task, reference) => [forbidden(`${reference}/emit`, 'emit is not supported in this version')], - call: (task, reference) => callRejections(task, reference), listen: (task, reference) => listenRejections(task, reference), raise: (task, reference, components) => raiseRejections(task, reference, components), wait: (task, reference) => durationRejections(field(task, 'wait'), `${reference}/wait`), @@ -56,10 +52,17 @@ const rejectionsByKind: Readonly> = { do: () => [], }; -export function ownRejections({ task, reference }: TaskEntry, components: Components): readonly Rejection[] { +export function ownRejections( + { task, reference }: TaskEntry, + components: Components, + functions: CallFunctions, +): readonly Rejection[] { const kind = kindOf(task); - return kind === undefined - ? [rejection(reference, 'The task has no type this runtime knows')] + if (kind === undefined) { + return [rejection(reference, 'The task has no type this runtime knows')]; + } + return kind === 'call' + ? callRejections(task, reference, functions) : rejectionsByKind[kind](task, reference, components); } @@ -90,36 +93,20 @@ function dataRejections(task: JsonObject, reference: string, part: string, trans return schema.concat(transformRejections(field(data, transform), `${reference}/${part}/${transform}`)); } -function callRejections(task: JsonObject, reference: string): readonly Rejection[] { +function callRejections(task: JsonObject, reference: string, functions: CallFunctions): readonly Rejection[] { const name = textField(task, 'call') ?? ''; if (outboundCalls.has(name)) { return [ forbidden( `${reference}/call`, - `call: ${name} is not allowed: a workflow reaches the world only through the specs of its brain; call ${executeSpecFunction}`, + `call: ${name} is not allowed: ${functions.howAWorkflowReachesTheWorld}; call ${namesOf(functions)}`, ), ]; } - return name === executeSpecFunction - ? executeSpecRejections(field(task, 'with'), `${reference}/with`) - : [forbidden(`${reference}/call`, `call: ${name} names no function; the one function is ${executeSpecFunction}`)]; -} - -function executeSpecRejections(arguments_: Json | undefined, pointer: string): readonly Rejection[] { - if (!isObject(arguments_)) { - return [rejection(pointer, `${executeSpecFunction} takes with: { primitive, name, input }`)]; - } - const unknown = Object.keys(arguments_) - .filter((key) => !executeSpecArguments.has(key)) - .map((key) => rejection(pointerTo(pointer, key), `${executeSpecFunction} takes no argument ${key}`)); - const missing = ['primitive', 'name'] - .filter((key) => typeof field(arguments_, key) !== 'string') - .map((key) => rejection(pointerTo(pointer, key), `${executeSpecFunction} needs a string ${key}`)); - const workflow = - field(arguments_, 'primitive') === 'orchestration' - ? [forbidden(`${pointer}/primitive`, 'A workflow cannot execute another workflow in this version')] - : []; - return unknown.concat(missing, workflow, templateRejections(arguments_, pointer)); + const argumentChecks = Object.hasOwn(functions.argumentChecks, name) ? functions.argumentChecks[name] : undefined; + return argumentChecks === undefined + ? [forbidden(`${reference}/call`, `call: ${name} names no function; ${theFunctions(functions)}`)] + : argumentChecks(field(task, 'with'), `${reference}/with`); } function listenRejections(task: JsonObject, reference: string): readonly Rejection[] { diff --git a/packages/workflow-engine/src/dsl/workflow-bounds.test.ts b/packages/workflow-engine/src/dsl/workflow-bounds.test.ts index 8b0181e14..219aca83c 100644 --- a/packages/workflow-engine/src/dsl/workflow-bounds.test.ts +++ b/packages/workflow-engine/src/dsl/workflow-bounds.test.ts @@ -1,8 +1,10 @@ import { describe, expect, it } from 'vitest'; -import { workflow } from '../testing/workflows.ts'; +import { testFunctions, workflow } from '../testing/workflows.ts'; import { durationLimitRejections } from './duration-limits.ts'; -import { rejectionsOf } from './policy.ts'; +import { policyOf } from './policy.ts'; + +const rejectionsOf = policyOf(testFunctions); const threeHours = 10_800_000; diff --git a/packages/workflow-engine/src/testing/workflows.ts b/packages/workflow-engine/src/testing/workflows.ts index 3b882730e..5e3356993 100644 --- a/packages/workflow-engine/src/testing/workflows.ts +++ b/packages/workflow-engine/src/testing/workflows.ts @@ -1,7 +1,9 @@ import { Schema } from 'effect'; import { parse } from 'yaml'; -import type { JsonObject } from '../dsl/json.ts'; +import type { CallFunctions } from '../dsl/call-functions.ts'; +import { isObject, type JsonObject } from '../dsl/json.ts'; +import { rejection, templateRejections } from '../dsl/policy-checks.ts'; export const header = { dsl: '1.0.3', namespace: 'acme', name: 'test', version: '1.0.0' }; @@ -11,6 +13,17 @@ export function yamlObject(source: string): JsonObject { return decodeObject(parse(source)); } +export const testFunctions: CallFunctions = { + argumentChecks: { + notify: (arguments_, pointer) => + isObject(arguments_) + ? templateRejections(arguments_, pointer) + : [rejection(pointer, 'notify takes with: { to }')], + }, + howAWorkflowReachesTheWorld: 'a workflow reaches the world only through the functions it is given', + howAWorkflowStarts: 'start it through its runtime', +}; + export function workflow(source: string): JsonObject { return { document: header, ...yamlObject(source) }; } diff --git a/primitives/orchestration/README.md b/primitives/orchestration/README.md index 3625ca8d5..0bfbe1650 100644 --- a/primitives/orchestration/README.md +++ b/primitives/orchestration/README.md @@ -252,7 +252,7 @@ A worker given only the workflow code bundles it with webpack and swc when it st ## Not in this version -Starting workflows from schedules or events, `run`, `emit`, outbound calls, catalogs and functions beyond `execute_spec`, executing a workflow from a workflow (or any spec that finishes later), checking inputs and outputs against their schemas, listening `until` a condition, correlating events, search attributes, and showing the progress of a run. A function is added by adding an activity to the activity contract (`src/workflow/activity-contract.ts`), a body to the `call` task (`src/interpreter/call-task.ts`), and its name to the policy (`packages/workflow-engine/src/dsl/task-policy.ts`). +Starting workflows from schedules or events, `run`, `emit`, outbound calls, catalogs and functions beyond `execute_spec`, executing a workflow from a workflow (or any spec that finishes later), checking inputs and outputs against their schemas, listening `until` a condition, correlating events, search attributes, and showing the progress of a run. A function is added by adding an activity to the activity contract (`src/workflow/activity-contract.ts`), a body to the `call` task (`src/interpreter/call-task.ts`), and its name and the checks of its arguments to the functions the policy is given (`src/document/workflow-functions.ts`). ## Testing diff --git a/primitives/orchestration/src/document/workflow-document.ts b/primitives/orchestration/src/document/workflow-document.ts index dd9b8057c..d4e37730f 100644 --- a/primitives/orchestration/src/document/workflow-document.ts +++ b/primitives/orchestration/src/document/workflow-document.ts @@ -3,12 +3,12 @@ import { InvalidInput, type Issue } from '@beonauto/operations'; import { durationLimitRejections } from '@beonauto/workflow-engine/dsl/duration-limits'; import { mostValueDepth, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; import { nestingRejections } from '@beonauto/workflow-engine/dsl/nesting'; -import { rejectionsOf } from '@beonauto/workflow-engine/dsl/policy'; import type { Rejection } from '@beonauto/workflow-engine/dsl/policy-checks'; import { Effect } from 'effect'; import { defaultMostDuration } from '../interpreter/workflow-run.ts'; import { dslProblems, type Problem } from './dsl-validation.ts'; +import { workflowPolicy } from './workflow-functions.ts'; const workflowYaml: YamlKind = { noun: 'a workflow document', @@ -32,7 +32,7 @@ export function parseWorkflowDocument( return Effect.fail(invalidDocument('The workflow document is not YAML this runtime reads', reading.problems)); } const { value, locate } = reading.document; - const rejections = rejectionsOf(value); + const rejections = workflowPolicy(value); const unreadable = nestingRejections(value).length > 0 || rejections.some(({ pointer }) => pointer === '/document/dsl'); const problems = [ diff --git a/primitives/orchestration/src/document/workflow-functions.test.ts b/primitives/orchestration/src/document/workflow-functions.test.ts new file mode 100644 index 000000000..47340024f --- /dev/null +++ b/primitives/orchestration/src/document/workflow-functions.test.ts @@ -0,0 +1,71 @@ +import { describe, expect, it } from 'vitest'; + +import { workflow } from '../testing/workflows.ts'; +import { workflowPolicy } from './workflow-functions.ts'; + +function rejectedIn(tasks: string): readonly string[] { + return workflowPolicy(workflow(`do:\n${tasks}`)).map(({ pointer, detail }) => `${pointer}: ${detail}`); +} + +describe('the functions a workflow calls', () => { + it('are execute_spec alone: a workflow reaches the world through the specs of its brain', () => { + expect( + rejectedIn(` + - fetch: { call: http, with: {} } + - log: { call: log, with: {} } +`), + ).toEqual([ + '/do/0/fetch/call: call: http is not allowed: a workflow reaches the world only through the specs of its brain; call execute_spec', + '/do/1/log/call: call: log names no function; the one function is execute_spec', + ]); + }); + + it('are named when a document brings catalogs or functions, or a schedule', () => { + expect( + workflowPolicy(workflow('schedule: { every: PT1H }\nuse: { catalogs: {}, functions: {} }\ndo: []')).map( + ({ detail }) => detail, + ), + ).toEqual([ + 'catalogs are not supported in this version: a workflow calls only execute_spec', + 'reusable functions are not supported in this version: call execute_spec directly', + 'schedules are not supported in this version: execute the spec to run it', + ]); + }); +}); + +describe('the policy of execute_spec', () => { + it('takes a primitive, a name and an input, as expressions or literals', () => { + expect( + rejectedIn(` + - summarize: + call: execute_spec + with: { primitive: inference, name: '\${ .spec }', input: { text: '\${ .text }' } } +`), + ).toEqual([]); + }); + + it('rejects arguments it does not take and arguments it lacks', () => { + expect( + rejectedIn(` + - nothing: { call: execute_spec } + - partial: { call: execute_spec, with: { name: 3, model: big } } +`), + ).toEqual([ + '/do/0/nothing/with: execute_spec takes with: { primitive, name, input }', + '/do/1/partial/with/model: execute_spec takes no argument model', + '/do/1/partial/with/primitive: execute_spec needs a string primitive', + '/do/1/partial/with/name: execute_spec needs a string name', + ]); + }); + + it('rejects executing another workflow, and broken expressions in its arguments', () => { + expect( + rejectedIn(` + - nested: { call: execute_spec, with: { primitive: orchestration, name: other, input: ['\${ .a + }'] } } +`), + ).toEqual([ + '/do/0/nested/with/primitive: A workflow cannot execute another workflow in this version', + expect.stringMatching(/^\/do\/0\/nested\/with\/input\/0: /u), + ]); + }); +}); diff --git a/primitives/orchestration/src/document/workflow-functions.ts b/primitives/orchestration/src/document/workflow-functions.ts new file mode 100644 index 000000000..3615188f5 --- /dev/null +++ b/primitives/orchestration/src/document/workflow-functions.ts @@ -0,0 +1,34 @@ +import type { CallFunctions } from '@beonauto/workflow-engine/dsl/call-functions'; +import { field, isObject, type Json } from '@beonauto/workflow-engine/dsl/json'; +import { policyOf } from '@beonauto/workflow-engine/dsl/policy'; +import { forbidden, rejection, templateRejections, type Rejection } from '@beonauto/workflow-engine/dsl/policy-checks'; +import { pointerTo } from '@beonauto/workflow-engine/dsl/tasks'; + +export const executeSpecFunction = 'execute_spec'; + +const executeSpecArguments = new Set(['primitive', 'name', 'input']); + +function executeSpecRejections(arguments_: Json | undefined, pointer: string): readonly Rejection[] { + if (!isObject(arguments_)) { + return [rejection(pointer, `${executeSpecFunction} takes with: { primitive, name, input }`)]; + } + const unknown = Object.keys(arguments_) + .filter((key) => !executeSpecArguments.has(key)) + .map((key) => rejection(pointerTo(pointer, key), `${executeSpecFunction} takes no argument ${key}`)); + const missing = ['primitive', 'name'] + .filter((key) => typeof field(arguments_, key) !== 'string') + .map((key) => rejection(pointerTo(pointer, key), `${executeSpecFunction} needs a string ${key}`)); + const workflow = + field(arguments_, 'primitive') === 'orchestration' + ? [forbidden(`${pointer}/primitive`, 'A workflow cannot execute another workflow in this version')] + : []; + return unknown.concat(missing, workflow, templateRejections(arguments_, pointer)); +} + +const workflowFunctions: CallFunctions = { + argumentChecks: { [executeSpecFunction]: executeSpecRejections }, + howAWorkflowReachesTheWorld: 'a workflow reaches the world only through the specs of its brain', + howAWorkflowStarts: 'execute the spec to run it', +}; + +export const workflowPolicy = policyOf(workflowFunctions); diff --git a/primitives/orchestration/src/interpreter/call-task.ts b/primitives/orchestration/src/interpreter/call-task.ts index ea18caace..5d216501d 100644 --- a/primitives/orchestration/src/interpreter/call-task.ts +++ b/primitives/orchestration/src/interpreter/call-task.ts @@ -1,6 +1,6 @@ import { field, isObject, jsonBytesOf, textField, type Json } from '@beonauto/workflow-engine/dsl/json'; -import { executeSpecFunction } from '@beonauto/workflow-engine/dsl/task-policy'; +import { executeSpecFunction } from '../document/workflow-functions.ts'; import { evaluateTemplate, placeOf } from './evaluation.ts'; import type { SpecCall, SpecCallResult } from './host.ts'; import type { Body, Invocation } from './invocation.ts'; diff --git a/primitives/orchestration/src/interpreter/interpreter.ts b/primitives/orchestration/src/interpreter/interpreter.ts index f0b9e8285..c527e4159 100644 --- a/primitives/orchestration/src/interpreter/interpreter.ts +++ b/primitives/orchestration/src/interpreter/interpreter.ts @@ -1,6 +1,6 @@ import { field, jsonBytesOf, objectField, type Json } from '@beonauto/workflow-engine/dsl/json'; -import { rejectionsOf } from '@beonauto/workflow-engine/dsl/policy'; +import { workflowPolicy } from '../document/workflow-functions.ts'; import { placeIn, transform } from './evaluation.ts'; import type { WorkflowHost } from './host.ts'; import { RaisedError, errorType } from './raised-error.ts'; @@ -94,7 +94,7 @@ async function outcomeOf(state: RunState): Promise { async function interpret(state: RunState): Promise { const { document, input } = state.run; - const rejections = rejectionsOf(document).filter(({ forbidden }) => forbidden); + const rejections = workflowPolicy(document).filter(({ forbidden }) => forbidden); if (rejections.length > 0) { throw new RaisedError({ type: errorType('configuration'), From 425d9b0045290ef5a848c2e199abbd0541376f82 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 19:51:55 +0100 Subject: [PATCH 13/23] refactor(orchestration): take the limits and DslError from the workflow engine The limits the interpreter repeated verbatim, tasks without waiting, expression work, work in one activation, tasks in one activation and the four limits on events, now live once in the engine's limits, which the package exports as @beonauto/workflow-engine/limits, a module of plain numbers that the bundled workflow code can import. DslError is the engine's DslErrorSchema type; the interpreter imports it as a type. The limits that differ under Temporal, the 16 MiB a workflow holds and its history, stay with the interpreter. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/package.json | 3 ++- .../src/inbox/received-event.ts | 10 +--------- packages/workflow-engine/src/index.ts | 17 +++++++---------- .../workflow-engine/src/machine/dsl-error.ts | 11 +++++++++++ packages/workflow-engine/src/machine/limits.ts | 10 ++++++++++ .../workflow-engine/src/machine/run-state.ts | 17 +---------------- .../src/events/send-execution-event.ts | 6 +++--- .../src/interpreter/error-tasks.ts | 3 ++- .../src/interpreter/evaluation.ts | 5 +++-- .../src/interpreter/event-limits.test.ts | 7 ++++++- .../orchestration/src/interpreter/inbox.ts | 14 ++++++-------- .../src/interpreter/limits.test.ts | 3 ++- .../src/interpreter/raised-error.ts | 9 +-------- .../orchestration/src/interpreter/run-state.ts | 18 ++++++++---------- .../src/interpreter/settlement.ts | 3 ++- 15 files changed, 65 insertions(+), 71 deletions(-) create mode 100644 packages/workflow-engine/src/machine/dsl-error.ts diff --git a/packages/workflow-engine/package.json b/packages/workflow-engine/package.json index 6c960bfde..7b28c7f4d 100644 --- a/packages/workflow-engine/package.json +++ b/packages/workflow-engine/package.json @@ -6,7 +6,8 @@ "type": "module", "exports": { ".": "./src/index.ts", - "./dsl/*": "./src/dsl/*.ts" + "./dsl/*": "./src/dsl/*.ts", + "./limits": "./src/machine/limits.ts" }, "scripts": { "lint": "oxlint -c ../../.oxlintrc.json", diff --git a/packages/workflow-engine/src/inbox/received-event.ts b/packages/workflow-engine/src/inbox/received-event.ts index 952de4281..9cecb9a66 100644 --- a/packages/workflow-engine/src/inbox/received-event.ts +++ b/packages/workflow-engine/src/inbox/received-event.ts @@ -1,14 +1,6 @@ import { Schema } from 'effect'; -export const mostEventIdLength = 256; - -export const mostWaitingEvents = 64; - -export const mostWaitingEventBytes = 1_048_576; - -export const mostReceivedEvents = 1024; - -export const mostReceivedEventBytes = 4_194_304; +import { mostEventIdLength } from '../machine/limits.ts'; const EventIdSchema = Schema.String.check(Schema.isMinLength(1), Schema.isMaxLength(mostEventIdLength)); diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts index edb5212af..244779870 100644 --- a/packages/workflow-engine/src/index.ts +++ b/packages/workflow-engine/src/index.ts @@ -18,15 +18,7 @@ export { export type { EnginePorts, Submission, SweepReport, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; export type { CallCancelReceipt, Executor, StartReceipt } from './executor/executor.ts'; -export { - ReceivedEventSchema, - mostEventIdLength, - mostReceivedEventBytes, - mostReceivedEvents, - mostWaitingEventBytes, - mostWaitingEvents, - type ReceivedEvent, -} from './inbox/received-event.ts'; +export { ReceivedEventSchema, type ReceivedEvent } from './inbox/received-event.ts'; export { RunMismatch, outcomeOf, @@ -34,17 +26,23 @@ export { type StaleReason, type SubmissionOutcome, } from './machine/admission.ts'; +export { DslErrorSchema, type DslError } from './machine/dsl-error.ts'; export { InputReceiptSchema, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; export { InstantSchema, clampedAt } from './machine/instant.ts'; export { mostCallArgumentsBytes, mostEventBytes, + mostEventIdLength, mostExpressionWork, mostHeldBytes, mostHistoryBytes, mostInputs, + mostReceivedEventBytes, + mostReceivedEvents, mostStepsWithoutWaiting, mostTasksPerInput, + mostWaitingEventBytes, + mostWaitingEvents, mostWorkPerInput, taskFrameBytes, } from './machine/limits.ts'; @@ -66,7 +64,6 @@ export { type ArmedTimer, type Branch, type CursorCurrent, - type DslError, type FrameBody, type HeldValue, type InboxState, diff --git a/packages/workflow-engine/src/machine/dsl-error.ts b/packages/workflow-engine/src/machine/dsl-error.ts new file mode 100644 index 000000000..f92800832 --- /dev/null +++ b/packages/workflow-engine/src/machine/dsl-error.ts @@ -0,0 +1,11 @@ +import { Schema } from 'effect'; + +export const DslErrorSchema = Schema.Struct({ + type: Schema.String, + status: Schema.Int, + instance: Schema.String, + title: Schema.optionalKey(Schema.String), + detail: Schema.optionalKey(Schema.String), +}); + +export type DslError = typeof DslErrorSchema.Type; diff --git a/packages/workflow-engine/src/machine/limits.ts b/packages/workflow-engine/src/machine/limits.ts index 1d22bf1f6..396c5a9ca 100644 --- a/packages/workflow-engine/src/machine/limits.ts +++ b/packages/workflow-engine/src/machine/limits.ts @@ -17,3 +17,13 @@ export const mostWorkPerInput = 16_000_000; export const mostInputs = 100_000; export const mostHistoryBytes = 536_870_912; + +export const mostEventIdLength = 256; + +export const mostWaitingEvents = 64; + +export const mostWaitingEventBytes = 1_048_576; + +export const mostReceivedEvents = 1024; + +export const mostReceivedEventBytes = 4_194_304; diff --git a/packages/workflow-engine/src/machine/run-state.ts b/packages/workflow-engine/src/machine/run-state.ts index fc6da503a..3585200a4 100644 --- a/packages/workflow-engine/src/machine/run-state.ts +++ b/packages/workflow-engine/src/machine/run-state.ts @@ -3,17 +3,10 @@ import { Schema } from 'effect'; import { CallKeySchema, type CallKey } from '../executor/call-key.ts'; import { ReceivedEventSchema, type ReceivedEvent } from '../inbox/received-event.ts'; import { TimerPurposeSchema, type TimerPurpose } from '../timers/timer-id.ts'; +import { DslErrorSchema, type DslError } from './dsl-error.ts'; import { InstantSchema } from './instant.ts'; import { RunLimitsSchema, type RunLimits } from './run-input.ts'; -export interface DslError { - readonly type: string; - readonly status: number; - readonly instance: string; - readonly title?: string; - readonly detail?: string; -} - export type ValueId = number; export interface HeldValue { @@ -138,14 +131,6 @@ const ValueIdSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)); const VariablesSchema = Schema.Record(Schema.String, ValueIdSchema); -const DslErrorSchema = Schema.Struct({ - type: Schema.String, - status: IntSchema, - instance: Schema.String, - title: Schema.optionalKey(Schema.String), - detail: Schema.optionalKey(Schema.String), -}); - const TaskFrameReference = Schema.suspend((): Schema.Codec => TaskFrameSchema); const CursorCurrentSchema: Schema.Codec = Schema.Union([ diff --git a/primitives/orchestration/src/events/send-execution-event.ts b/primitives/orchestration/src/events/send-execution-event.ts index abaa68381..f3d08510d 100644 --- a/primitives/orchestration/src/events/send-execution-event.ts +++ b/primitives/orchestration/src/events/send-execution-event.ts @@ -3,14 +3,14 @@ import { randomUUID } from 'node:crypto'; import { BrainContext, defineCommand, NotFound, quoted } from '@beonauto/operations'; import { getExecution } from '@beonauto/specs'; import { jsonBytesOf } from '@beonauto/workflow-engine/dsl/json'; -import { DateTime, Effect, Schema, SchemaTransformation } from 'effect'; - import { mostReceivedEventBytes, mostReceivedEvents, mostWaitingEventBytes, mostWaitingEvents, -} from '../interpreter/inbox.ts'; +} from '@beonauto/workflow-engine/limits'; +import { DateTime, Effect, Schema, SchemaTransformation } from 'effect'; + import type { OrchestrationClient } from '../primitive/orchestration-client.ts'; const mostEventBytes = 262_144; diff --git a/primitives/orchestration/src/interpreter/error-tasks.ts b/primitives/orchestration/src/interpreter/error-tasks.ts index 8946f52a0..02a131d25 100644 --- a/primitives/orchestration/src/interpreter/error-tasks.ts +++ b/primitives/orchestration/src/interpreter/error-tasks.ts @@ -1,3 +1,4 @@ +import type { DslError } from '@beonauto/workflow-engine'; import type { Variables } from '@beonauto/workflow-engine/dsl/expressions'; import { entriesOf, @@ -11,7 +12,7 @@ import { import { evaluateTemplate, holds, placeOf } from './evaluation.ts'; import { bodyOf, type Body, type Invocation } from './invocation.ts'; -import { RaisedError, errorAsJson, errorFromJson, raised, type DslError } from './raised-error.ts'; +import { RaisedError, errorAsJson, errorFromJson, raised } from './raised-error.ts'; import { attemptDuration, retryDelay, retryPolicyOf, type RetryContext, type RetryState } from './retry-policy.ts'; import { withTimeout } from './timeouts.ts'; diff --git a/primitives/orchestration/src/interpreter/evaluation.ts b/primitives/orchestration/src/interpreter/evaluation.ts index 632c89534..9733648c8 100644 --- a/primitives/orchestration/src/interpreter/evaluation.ts +++ b/primitives/orchestration/src/interpreter/evaluation.ts @@ -14,10 +14,11 @@ import { type JsonEntry, type JsonObject, } from '@beonauto/workflow-engine/dsl/json'; +import { mostExpressionWork, mostWorkPerInput } from '@beonauto/workflow-engine/limits'; import type { Invocation, Place } from './invocation.ts'; import { RaisedError, errorType, raised } from './raised-error.ts'; -import { mostActivationWork, mostExpressionWork, type RunState } from './run-state.ts'; +import type { RunState } from './run-state.ts'; export function evaluate(source: string, data: Json, variables: Variables, place: Place): Json { const mostWork = place.meter.allowance(); @@ -44,7 +45,7 @@ export function placeIn(state: RunState, reference: string): Place { function exhaustionOf(problem: string, mostWork: number): string { return mostWork < mostExpressionWork - ? `${problem}: the workflow did ${mostActivationWork} units of expression work in one activation; it lets other workflows run between tasks, not within one` + ? `${problem}: the workflow did ${mostWorkPerInput} units of expression work in one activation; it lets other workflows run between tasks, not within one` : `${problem}: an expression may do ${mostExpressionWork} units of work`; } diff --git a/primitives/orchestration/src/interpreter/event-limits.test.ts b/primitives/orchestration/src/interpreter/event-limits.test.ts index 159cdd120..226d47507 100644 --- a/primitives/orchestration/src/interpreter/event-limits.test.ts +++ b/primitives/orchestration/src/interpreter/event-limits.test.ts @@ -1,9 +1,14 @@ +import { + mostReceivedEventBytes, + mostReceivedEvents, + mostWaitingEventBytes, + mostWaitingEvents, +} from '@beonauto/workflow-engine/limits'; import { describe, expect, it } from 'vitest'; import type { FakeHost } from '../testing/fake-host.ts'; import { interpret, workflow } from '../testing/workflows.ts'; import type { RunSettlement } from './host.ts'; -import { mostReceivedEventBytes, mostReceivedEvents, mostWaitingEventBytes, mostWaitingEvents } from './inbox.ts'; import type { WorkflowStart } from './interpreter.ts'; const waitingForever = workflow(` diff --git a/primitives/orchestration/src/interpreter/inbox.ts b/primitives/orchestration/src/interpreter/inbox.ts index b0f0925c7..4134f0cc7 100644 --- a/primitives/orchestration/src/interpreter/inbox.ts +++ b/primitives/orchestration/src/interpreter/inbox.ts @@ -1,4 +1,10 @@ import { isJson, isObject, jsonBytesOf, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; +import { + mostReceivedEventBytes, + mostReceivedEvents, + mostWaitingEventBytes, + mostWaitingEvents, +} from '@beonauto/workflow-engine/limits'; import { raised, type RaisedError } from './raised-error.ts'; import { retainedBytesOf } from './retained-size.ts'; @@ -12,14 +18,6 @@ export interface Inbox { readonly overflow: () => RaisedError | undefined; } -export const mostWaitingEvents = 64; - -export const mostWaitingEventBytes = 1_048_576; - -export const mostReceivedEvents = 1024; - -export const mostReceivedEventBytes = 4_194_304; - interface Waiting { readonly event: JsonObject; readonly bytes: number; diff --git a/primitives/orchestration/src/interpreter/limits.test.ts b/primitives/orchestration/src/interpreter/limits.test.ts index d689ab4b9..d0df1154f 100644 --- a/primitives/orchestration/src/interpreter/limits.test.ts +++ b/primitives/orchestration/src/interpreter/limits.test.ts @@ -1,7 +1,8 @@ +import { mostStepsWithoutWaiting } from '@beonauto/workflow-engine/limits'; import { describe, expect, it } from 'vitest'; import { interpret, workflow } from '../testing/workflows.ts'; -import { mostHistoryBytes, mostHistoryEvents, mostStepsWithoutWaiting } from './run-state.ts'; +import { mostHistoryBytes, mostHistoryEvents } from './run-state.ts'; const calling = workflow('do:\n - fetch: { call: execute_spec, with: { primitive: inference, name: lookup } }'); diff --git a/primitives/orchestration/src/interpreter/raised-error.ts b/primitives/orchestration/src/interpreter/raised-error.ts index a823e321c..cc6ed6a20 100644 --- a/primitives/orchestration/src/interpreter/raised-error.ts +++ b/primitives/orchestration/src/interpreter/raised-error.ts @@ -1,3 +1,4 @@ +import type { DslError } from '@beonauto/workflow-engine'; import { field, textField, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; export type ErrorKind = @@ -10,14 +11,6 @@ export type ErrorKind = | 'communication' | 'runtime'; -export interface DslError { - readonly type: string; - readonly status: number; - readonly instance: string; - readonly title?: string; - readonly detail?: string; -} - export class RaisedError extends Error { readonly error: DslError; diff --git a/primitives/orchestration/src/interpreter/run-state.ts b/primitives/orchestration/src/interpreter/run-state.ts index 68f77c32e..72c70257d 100644 --- a/primitives/orchestration/src/interpreter/run-state.ts +++ b/primitives/orchestration/src/interpreter/run-state.ts @@ -1,5 +1,11 @@ import { measureOf, mostValueDepth, objectField, type Json, type JsonObject } from '@beonauto/workflow-engine/dsl/json'; import type { Components } from '@beonauto/workflow-engine/dsl/policy-checks'; +import { + mostExpressionWork, + mostStepsWithoutWaiting, + mostTasksPerInput, + mostWorkPerInput, +} from '@beonauto/workflow-engine/limits'; import { makeHolding, type Hold } from './holding.ts'; import type { WorkflowHost } from './host.ts'; @@ -39,14 +45,6 @@ export const mostHistoryBytes = 8_388_608; export const mostHistoryEvents = 40_000; -export const mostStepsWithoutWaiting = 10_000; - -export const mostExpressionWork = 8_000_000; - -export const mostActivationWork = 16_000_000; - -const mostTasksPerActivation = 100; - const mostValueWork = mostExpressionWork; export function dateTimeOf(milliseconds: number): JsonObject { @@ -145,7 +143,7 @@ function makeMeter(host: WorkflowHost): Meter { return { allowance: () => { current(); - return Math.min(mostExpressionWork, mostActivationWork - work); + return Math.min(mostExpressionWork, mostWorkPerInput - work); }, record: (done) => { current(); @@ -153,7 +151,7 @@ function makeMeter(host: WorkflowHost): Meter { }, shouldYield: () => { current(); - return work >= mostExpressionWork || tasks >= mostTasksPerActivation; + return work >= mostExpressionWork || tasks >= mostTasksPerInput; }, countTask: () => { current(); diff --git a/primitives/orchestration/src/interpreter/settlement.ts b/primitives/orchestration/src/interpreter/settlement.ts index 04844074e..b99eb6689 100644 --- a/primitives/orchestration/src/interpreter/settlement.ts +++ b/primitives/orchestration/src/interpreter/settlement.ts @@ -1,7 +1,8 @@ +import type { DslError } from '@beonauto/workflow-engine'; import type { Json } from '@beonauto/workflow-engine/dsl/json'; import type { RunSettlement } from './host.ts'; -import { describeError, type DslError } from './raised-error.ts'; +import { describeError } from './raised-error.ts'; export type RunOutcome = | { readonly kind: 'completed'; readonly output: Json } From 8f18fcafb0231841f5c67efee999326496e581f5 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:21:21 +0100 Subject: [PATCH 14/23] test(workflow-engine): scan the DSL and the jq library the machine runs The purity scan now covers the DSL beside the machine and the run log, and a scan of jq-ts's distribution, the one library the machine runs expressions with, finds no Node-only API, code generation, host timer, clock read or random source. jq reaches the host's time zone only through localtime and strflocaltime, the two builtins it builds with local time, and the DSL refuses both; the test checks both facts. Co-Authored-By: Claude Opus 5.5 --- .../src/engine/portability.test.ts | 42 +++++++++++++++++-- 1 file changed, 38 insertions(+), 4 deletions(-) diff --git a/packages/workflow-engine/src/engine/portability.test.ts b/packages/workflow-engine/src/engine/portability.test.ts index 61e620407..96b2de24c 100644 --- a/packages/workflow-engine/src/engine/portability.test.ts +++ b/packages/workflow-engine/src/engine/portability.test.ts @@ -5,6 +5,8 @@ import { fileURLToPath } from 'node:url'; import { describe, expect, it } from 'vitest'; +import { checkExpression } from '../dsl/expressions.ts'; + interface Forbidden { readonly what: string; readonly pattern: Readonly; @@ -25,12 +27,23 @@ const nodeOnly: readonly Forbidden[] = [ { what: 'the Temporal global', pattern: /\bTemporal\./u }, ]; -const impure: readonly Forbidden[] = [ - { what: 'the clock', pattern: /\bDate\.now\(|\bnew Date\b|\bperformance\.now\(/u }, +const clockOrRandom: readonly Forbidden[] = [ + { what: 'the clock', pattern: /\bDate\.now\(|\bperformance\.now\(/u }, { what: 'a random source', pattern: /\bMath\.random\(|\bcrypto\.getRandomValues\(|\brandomUUID\(/u }, +]; + +const impure: readonly Forbidden[] = [ + ...clockOrRandom, + { what: 'a date of the host', pattern: /\bnew Date\b/u }, { what: 'the host locale or time zone', pattern: /\bIntl\./u }, ]; +const blockComments = /\/\*[\s\S]*?\*\//gu; + +const localTimeBuiltins = /Builtin\("\w+", false\)/gu; + +const jq = readFileSync(fileURLToPath(import.meta.resolve('@gabrielbryk/jq-ts')), 'utf8').replaceAll(blockComments, ''); + const growingCache: readonly Forbidden[] = [ { what: 'a module-level collection', @@ -59,13 +72,19 @@ function findingsIn(files: readonly string[], forbidden: readonly Forbidden[]): }); } +function localTimeBuiltinsIn(text: string): readonly string[] { + return (text.match(localTimeBuiltins) ?? []).map((call: string) => + call.slice('Builtin("'.length, call.indexOf('",')), + ); +} + function caught(text: string, forbidden: readonly Forbidden[]): readonly string[] { return forbidden.filter(({ pattern }: Forbidden) => pattern.test(text)).map(({ what }: Forbidden) => what); } const everySource = sourcesUnder('.'); -const machineAndRunLog = [...sourcesUnder('machine'), ...sourcesUnder('run-log')]; +const machineAndRunLog = [...sourcesUnder('machine'), ...sourcesUnder('run-log'), ...sourcesUnder('dsl')]; describe('the engine core', () => { it('uses no Node-only API, no dynamic import, no code generation and no Temporal, so it runs in workerd as in Node', () => { @@ -78,13 +97,28 @@ describe('the engine core', () => { }); }); -describe('the machine and the run log', () => { +describe('the machine, the run log and the DSL', () => { it('read no clock, no random source and no locale, so the same state and input decide the same events', () => { expect(machineAndRunLog.length).toBeGreaterThan(10); expect(findingsIn(machineAndRunLog, impure)).toEqual([]); }); }); +describe('the jq library the machine runs expressions with', () => { + it('uses no Node-only API, no code generation and no host timer, and reads no clock or random source', () => { + expect(jq.length).toBeGreaterThan(100_000); + expect(caught(jq, [...nodeOnly, ...clockOrRandom])).toEqual([]); + }); + + it('reaches the host time zone only through localtime and strflocaltime, which the DSL refuses', () => { + expect(localTimeBuiltinsIn(jq)).toEqual(['localtime', 'strflocaltime']); + expect([checkExpression('now | localtime'), checkExpression('now | strflocaltime("%H")')]).toEqual([ + expect.stringContaining("localtime reads the host's time zone"), + expect.stringContaining("strflocaltime reads the host's time zone"), + ]); + }); +}); + describe('the checks of purity', () => { it('catch what they are there to catch', () => { expect( From ee84e71f2c7134271448796e90526a0f0eb59239 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:21:51 +0100 Subject: [PATCH 15/23] feat(workflow-engine): second revision of the contract after the audit - History bytes live in the state. withHistoryBytes adds the event's own replace of /historyBytes, solved for the size of the event that carries it, so decide can bound the history as it bounds inputs, and a load dies when the state's count is not the bytes of its events. - The whole-state decode refuses members the format does not know. - Each event is folded under its own state format, the state is upcast where the format changes, formats never go back within a stream, and corpus/format-1.json, a committed stream and snapshot, must load as a test. - A call frame holds the opaque function, its arguments as a value id and a label; the timer purpose call_deadline pairs every started call with a deadline; start receipts tell a call started again from one still running or answered again. - Values are held while reachable: withReachableValuesOnly sweeps the value table after a decision, and a property test over 400 states checks heldBytes against the bytes the frames, context and input reach. - The record notes each live run's next due time and whether its dispatch fell behind, so a sweep touches only overdue runs; troubling settle receipts go to a RunReporter. - A second start with the same document and another input dies. - A fired timer's input takes at least its due time. - Snapshots are due by bytes alone and chunked with encodeInto. - runLoopOf runs the ledger's decisionLoop with a load from a snapshot and its tail. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/corpus/format-1.json | 283 ++++++++++++++++++ .../src/dispatch/dispatch-watermark.test.ts | 1 - .../src/dispatch/run-due.test.ts | 38 +++ .../workflow-engine/src/dispatch/run-due.ts | 16 + .../src/engine/run-loop.test.ts | 71 +++++ .../workflow-engine/src/engine/run-loop.ts | 31 ++ .../src/engine/workflow-engine.ts | 6 +- .../workflow-engine/src/executor/executor.ts | 2 +- packages/workflow-engine/src/index.ts | 37 ++- .../src/machine/admission.test.ts | 5 +- .../workflow-engine/src/machine/admission.ts | 6 +- .../src/machine/held-values.test.ts | 198 ++++++++++++ .../src/machine/held-values.ts | 86 ++++++ .../src/machine/input-receipt.test.ts | 22 +- .../src/machine/input-receipt.ts | 13 +- .../workflow-engine/src/machine/run-state.ts | 18 +- .../src/run-log/corpus.test.ts | 59 ++++ .../src/run-log/run-event.test.ts | 19 +- .../workflow-engine/src/run-log/run-event.ts | 14 +- .../src/run-log/run-fold.test.ts | 221 ++++++++++---- .../workflow-engine/src/run-log/run-fold.ts | 104 +++++-- .../workflow-engine/src/run-log/run-store.ts | 7 +- .../src/run-log/snapshot.test.ts | 49 ++- .../workflow-engine/src/run-log/snapshot.ts | 56 ++-- .../src/run-log/state-format.ts | 16 +- .../src/settlement/record-store.ts | 25 +- .../workflow-engine/src/testing/run-store.ts | 66 ++++ packages/workflow-engine/src/testing/runs.ts | 16 +- .../workflow-engine/src/testing/streams.ts | 61 ++++ .../workflow-engine/src/timers/timer-id.ts | 1 + 30 files changed, 1351 insertions(+), 196 deletions(-) create mode 100644 packages/workflow-engine/corpus/format-1.json create mode 100644 packages/workflow-engine/src/dispatch/run-due.test.ts create mode 100644 packages/workflow-engine/src/dispatch/run-due.ts create mode 100644 packages/workflow-engine/src/engine/run-loop.test.ts create mode 100644 packages/workflow-engine/src/engine/run-loop.ts create mode 100644 packages/workflow-engine/src/machine/held-values.test.ts create mode 100644 packages/workflow-engine/src/machine/held-values.ts create mode 100644 packages/workflow-engine/src/run-log/corpus.test.ts create mode 100644 packages/workflow-engine/src/testing/run-store.ts create mode 100644 packages/workflow-engine/src/testing/streams.ts diff --git a/packages/workflow-engine/corpus/format-1.json b/packages/workflow-engine/corpus/format-1.json new file mode 100644 index 000000000..f4d12a199 --- /dev/null +++ b/packages/workflow-engine/corpus/format-1.json @@ -0,0 +1,283 @@ +{ + "format": 1, + "stream": [ + { + "version": 1, + "event": { + "type": "input_applied", + "format": 1, + "receipt": { + "kind": "started", + "key": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a", + "at": 1791100060000 + }, + "steps": [], + "patch": [ + { + "op": "replace", + "path": "/executionId", + "value": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a" + }, + { + "op": "replace", + "path": "/status", + "value": "running" + }, + { + "op": "replace", + "path": "/workflow", + "value": { + "document": { + "document": { + "dsl": "1.0.3", + "namespace": "acme", + "name": "triage", + "version": "1.0.0" + }, + "do": [] + }, + "input": 1 + } + }, + { + "op": "add", + "path": "/machine/values/1", + "value": { + "value": { + "ticket": 7 + }, + "bytes": 12 + } + }, + { + "op": "replace", + "path": "/inputs", + "value": 1 + }, + { + "op": "replace", + "path": "/startedAt", + "value": 1791100060000 + }, + { + "op": "replace", + "path": "/lastInputAt", + "value": 1791100060000 + }, + { + "op": "replace", + "path": "/historyBytes", + "value": 755 + } + ], + "outputs": [] + } + }, + { + "version": 2, + "event": { + "type": "input_applied", + "format": 1, + "receipt": { + "kind": "event_received", + "key": "event-1", + "at": 1791100060010, + "eventType": "com.acme.tick" + }, + "steps": [], + "patch": [ + { + "op": "add", + "path": "/timers/armed/0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a~1timers~11", + "value": { + "purpose": "wait", + "reference": "/do/0", + "dueAt": 1791100061000 + } + }, + { + "op": "replace", + "path": "/timers/next", + "value": 2 + }, + { + "op": "replace", + "path": "/inputs", + "value": 2 + }, + { + "op": "replace", + "path": "/lastInputAt", + "value": 1791100060010 + }, + { + "op": "replace", + "path": "/historyBytes", + "value": 1453 + } + ], + "outputs": [ + { + "kind": "arm_timer", + "executionId": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a", + "timerId": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a/timers/1", + "dueAt": 1791100061000, + "purpose": "wait" + } + ] + } + }, + { + "version": 3, + "event": { + "type": "input_applied", + "format": 1, + "receipt": { + "kind": "timer_fired", + "key": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a/timers/1", + "at": 1791100061000 + }, + "steps": [], + "patch": [ + { + "op": "remove", + "path": "/timers/armed/0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a~1timers~11" + }, + { + "op": "replace", + "path": "/machine/context", + "value": 1 + }, + { + "op": "replace", + "path": "/inputs", + "value": 3 + }, + { + "op": "replace", + "path": "/lastInputAt", + "value": 1791100061000 + }, + { + "op": "replace", + "path": "/historyBytes", + "value": 1926 + } + ], + "outputs": [] + } + } + ], + "snapshot": { + "chunks": [ + "{\"format\":1,\"executionId\":\"0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a\",\"version\":2,\"historyBytes\":1453,\"state\":{\"executionId\":\"0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a\",\"status\":\"running\",\"workflow\":{\"document\":{\"document\":{\"dsl\":\"1.0.3\",\"namespace\":\"acme\",\"name\":\"triage\",\"version\":\"1.0.0\"},\"do\":[]},\"input\":1},\"attributes\":{},\"limits\":{\"mostDurationMs\":1,\"longestCallMs\":1},\"startedAt\":1791100060000,\"lastInputAt\":1791100060010,\"inputs\":2,\"random\":{\"seed\":0,\"draws\":0},\"runs\":{},\"timers\":{\"next\":2,\"armed\":{\"0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a/timers/1\":{\"purpose\":\"wait\",\"reference\":\"/do/0\",\"dueAt\":1791100061000}}},\"calls\":{},\"inbox\":{\"waiting\":[],\"waitingBytes\":0,\"receivedIds\":[],\"received\":0,\"receivedBytes\":0,\"overflow\":null},\"heldBytes\":0,\"historyBytes\":1453,\"stepsWithoutWaiting\":0,\"cancelRequested\":false,\"machine\":{\"values\":{\"0\":{\"value\":{},\"bytes\":2},\"1\":{\"value\":{\"ticket\":7},\"bytes\":12}},\"nextValue\":1,\"context\":0,\"root\":null},\"outcome\":null}}" + ], + "bytes": 949, + "tail": [ + { + "version": 3, + "event": { + "type": "input_applied", + "format": 1, + "receipt": { + "kind": "timer_fired", + "key": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a/timers/1", + "at": 1791100061000 + }, + "steps": [], + "patch": [ + { + "op": "remove", + "path": "/timers/armed/0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a~1timers~11" + }, + { + "op": "replace", + "path": "/machine/context", + "value": 1 + }, + { + "op": "replace", + "path": "/inputs", + "value": 3 + }, + { + "op": "replace", + "path": "/lastInputAt", + "value": 1791100061000 + }, + { + "op": "replace", + "path": "/historyBytes", + "value": 1926 + } + ], + "outputs": [] + } + } + ] + }, + "state": { + "executionId": "0199a3c4-7d2e-7c1a-9b3f-2f1e0d9c8b7a", + "status": "running", + "workflow": { + "document": { + "document": { + "dsl": "1.0.3", + "namespace": "acme", + "name": "triage", + "version": "1.0.0" + }, + "do": [] + }, + "input": 1 + }, + "attributes": {}, + "limits": { + "mostDurationMs": 1, + "longestCallMs": 1 + }, + "startedAt": 1791100060000, + "lastInputAt": 1791100061000, + "inputs": 3, + "random": { + "seed": 0, + "draws": 0 + }, + "runs": {}, + "timers": { + "next": 2, + "armed": {} + }, + "calls": {}, + "inbox": { + "waiting": [], + "waitingBytes": 0, + "receivedIds": [], + "received": 0, + "receivedBytes": 0, + "overflow": null + }, + "heldBytes": 0, + "historyBytes": 1926, + "stepsWithoutWaiting": 0, + "cancelRequested": false, + "machine": { + "values": { + "0": { + "value": {}, + "bytes": 2 + }, + "1": { + "value": { + "ticket": 7 + }, + "bytes": 12 + } + }, + "nextValue": 1, + "context": 1, + "root": null + }, + "outcome": null + } +} diff --git a/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts index 9cd60336c..35542fc50 100644 --- a/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts +++ b/packages/workflow-engine/src/dispatch/dispatch-watermark.test.ts @@ -25,7 +25,6 @@ const settle: RunOutput = { kind: 'settle', executionId, settlement: { status: ' function applied(version: number, outputs: readonly RunOutput[]): PositionedEvent { return { version, - bytes: 100, event: { type: 'input_applied', format: stateFormat, diff --git a/packages/workflow-engine/src/dispatch/run-due.test.ts b/packages/workflow-engine/src/dispatch/run-due.test.ts new file mode 100644 index 000000000..b70fe4d50 --- /dev/null +++ b/packages/workflow-engine/src/dispatch/run-due.test.ts @@ -0,0 +1,38 @@ +import { describe, expect, it } from 'vitest'; + +import { changesTimers, isTroubling, nextDueAtOf, runDueOf } from '../index.ts'; +import { at, executionId, runningState } from '../testing/runs.ts'; +import { exampleStream, streamOf } from '../testing/streams.ts'; + +describe('the next time a run is due', () => { + it('is the earliest of its armed timers, or none when nothing is armed', () => { + expect(nextDueAtOf(runningState)).toBe(at + 60_000); + expect(nextDueAtOf({ ...runningState, timers: { next: 3, armed: {} } })).toBeNull(); + }); + + it('is noted in the record by version, with whether the run fell behind its dispatch', () => { + expect(runDueOf(runningState, 7, true)).toEqual({ executionId, version: 7, nextDueAt: at + 60_000, behind: true }); + }); + + it('changes only with an event that arms or cancels a timer', () => { + const cancelling = streamOf([ + { + receipt: { kind: 'cancel_requested', key: executionId, at }, + patch: [], + outputs: [{ kind: 'cancel_timer', executionId, timerId: `${executionId}/timers/1` }], + }, + ]); + + expect([...exampleStream, ...cancelling].map((event) => changesTimers(event))).toEqual([false, true, false, true]); + }); +}); + +describe('a settle receipt', () => { + it('is troubling when the record was settled otherwise or names no execution, and is then reported', () => { + expect( + (['recorded', 'already_recorded', 'settled_otherwise', 'unknown_execution'] as const).map((receipt) => + isTroubling(receipt), + ), + ).toEqual([false, false, true, true]); + }); +}); diff --git a/packages/workflow-engine/src/dispatch/run-due.ts b/packages/workflow-engine/src/dispatch/run-due.ts new file mode 100644 index 000000000..45b21ba9b --- /dev/null +++ b/packages/workflow-engine/src/dispatch/run-due.ts @@ -0,0 +1,16 @@ +import type { RunState } from '../machine/run-state.ts'; +import type { PositionedEvent } from '../run-log/run-event.ts'; +import type { RunDue } from '../settlement/record-store.ts'; + +export function nextDueAtOf(state: RunState): number | null { + const dueTimes = Object.values(state.timers.armed).map(({ dueAt }) => dueAt); + return dueTimes.length === 0 ? null : Math.min(...dueTimes); +} + +export function changesTimers({ event }: PositionedEvent): boolean { + return event.outputs.some(({ kind }) => kind === 'arm_timer' || kind === 'cancel_timer'); +} + +export function runDueOf(state: RunState, version: number, behind: boolean): RunDue { + return { executionId: state.executionId, version, nextDueAt: nextDueAtOf(state), behind }; +} diff --git a/packages/workflow-engine/src/engine/run-loop.test.ts b/packages/workflow-engine/src/engine/run-loop.test.ts new file mode 100644 index 000000000..81f3b265b --- /dev/null +++ b/packages/workflow-engine/src/engine/run-loop.test.ts @@ -0,0 +1,71 @@ +import { Conflict } from '@beonauto/operations'; +import { Effect, Result } from 'effect'; +import { describe, expect, it } from 'vitest'; + +import { eventBytesOf, runLoopOf, snapshotOf, submissionOf, type RunInput } from '../index.ts'; +import { countingDecider, memoryRunStore } from '../testing/run-store.ts'; +import { at, executionId, runningState, started } from '../testing/runs.ts'; + +const cancelled: RunInput = { kind: 'cancel_requested', executionId, at: at + 1 }; + +function submitted(store: ReturnType, input: RunInput) { + return runLoopOf(store, countingDecider)(input.executionId, input).pipe( + Effect.map((decided) => submissionOf(decided, input)), + ); +} + +describe('the engine on the ledger loop', () => { + it('loads the run from its stream, decides, appends one event with the version it read, and answers applied', async () => { + const store = memoryRunStore(); + + const answers = await Effect.runPromise( + Effect.all([submitted(store, started), submitted(store, cancelled), submitted(store, cancelled)]), + ); + + expect(answers).toEqual([ + { outcome: 'applied', version: 1 }, + { outcome: 'applied', version: 2 }, + { outcome: 'stale', version: 2 }, + ]); + expect(store.events(executionId).map(({ version }) => version)).toEqual([1, 2]); + }); + + it('answers not_started for an input to a run whose start has not arrived, and appends nothing', async () => { + const store = memoryRunStore(); + + expect(await Effect.runPromise(submitted(store, cancelled))).toEqual({ outcome: 'not_started', version: 0 }); + expect(store.events(executionId)).toEqual([]); + }); + + it('counts the bytes of every event it appends in the state it folds', async () => { + const store = memoryRunStore(); + + const { state } = await Effect.runPromise(runLoopOf(store, countingDecider)(executionId, started)); + + expect(state.historyBytes).toBe(store.events(executionId).reduce((sum, { event }) => sum + eventBytesOf(event), 0)); + }); + + it('loads and decides again after a version conflict, and fails with the ledger Conflict after three more', async () => { + const recovered = memoryRunStore(3); + const lost = memoryRunStore(4); + + expect(await Effect.runPromise(submitted(recovered, started))).toEqual({ outcome: 'applied', version: 1 }); + expect(await Effect.runPromise(Effect.result(submitted(lost, started)))).toEqual( + Result.fail( + new Conflict({ detail: 'The state changed while the command was decided', kind: 'concurrent_change' }), + ), + ); + }); +}); + +describe('the run store the tests use', () => { + it('reads back the events after a version for dispatch, and takes a snapshot', async () => { + const store = memoryRunStore(); + await Effect.runPromise(Effect.all([submitted(store, started), submitted(store, cancelled)])); + + const after = await Effect.runPromise(store.eventsAfter(executionId, 1)); + + expect(after.map(({ version }) => version)).toEqual([2]); + await Effect.runPromise(store.saveSnapshot(snapshotOf(runningState, 2))); + }); +}); diff --git a/packages/workflow-engine/src/engine/run-loop.ts b/packages/workflow-engine/src/engine/run-loop.ts new file mode 100644 index 000000000..400324d0e --- /dev/null +++ b/packages/workflow-engine/src/engine/run-loop.ts @@ -0,0 +1,31 @@ +import { decisionLoop, type Decided, type DecisionLoop } from '@beonauto/ledger'; +import { Effect } from 'effect'; + +import { outcomeOf, staleReasonOf } from '../machine/admission.ts'; +import type { RunDecider } from '../machine/run-decider.ts'; +import type { RunInput } from '../machine/run-input.ts'; +import type { RunState } from '../machine/run-state.ts'; +import type { RunEvent } from '../run-log/run-event.ts'; +import { loadedRunOf, type LoadedRun } from '../run-log/run-fold.ts'; +import type { RunStore } from '../run-log/run-store.ts'; +import type { Submission } from './workflow-engine.ts'; + +export type RunDecision = Decided; + +export function runLoopOf( + runStore: RunStore, + decider: RunDecider, +): DecisionLoop { + return decisionLoop( + (executionId: string) => Effect.map(runStore.load(executionId), (stored) => loadedRunOf(stored)), + (executionId, events, expectedVersion) => + Effect.forEach(events, (event, index) => runStore.append(executionId, event, expectedVersion + index), { + discard: true, + }), + decider, + ); +} + +export function submissionOf({ events, loaded, version }: RunDecision, input: RunInput): Submission { + return { outcome: events.length > 0 ? 'applied' : outcomeOf(staleReasonOf(loaded.state, input)), version }; +} diff --git a/packages/workflow-engine/src/engine/workflow-engine.ts b/packages/workflow-engine/src/engine/workflow-engine.ts index 2d574e63e..fae2e7104 100644 --- a/packages/workflow-engine/src/engine/workflow-engine.ts +++ b/packages/workflow-engine/src/engine/workflow-engine.ts @@ -7,7 +7,7 @@ import type { SubmissionOutcome } from '../machine/admission.ts'; import type { RunInput } from '../machine/run-input.ts'; import type { RunStore } from '../run-log/run-store.ts'; import type { RunSerialiser } from '../serialisation/run-serialiser.ts'; -import type { RecordStore } from '../settlement/record-store.ts'; +import type { RecordStore, RunReporter } from '../settlement/record-store.ts'; import type { Timers } from '../timers/timers.ts'; export interface EnginePorts { @@ -16,6 +16,7 @@ export interface EnginePorts { readonly timers: Timers; readonly executor: Executor; readonly recordStore: RecordStore; + readonly reporter: RunReporter; readonly serialiser: RunSerialiser; } @@ -31,12 +32,11 @@ export interface Wake { export interface SweepReport { readonly runs: number; - readonly behind: number; readonly timersArmedAgain: number; } export interface WorkflowEngine { readonly submit: (input: RunInput) => Effect.Effect; readonly wake: (executionId: string) => Effect.Effect; - readonly sweep: () => Effect.Effect; + readonly sweep: (before: number) => Effect.Effect; } diff --git a/packages/workflow-engine/src/executor/executor.ts b/packages/workflow-engine/src/executor/executor.ts index 7cccf5325..96d7b1be1 100644 --- a/packages/workflow-engine/src/executor/executor.ts +++ b/packages/workflow-engine/src/executor/executor.ts @@ -3,7 +3,7 @@ import type { Effect } from 'effect'; import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark.ts'; import type { CancelCall, StartCall } from '../dispatch/run-output.ts'; -export type StartReceipt = 'started' | 'already_started' | 'refused_after_cancel'; +export type StartReceipt = 'started' | 'started_again' | 'running' | 'answered_again' | 'refused_after_cancel'; export type CallCancelReceipt = 'cancelled' | 'already_answered' | 'tombstoned'; diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts index 244779870..7b556c55b 100644 --- a/packages/workflow-engine/src/index.ts +++ b/packages/workflow-engine/src/index.ts @@ -15,6 +15,8 @@ export { type Settle, type StartCall, } from './dispatch/run-output.ts'; +export { changesTimers, nextDueAtOf, runDueOf } from './dispatch/run-due.ts'; +export { runLoopOf, submissionOf, type RunDecision } from './engine/run-loop.ts'; export type { EnginePorts, Submission, SweepReport, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; export type { CallCancelReceipt, Executor, StartReceipt } from './executor/executor.ts'; @@ -27,7 +29,14 @@ export { type SubmissionOutcome, } from './machine/admission.ts'; export { DslErrorSchema, type DslError } from './machine/dsl-error.ts'; -export { InputReceiptSchema, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; +export { + MissingValue, + heldBytesOf, + heldValueOf, + reachableValueIds, + withReachableValuesOnly, +} from './machine/held-values.ts'; +export { InputReceiptSchema, inputTimeOf, receiptOf, type InputReceipt } from './machine/input-receipt.ts'; export { InstantSchema, clampedAt } from './machine/instant.ts'; export { mostCallArgumentsBytes, @@ -82,24 +91,31 @@ export { StepSchema, eventBytesOf, fitsInOneEvent, + withHistoryBytes, type PositionedEvent, type RunEvent, type Step, } from './run-log/run-event.ts'; -export { StreamGap, evolveRun, loadedRunOf, type LoadedRun } from './run-log/run-fold.ts'; -export type { AppendedEvent, RunStore, StoredRun, StoredSnapshot } from './run-log/run-store.ts'; +export { UnreadableRun, evolveRun, loadedRunOf, stateInCurrentFormat, type LoadedRun } from './run-log/run-fold.ts'; +export type { RunStore, StoredRun, StoredSnapshot } from './run-log/run-store.ts'; export { SnapshotSchema, isSnapshotDue, mostSnapshotChunkBytes, snapshotChunks, snapshotEveryBytes, - snapshotEveryInputs, snapshotFromChunks, + snapshotOf, type SinceSnapshot, type Snapshot, } from './run-log/snapshot.ts'; -export { StateFormatSchema, stateFormat } from './run-log/state-format.ts'; +export { + StateFormatSchema, + stateFormat, + stateFormats, + type OlderFormat, + type StateFormats, +} from './run-log/state-format.ts'; export { PatchFailed, PatchOperationSchema, @@ -108,6 +124,15 @@ export { type StatePatch, } from './run-log/state-patch.ts'; export type { RunSerialiser } from './serialisation/run-serialiser.ts'; -export type { RecordStore, SettleReceipt, SettleRequest } from './settlement/record-store.ts'; +export { + isTroubling, + type RecordStore, + type RunDue, + type RunReporter, + type SettleReceipt, + type SettleRequest, + type TroublingReceipt, + type UnsettledReport, +} from './settlement/record-store.ts'; export { TimerPurposeSchema, timerIdOf, type TimerPurpose } from './timers/timer-id.ts'; export type { ArmReceipt, TimerCancelReceipt, Timers } from './timers/timers.ts'; diff --git a/packages/workflow-engine/src/machine/admission.test.ts b/packages/workflow-engine/src/machine/admission.test.ts index 3bd9b8f42..a190fc551 100644 --- a/packages/workflow-engine/src/machine/admission.test.ts +++ b/packages/workflow-engine/src/machine/admission.test.ts @@ -70,12 +70,15 @@ describe('a stale input', () => { }); describe('an input that belongs to no run like this one', () => { - it('dies instead of being taken as stale: another execution, or a start with another document', () => { + it('dies instead of being taken as stale: another execution, or a start with another document or input', () => { expect(() => staleReasonOf(runningState, { ...fired(armedTimer), executionId: 'another' })).toThrow(RunMismatch); expect(() => staleReasonOf(runningState, { ...started, executionId: 'another' })).toThrow(RunMismatch); expect(() => staleReasonOf(runningState, { ...started, document: { do: [{ other: {} }] } })).toThrow( new RunMismatch({ detail: `The run of ${executionId} was started again with another document` }), ); + expect(() => staleReasonOf(runningState, { ...started, input: { ticket: 8 } })).toThrow( + new RunMismatch({ detail: `The run of ${executionId} was started again with another input` }), + ); expect(() => staleReasonOf({ ...runningState, workflow: null }, started)).toThrow(RunMismatch); }); }); diff --git a/packages/workflow-engine/src/machine/admission.ts b/packages/workflow-engine/src/machine/admission.ts index 2f4eb401a..6ac1800af 100644 --- a/packages/workflow-engine/src/machine/admission.ts +++ b/packages/workflow-engine/src/machine/admission.ts @@ -1,6 +1,7 @@ import { Data } from 'effect'; import { callKeyText } from '../executor/call-key.ts'; +import { heldValueOf } from './held-values.ts'; import type { RunInput } from './run-input.ts'; import type { RunState } from './run-state.ts'; import { sameJson } from './same-json.ts'; @@ -32,9 +33,12 @@ function startedReason( return undefined; } requireSameRun(state, input); - if (!sameJson(state.workflow?.document ?? null, input.document)) { + if (state.workflow === null || !sameJson(state.workflow.document, input.document)) { throw new RunMismatch({ detail: `The run of ${state.executionId} was started again with another document` }); } + if (!sameJson(heldValueOf(state, state.workflow.input).value, input.input)) { + throw new RunMismatch({ detail: `The run of ${state.executionId} was started again with another input` }); + } return 'started_before'; } diff --git a/packages/workflow-engine/src/machine/held-values.test.ts b/packages/workflow-engine/src/machine/held-values.test.ts new file mode 100644 index 000000000..e372bf112 --- /dev/null +++ b/packages/workflow-engine/src/machine/held-values.test.ts @@ -0,0 +1,198 @@ +import { describe, expect, it } from 'vitest'; + +import { + heldBytesOf, + MissingValue, + newRun, + reachableValueIds, + taskFrameBytes, + withReachableValuesOnly, + type Branch, + type FrameBody, + type HeldValue, + type ListCursor, + type RunState, + type TaskFrame, +} from '../index.ts'; +import { runningState } from '../testing/runs.ts'; + +type Choose = (modulus: number) => number; + +const valueCount = 16; + +const valueKeys = new Set(['rawInput', 'input', 'data', 'items', 'output', 'arguments']); + +const utf8 = new TextEncoder(); + +function chooserOf(seed: number): Choose { + const drawn = { state: seed }; + return (modulus) => { + drawn.state = (drawn.state * 48_271) % 2_147_483_647; + return drawn.state % modulus; + }; +} + +function cursorOf(choose: Choose, depth: number): ListCursor { + const running = depth > 0 && choose(3) > 0; + const idle = choose(2) === 0 ? null : { kind: 'yielding' as const, timer: 't' }; + return { + pointer: '/do', + position: choose(4), + data: choose(valueCount), + variables: { item: choose(valueCount) }, + current: running ? { kind: 'running', task: frameOf(choose, depth - 1) } : idle, + }; +} + +function branchOf(choose: Choose, depth: number): Branch { + const kind = choose(3); + if (kind === 0) { + return { state: 'running', task: frameOf(choose, Math.max(depth - 1, 0)) }; + } + return kind === 1 + ? { state: 'finished', output: choose(valueCount), flow: 'continue' } + : { state: 'failed', error: { type: 'runtime', status: 500, instance: '/do/0' } }; +} + +function bodyOf(choose: Choose, depth: number): FrameBody { + const kinds: readonly (() => FrameBody)[] = [ + () => ({ kind: 'list', list: cursorOf(choose, depth) }), + () => ({ + kind: 'for', + items: choose(valueCount), + index: 0, + data: choose(valueCount), + list: choose(2) === 0 ? null : cursorOf(choose, depth), + }), + () => ({ kind: 'fork', compete: false, branches: [branchOf(choose, depth), branchOf(choose, depth)] }), + () => ({ + kind: 'try', + attempt: 1, + startedAt: 0, + phase: { kind: 'trying', list: cursorOf(choose, depth), attemptLimit: null }, + }), + () => ({ kind: 'try', attempt: 2, startedAt: 0, phase: { kind: 'recovering', list: cursorOf(choose, depth) } }), + () => ({ + kind: 'try', + attempt: 2, + startedAt: 0, + phase: { kind: 'backing_off', timer: 't', error: { type: 'runtime', status: 500, instance: '/do/0' } }, + }), + () => ({ kind: 'wait', timer: 't' }), + () => ({ + kind: 'call', + key: { executionId: 'e', reference: '/do/0', run: 1 }, + function: 'notify', + arguments: choose(valueCount), + label: 'notify', + }), + () => ({ kind: 'listen', consumed: [choose(valueCount), choose(valueCount)] }), + ]; + return kinds[choose(kinds.length)]?.() ?? { kind: 'wait', timer: 't' }; +} + +function frameOf(choose: Choose, depth: number): TaskFrame { + return { + reference: '/do/0', + run: 1, + rawInput: choose(valueCount), + input: choose(valueCount), + variables: { attempt: choose(valueCount) }, + timeout: null, + body: bodyOf(choose, depth), + }; +} + +function stateOf(seed: number): RunState { + const choose = chooserOf(seed); + const values: Record = Object.fromEntries( + Array.from({ length: valueCount }, (_, id) => [id, { value: `v${id}`, bytes: 10 + choose(1000) }]), + ); + return { + ...newRun, + executionId: 'e', + status: 'running', + workflow: { document: { do: [{ seed }] }, input: choose(valueCount) }, + machine: { values, nextValue: valueCount, context: choose(valueCount), root: frameOf(choose, 3) }, + }; +} + +function idsHeldIn(node: unknown): readonly number[] { + if (Array.isArray(node)) { + return node.flatMap((item: unknown) => idsHeldIn(item)); + } + if (typeof node !== 'object' || node === null) { + return []; + } + return Object.entries(node).flatMap(([key, item]: readonly [string, unknown]) => { + if (valueKeys.has(key) && typeof item === 'number') { + return [item]; + } + if ((key === 'variables' || key === 'consumed') && typeof item === 'object' && item !== null) { + return Object.values(item).filter((id): id is number => typeof id === 'number'); + } + return idsHeldIn(item); + }); +} + +function framesIn(node: unknown): number { + if (Array.isArray(node)) { + return node.reduce((sum: number, item: unknown) => sum + framesIn(item), 0); + } + if (typeof node !== 'object' || node === null) { + return 0; + } + const own = Object.hasOwn(node, 'rawInput') ? 1 : 0; + return Object.values(node).reduce((sum: number, item: unknown) => sum + framesIn(item), own); +} + +function expectedHeldBytes(state: RunState): number { + const ids = new Set([ + state.machine.context, + ...(state.workflow === null ? [] : [state.workflow.input]), + ...idsHeldIn(state.machine.root), + ]); + const valueBytes = [...ids].reduce((sum, id) => sum + (state.machine.values[id]?.bytes ?? 0), 0); + const documentBytes = utf8.encode(JSON.stringify(state.workflow?.document)).byteLength; + return valueBytes + framesIn(state.machine.root) * taskFrameBytes + documentBytes; +} + +const seeds = Array.from({ length: 400 }, (_, index) => index + 1); + +describe('the data a run holds', () => { + it('is, for any state, the bytes of the values its frames, context and input reach, 4 KiB a frame, and the document', () => { + const states = seeds.map((seed) => stateOf(seed)); + + expect(states.map((state) => heldBytesOf(state))).toEqual(states.map((state) => expectedHeldBytes(state))); + }); + + it('keeps, after the sweep that follows each decision, exactly the values that are reached, and the same bytes', () => { + const swept = seeds.map((seed) => withReachableValuesOnly(stateOf(seed))); + + expect(swept.map((state) => Object.keys(state.machine.values).map(Number))).toEqual( + seeds.map((seed) => reachableValueIds(stateOf(seed))), + ); + expect(swept.map((state) => state.heldBytes)).toEqual(swept.map((state) => expectedHeldBytes(state))); + expect(swept.some((state) => Object.keys(state.machine.values).length < valueCount)).toBe(true); + }); + + it('of a new run is its empty context alone: no document, no input and no frame yet', () => { + expect([reachableValueIds(newRun), heldBytesOf(newRun)]).toEqual([[0], 2]); + }); + + it('lets go of a value nothing reaches any more, such as the answer of a call once it moved on', () => { + const answered = { ...runningState.machine.values, 7: { value: { answer: 1 }, bytes: 11 } }; + + const swept = withReachableValuesOnly({ ...runningState, machine: { ...runningState.machine, values: answered } }); + + expect(Object.keys(swept.machine.values)).toEqual(['0', '1', '2']); + }); + + it('dies on a value that is reached but missing, a state no decision may leave', () => { + const { 2: _approval, ...values } = runningState.machine.values; + + expect(() => heldBytesOf({ ...runningState, machine: { ...runningState.machine, values } })).toThrow( + new MissingValue({ value: 2 }), + ); + }); +}); diff --git a/packages/workflow-engine/src/machine/held-values.ts b/packages/workflow-engine/src/machine/held-values.ts new file mode 100644 index 000000000..6502e0d69 --- /dev/null +++ b/packages/workflow-engine/src/machine/held-values.ts @@ -0,0 +1,86 @@ +import { Data } from 'effect'; + +import { jsonBytesOf } from '../dsl/json.ts'; +import { taskFrameBytes } from './limits.ts'; +import type { Branch, FrameBody, HeldValue, ListCursor, RunState, TaskFrame, ValueId } from './run-state.ts'; + +interface Reach { + readonly values: readonly ValueId[]; + readonly frames: number; +} + +export class MissingValue extends Data.TaggedError('missing_value')<{ readonly value: ValueId }> {} + +const nothing: Reach = { values: [], frames: 0 }; + +function together(reaches: readonly Reach[]): Reach { + return { + values: reaches.flatMap(({ values }: Reach) => values), + frames: reaches.reduce((sum, { frames }: Reach) => sum + frames, 0), + }; +} + +function cursorReach(cursor: ListCursor): Reach { + const own: Reach = { values: [cursor.data, ...Object.values(cursor.variables)], frames: 0 }; + return cursor.current?.kind === 'running' ? together([own, frameReach(cursor.current.task)]) : own; +} + +function branchReach(branch: Branch): Reach { + if (branch.state === 'running') { + return frameReach(branch.task); + } + return branch.state === 'finished' ? { values: [branch.output], frames: 0 } : nothing; +} + +function bodyReach(body: FrameBody): Reach { + if (body.kind === 'list') { + return cursorReach(body.list); + } + if (body.kind === 'for') { + const own: Reach = { values: [body.items, body.data], frames: 0 }; + return body.list === null ? own : together([own, cursorReach(body.list)]); + } + if (body.kind === 'fork') { + return together(body.branches.map((branch: Branch) => branchReach(branch))); + } + if (body.kind === 'try') { + return body.phase.kind === 'backing_off' ? nothing : cursorReach(body.phase.list); + } + if (body.kind === 'call') { + return { values: [body.arguments], frames: 0 }; + } + return body.kind === 'listen' ? { values: body.consumed, frames: 0 } : nothing; +} + +function frameReach(frame: TaskFrame): Reach { + const own: Reach = { values: [frame.rawInput, frame.input, ...Object.values(frame.variables)], frames: 1 }; + return together([own, bodyReach(frame.body)]); +} + +function runReach({ machine, workflow }: RunState): Reach { + const roots: Reach = { values: workflow === null ? [machine.context] : [machine.context, workflow.input], frames: 0 }; + return machine.root === null ? roots : together([roots, frameReach(machine.root)]); +} + +export function heldValueOf(state: RunState, id: ValueId): HeldValue { + const held = Object.hasOwn(state.machine.values, id) ? state.machine.values[id] : undefined; + if (held === undefined) { + throw new MissingValue({ value: id }); + } + return held; +} + +export function reachableValueIds(state: RunState): readonly ValueId[] { + return [...new Set(runReach(state).values)].toSorted((first, second) => first - second); +} + +export function heldBytesOf(state: RunState): number { + const valueBytes = reachableValueIds(state).reduce((sum, id) => sum + heldValueOf(state, id).bytes, 0); + const documentBytes = state.workflow === null ? 0 : jsonBytesOf(state.workflow.document); + return valueBytes + runReach(state).frames * taskFrameBytes + documentBytes; +} + +export function withReachableValuesOnly(state: RunState): RunState { + const values = Object.fromEntries(reachableValueIds(state).map((id) => [id, heldValueOf(state, id)])); + return { ...state, heldBytes: heldBytesOf(state), machine: { ...state.machine, values } }; +} diff --git a/packages/workflow-engine/src/machine/input-receipt.test.ts b/packages/workflow-engine/src/machine/input-receipt.test.ts index aace3bc8a..d15f8c521 100644 --- a/packages/workflow-engine/src/machine/input-receipt.test.ts +++ b/packages/workflow-engine/src/machine/input-receipt.test.ts @@ -1,7 +1,7 @@ import { describe, expect, it } from 'vitest'; -import { callKeyText, clampedAt, receiptOf, type RunInput } from '../index.ts'; -import { armedTimer, at, executionId, openCall, started } from '../testing/runs.ts'; +import { callKeyText, clampedAt, inputTimeOf, receiptOf, type RunInput } from '../index.ts'; +import { armedTimer, at, executionId, openCall, runningState, started } from '../testing/runs.ts'; const inputs: readonly RunInput[] = [ started, @@ -13,7 +13,7 @@ const inputs: readonly RunInput[] = [ describe('the receipt of an input', () => { it('names the input by the key it is deduplicated by, with the status of an answer or the type of an event', () => { - expect(inputs.map((input) => receiptOf(input, 0))).toEqual([ + expect(inputs.map((input) => receiptOf(input, at))).toEqual([ { kind: 'started', key: executionId, at }, { kind: 'timer_fired', key: armedTimer, at }, { kind: 'call_answered', key: callKeyText(openCall), at, status: 'failed' }, @@ -21,9 +21,21 @@ describe('the receipt of an input', () => { { kind: 'cancel_requested', key: executionId, at }, ]); }); +}); + +describe('the time of an input', () => { + it('never goes back: an input whose clock is behind the last input takes the time of that input', () => { + const later = { ...runningState, lastInputAt: at + 5000 }; - it('never goes back in time: an input whose clock is behind the last input takes the time of that input', () => { - expect(receiptOf(started, at + 5000).at).toBe(at + 5000); + expect(inputTimeOf(later, started)).toBe(at + 5000); + expect(inputTimeOf(runningState, { kind: 'cancel_requested', executionId, at: at + 1 })).toBe(at + 1); expect([clampedAt(at, at - 1), clampedAt(at, at + 1)]).toEqual([at, at + 1]); }); + + it('is never before the time a fired timer was due, so a timer that fires early still fires at its time', () => { + const early: RunInput = { kind: 'timer_fired', executionId, at, timerId: armedTimer }; + const unknown: RunInput = { kind: 'timer_fired', executionId, at, timerId: `${executionId}/timers/9` }; + + expect([inputTimeOf(runningState, early), inputTimeOf(runningState, unknown)]).toEqual([at + 60_000, at]); + }); }); diff --git a/packages/workflow-engine/src/machine/input-receipt.ts b/packages/workflow-engine/src/machine/input-receipt.ts index 11860055c..5ff071ce1 100644 --- a/packages/workflow-engine/src/machine/input-receipt.ts +++ b/packages/workflow-engine/src/machine/input-receipt.ts @@ -3,6 +3,7 @@ import { Schema } from 'effect'; import { callKeyText } from '../executor/call-key.ts'; import { clampedAt, InstantSchema } from './instant.ts'; import type { RunInput } from './run-input.ts'; +import type { RunState } from './run-state.ts'; const keyed = { key: Schema.String, @@ -23,8 +24,16 @@ export const InputReceiptSchema = Schema.Union([ export type InputReceipt = typeof InputReceiptSchema.Type; -export function receiptOf(input: RunInput, lastInputAt: number): InputReceipt { - const at = clampedAt(lastInputAt, input.at); +export function inputTimeOf(state: RunState, input: RunInput): number { + const at = clampedAt(state.lastInputAt, input.at); + const armed = + input.kind === 'timer_fired' && Object.hasOwn(state.timers.armed, input.timerId) + ? state.timers.armed[input.timerId] + : undefined; + return armed === undefined ? at : Math.max(at, armed.dueAt); +} + +export function receiptOf(input: RunInput, at: number): InputReceipt { if (input.kind === 'timer_fired') { return { kind: input.kind, key: input.timerId, at }; } diff --git a/packages/workflow-engine/src/machine/run-state.ts b/packages/workflow-engine/src/machine/run-state.ts index 3585200a4..d38f93898 100644 --- a/packages/workflow-engine/src/machine/run-state.ts +++ b/packages/workflow-engine/src/machine/run-state.ts @@ -12,7 +12,6 @@ export type ValueId = number; export interface HeldValue { readonly value: Schema.Json; readonly bytes: number; - readonly holders: number; } export type Variables = Readonly>; @@ -54,8 +53,9 @@ export type FrameBody = | { readonly kind: 'call'; readonly key: CallKey; - readonly primitive: string | null; - readonly name: string | null; + readonly function: string; + readonly arguments: ValueId; + readonly label: string; } | { readonly kind: 'listen'; readonly consumed: readonly ValueId[] }; @@ -119,6 +119,7 @@ export interface RunState { readonly calls: Readonly>; readonly inbox: InboxState; readonly heldBytes: number; + readonly historyBytes: number; readonly stepsWithoutWaiting: number; readonly cancelRequested: boolean; readonly machine: MachineState; @@ -173,8 +174,9 @@ const FrameBodySchema: Schema.Codec = Schema.Union([ Schema.Struct({ kind: Schema.Literal('call'), key: CallKeySchema, - primitive: Schema.NullOr(Schema.String), - name: Schema.NullOr(Schema.String), + function: Schema.NonEmptyString, + arguments: ValueIdSchema, + label: Schema.String, }), Schema.Struct({ kind: Schema.Literal('listen'), consumed: Schema.Array(ValueIdSchema) }), ]); @@ -226,10 +228,11 @@ export const RunStateSchema: Schema.Codec = Schema.Struct({ overflow: Schema.NullOr(DslErrorSchema), }), heldBytes: IntSchema, + historyBytes: IntSchema, stepsWithoutWaiting: IntSchema, cancelRequested: Schema.Boolean, machine: Schema.Struct({ - values: Schema.Record(Schema.String, Schema.Struct({ value: Schema.Json, bytes: IntSchema, holders: IntSchema })), + values: Schema.Record(Schema.String, Schema.Struct({ value: Schema.Json, bytes: IntSchema })), nextValue: ValueIdSchema, context: ValueIdSchema, root: Schema.NullOr(TaskFrameSchema), @@ -252,8 +255,9 @@ export const newRun: RunState = { calls: {}, inbox: { waiting: [], waitingBytes: 0, receivedIds: [], received: 0, receivedBytes: 0, overflow: null }, heldBytes: 0, + historyBytes: 0, stepsWithoutWaiting: 0, cancelRequested: false, - machine: { values: { 0: { value: {}, bytes: 2, holders: 1 } }, nextValue: 1, context: 0, root: null }, + machine: { values: { 0: { value: {}, bytes: 2 } }, nextValue: 1, context: 0, root: null }, outcome: null, }; diff --git a/packages/workflow-engine/src/run-log/corpus.test.ts b/packages/workflow-engine/src/run-log/corpus.test.ts new file mode 100644 index 000000000..b9fb0cae3 --- /dev/null +++ b/packages/workflow-engine/src/run-log/corpus.test.ts @@ -0,0 +1,59 @@ +import { readdirSync, readFileSync } from 'node:fs'; +import { fileURLToPath } from 'node:url'; + +import { Result, Schema } from 'effect'; +import { describe, expect, it } from 'vitest'; + +import { + loadedRunOf, + RunEventSchema, + snapshotFromChunks, + stateFormats, + stateInCurrentFormat, + StateFormatSchema, +} from '../index.ts'; + +const StoredEventsSchema = Schema.Array(Schema.Struct({ version: Schema.Int, event: RunEventSchema })); + +const CorpusSchema = Schema.Struct({ + format: StateFormatSchema, + stream: StoredEventsSchema, + snapshot: Schema.Struct({ chunks: Schema.Array(Schema.String), bytes: Schema.Int, tail: StoredEventsSchema }), + state: Schema.Json, +}); + +type Corpus = typeof CorpusSchema.Type; + +const decodeCorpus = Schema.decodeUnknownSync(Schema.fromJsonString(CorpusSchema)); + +const directory = fileURLToPath(new URL('../../corpus/', import.meta.url)); + +const corpora: readonly Corpus[] = readdirSync(directory) + .filter((name) => name.endsWith('.json')) + .map((name) => decodeCorpus(readFileSync(`${directory}${name}`, 'utf8'))); + +function loadedBothWays({ stream, snapshot }: Corpus): readonly unknown[] { + const whole = loadedRunOf({ snapshot: null, tail: stream }); + const fromSnapshot = loadedRunOf({ + snapshot: { snapshot: Result.getOrThrow(snapshotFromChunks(snapshot.chunks)), bytes: snapshot.bytes }, + tail: snapshot.tail, + }); + return [whole.state, fromSnapshot.state, whole.version, fromSnapshot.version]; +} + +describe('the committed corpus of past state formats', () => { + it('holds a stream and a snapshot for every state format up to the current one', () => { + expect(corpora.map(({ format }) => format).toSorted((first, second) => first - second)).toEqual( + Array.from({ length: stateFormats.current }, (_, index) => index + 1), + ); + }); + + it('still loads, from the whole stream and from the snapshot and its tail, to the state it recorded', () => { + expect(corpora.map((corpus) => loadedBothWays(corpus))).toEqual( + corpora.map(({ format, state, stream }) => { + const recorded = stateInCurrentFormat(format, state); + return [recorded, recorded, stream.length, stream.length]; + }), + ); + }); +}); diff --git a/packages/workflow-engine/src/run-log/run-event.test.ts b/packages/workflow-engine/src/run-log/run-event.test.ts index 85ceb12a8..2dfa8d9fc 100644 --- a/packages/workflow-engine/src/run-log/run-event.test.ts +++ b/packages/workflow-engine/src/run-log/run-event.test.ts @@ -10,6 +10,7 @@ import { RunInputSchema, RunStateSchema, stateFormat, + withHistoryBytes, type RunEvent, type RunInput, } from '../index.ts'; @@ -47,6 +48,11 @@ const event: RunEvent = { ], }; +function countedAfter(before: number): readonly [unknown, unknown] { + const applied = withHistoryBytes(event, before); + return [applied.patch.at(-1), { op: 'replace', path: '/historyBytes', value: before + eventBytesOf(applied) }]; +} + function near(text: string): RunEvent { return { ...event, patch: [{ op: 'replace', path: '/machine/context', value: text }] }; } @@ -70,9 +76,9 @@ describe('a run event', () => { expect(readBack(RunEventSchema, asStored(RunEventSchema, event))).toEqual(Result.succeed(event)); }); - it('is refused when it names another state format, or holds an output the engine does not dispatch', () => { + it('is refused when it names no state format, or holds an output the engine does not dispatch', () => { expect([ - Result.isFailure(readBack(RunEventSchema, { ...event, format: stateFormat + 1 })), + Result.isFailure(readBack(RunEventSchema, { ...event, format: 0 })), Result.isFailure(readBack(RunEventSchema, { ...event, outputs: [{ kind: 'send_email', to: 'someone' }] })), ]).toEqual([true, true]); }); @@ -84,6 +90,15 @@ describe('a run event', () => { expect(fitsInOneEvent(near('x'.repeat(mostEventBytes - overhead)))).toBe(true); expect(fitsInOneEvent(near('x'.repeat(mostEventBytes - overhead + 1)))).toBe(false); }); + + it('sets the bytes of history to those before it and its own, which it knows before it is appended', () => { + const befores = [0, 1, 9_999_000, 99_999_000, 999_999_000, 9_999_999_000, 536_870_000]; + + const counted = befores.map((before) => countedAfter(before)); + + expect(counted.map(([last]) => last)).toEqual(counted.map(([, expected]) => expected)); + expect(withHistoryBytes(event, 0).patch.slice(0, -1)).toEqual(event.patch); + }); }); describe('a run input', () => { diff --git a/packages/workflow-engine/src/run-log/run-event.ts b/packages/workflow-engine/src/run-log/run-event.ts index fd8a6ac76..ee546ba39 100644 --- a/packages/workflow-engine/src/run-log/run-event.ts +++ b/packages/workflow-engine/src/run-log/run-event.ts @@ -27,7 +27,6 @@ export type RunEvent = typeof RunEventSchema.Type; export interface PositionedEvent { readonly version: number; - readonly bytes: number; readonly event: RunEvent; } @@ -40,3 +39,16 @@ export function eventBytesOf(event: RunEvent): number { export function fitsInOneEvent(event: RunEvent): boolean { return eventBytesOf(event) <= mostEventBytes; } + +function withHistoryBytesOf(event: RunEvent, historyBytes: number): RunEvent { + return { ...event, patch: [...event.patch, { op: 'replace', path: '/historyBytes', value: historyBytes }] }; +} + +export function withHistoryBytes(event: RunEvent, before: number): RunEvent { + const settled = (guess: number): RunEvent => { + const counted = withHistoryBytesOf(event, before + guess); + const bytes = eventBytesOf(counted); + return bytes === guess ? counted : settled(bytes); + }; + return settled(0); +} diff --git a/packages/workflow-engine/src/run-log/run-fold.test.ts b/packages/workflow-engine/src/run-log/run-fold.test.ts index 3a915e10b..7d04ddc64 100644 --- a/packages/workflow-engine/src/run-log/run-fold.test.ts +++ b/packages/workflow-engine/src/run-log/run-fold.test.ts @@ -1,97 +1,194 @@ +import { Schema } from 'effect'; import { describe, expect, it } from 'vitest'; import { + eventBytesOf, evolveRun, loadedRunOf, newRun, - stateFormat, - StreamGap, + snapshotOf, + stateInCurrentFormat, + UnreadableRun, + withHistoryBytes, + type OlderFormat, type PositionedEvent, type RunEvent, - type Snapshot, + type StateFormats, type StatePatch, } from '../index.ts'; -import { at, document, executionId } from '../testing/runs.ts'; - -const timer = `${executionId}/timers/1`; - -function applied(version: number, bytes: number, patch: StatePatch): PositionedEvent { - const event: RunEvent = { - type: 'input_applied', - format: stateFormat, - receipt: { kind: 'timer_fired', key: timer, at }, - steps: [], - patch, - outputs: [], - }; - return { version, bytes, event }; +import { at, executionId } from '../testing/runs.ts'; +import { exampleStream } from '../testing/streams.ts'; + +function bytesOf(events: readonly PositionedEvent[]): number { + return events.reduce((sum, { event }: PositionedEvent) => sum + eventBytesOf(event), 0); +} + +function eventIn(format: number, patch: StatePatch, before: number): RunEvent { + return withHistoryBytes( + { + type: 'input_applied', + format, + receipt: { kind: 'cancel_requested', key: executionId, at }, + steps: [], + patch, + outputs: [], + }, + before, + ); +} + +function streamIn(changes: readonly (readonly [number, StatePatch])[]): readonly PositionedEvent[] { + return changes.reduce( + (events: readonly PositionedEvent[], [format, patch]: readonly [number, StatePatch], index) => [ + ...events, + { version: index + 1, event: eventIn(format, patch, bytesOf(events)) }, + ], + [], + ); } -const startedRun = applied(1, 900, [ +const { inputs: _inputs, ...withoutInputs } = newRun; + +function isCounting(state: unknown): state is { readonly applied: number } { + return typeof state === 'object' && state !== null && typeof Reflect.get(state, 'applied') === 'number'; +} + +const countingApplied: OlderFormat = { + format: 1, + initial: { ...withoutInputs, applied: 0 }, + read: (state) => { + if (!isCounting(state)) { + throw new TypeError('A state of format 1 counts its inputs as applied'); + } + return state; + }, + upcast: (state) => { + const { applied, ...rest } = isCounting(state) ? state : { applied: 0 }; + return { ...rest, inputs: applied }; + }, +}; + +const twoFormats: StateFormats = { current: 2, older: [countingApplied] }; + +const started: StatePatch = [ { op: 'replace', path: '/executionId', value: executionId }, { op: 'replace', path: '/status', value: 'running' }, - { op: 'replace', path: '/workflow', value: { document, input: 1 } }, - { op: 'add', path: '/machine/values/1', value: { value: { ticket: 7 }, bytes: 12, holders: 1 } }, - { op: 'replace', path: '/inputs', value: 1 }, - { op: 'replace', path: '/startedAt', value: at }, - { op: 'replace', path: '/lastInputAt', value: at }, -]); - -const armed = applied(2, 300, [ - { - op: 'add', - path: `/timers/armed/${timer.replaceAll('/', '~1')}`, - value: { purpose: 'wait', reference: '/do/0', dueAt: at + 1000 }, - }, - { op: 'replace', path: '/timers/next', value: 2 }, - { op: 'replace', path: '/inputs', value: 2 }, -]); +]; -const fired = applied(3, 200, [ - { op: 'remove', path: `/timers/armed/${timer.replaceAll('/', '~1')}` }, - { op: 'replace', path: '/machine/context', value: 1 }, - { op: 'replace', path: '/inputs', value: 3 }, - { op: 'replace', path: '/lastInputAt', value: at + 1000 }, -]); +function patched(patch: StatePatch): RunEvent { + return eventIn(1, patch, 0); +} describe('a run loaded from its stream', () => { it('is the fold of its events from a new run, with its version and the bytes of its history', () => { - const loaded = loadedRunOf({ snapshot: null, tail: [startedRun, armed, fired] }); + const loaded = loadedRunOf({ snapshot: null, tail: exampleStream }); expect(loaded).toMatchObject({ version: 3, - historyBytes: 1400, - sinceSnapshot: { inputs: 3, bytes: 1400, snapshotBytes: 0 }, - state: { executionId, status: 'running', inputs: 3, lastInputAt: at + 1000, timers: { next: 2, armed: {} } }, + sinceSnapshot: { bytes: bytesOf(exampleStream), snapshotBytes: 0 }, + state: { executionId, status: 'running', inputs: 3, historyBytes: bytesOf(exampleStream), timers: { armed: {} } }, }); expect(loadedRunOf({ snapshot: null, tail: [] }).state).toEqual(newRun); }); +}); +describe('a run loaded from a snapshot', () => { it('is the same from its latest snapshot and the events after it as from its whole stream', () => { - const atTwo = loadedRunOf({ snapshot: null, tail: [startedRun, armed] }); - const snapshot: Snapshot = { - format: stateFormat, - executionId, - version: 2, - historyBytes: atTwo.historyBytes, - state: atTwo.state, - }; + const atTwo = loadedRunOf({ snapshot: null, tail: exampleStream.slice(0, 2) }); + const tail = exampleStream.slice(2); - const fromSnapshot = loadedRunOf({ snapshot: { snapshot, bytes: 5000 }, tail: [fired] }); + const fromSnapshot = loadedRunOf({ snapshot: { snapshot: snapshotOf(atTwo.state, 2), bytes: 5000 }, tail }); expect(fromSnapshot).toEqual({ - ...loadedRunOf({ snapshot: null, tail: [startedRun, armed, fired] }), - sinceSnapshot: { inputs: 1, bytes: 200, snapshotBytes: 5000 }, + ...loadedRunOf({ snapshot: null, tail: exampleStream }), + sinceSnapshot: { bytes: bytesOf(tail), snapshotBytes: 5000 }, }); - expect(evolveRun(atTwo.state, fired.event)).toEqual(fromSnapshot.state); + expect(tail.reduce((state, { event }: PositionedEvent) => evolveRun(state, event), atTwo.state)).toEqual( + fromSnapshot.state, + ); + }); +}); + +describe('a run that cannot be read', () => { + it('dies on a gap in its events, or on events whose sizes are not the bytes of history the state counts', () => { + const afterAGap = streamIn([ + [1, started], + [1, []], + ]).slice(1); + const miscounted = streamIn([[1, started]]).map(({ version, event }) => ({ + version, + event: { + ...event, + patch: [...event.patch.slice(0, -1), { op: 'replace' as const, path: '/historyBytes', value: 1 }], + }, + })); + + expect(() => loadedRunOf({ snapshot: null, tail: afterAGap })).toThrow( + new UnreadableRun({ detail: 'Event 1 is missing; the next event read is 2' }), + ); + expect(() => loadedRunOf({ snapshot: null, tail: miscounted })).toThrow(UnreadableRun); + }); + + it('dies on a patch that leaves a state this format does not describe, an unknown member included', () => { + expect(() => evolveRun(newRun, patched([{ op: 'replace', path: '/status', value: 'paused' }]))).toThrow( + /Expected "new" \| "running" \| "ended"/u, + ); + expect(() => evolveRun(newRun, patched([{ op: 'add', path: '/machine/extra', value: 1 }]))).toThrow(/extra/u); + }); +}); + +describe('the state formats of a stream', () => { + it('fold each event under its own format and upcast the state where the format changes', () => { + const stream = streamIn([ + [1, [...started, { op: 'replace', path: '/applied', value: 1 }]], + [1, [{ op: 'replace', path: '/applied', value: 2 }]], + [2, [{ op: 'replace', path: '/inputs', value: 3 }]], + ]); + + expect(loadedRunOf({ snapshot: null, tail: stream }, twoFormats).state).toMatchObject({ + inputs: 3, + status: 'running', + }); + expect(loadedRunOf({ snapshot: null, tail: stream.slice(0, 2) }, twoFormats).state).toMatchObject({ inputs: 2 }); + }); + + it('upcast a snapshot of an older format before the events after it', () => { + const olderState = Schema.decodeUnknownSync(Schema.Json)({ ...withoutInputs, applied: 4 }); + const snapshot = { format: 1, executionId, version: 4, historyBytes: 0, state: olderState }; + const tail = streamIn([[2, [{ op: 'replace', path: '/inputs', value: 5 }]]]).map(({ event }) => ({ + version: 5, + event, + })); + + expect(loadedRunOf({ snapshot: { snapshot, bytes: 100 }, tail }, twoFormats).state).toMatchObject({ inputs: 5 }); + expect(stateInCurrentFormat(1, olderState, twoFormats)).toMatchObject({ inputs: 4 }); + }); +}); + +describe('a state format this code does not read', () => { + it('is refused: formats never go back within a stream, and a format must be one this code knows', () => { + const goingBack = streamIn([ + [2, started], + [1, []], + ]); + const newer = streamIn([[3, started]]); + const unknown = streamIn([[1, started]]); + + expect(() => loadedRunOf({ snapshot: null, tail: goingBack }, twoFormats)).toThrow( + new UnreadableRun({ detail: 'State format 1 follows format 2; formats never go back' }), + ); + expect(() => loadedRunOf({ snapshot: null, tail: newer }, twoFormats)).toThrow( + new UnreadableRun({ detail: 'State format 3 is newer than 2, the newest this code reads' }), + ); + expect(() => loadedRunOf({ snapshot: null, tail: unknown }, { current: 2, older: [] })).toThrow( + new UnreadableRun({ detail: 'This code reads no state of format 1' }), + ); + expect(() => evolveRun(newRun, eventIn(2, [], 0))).toThrow(UnreadableRun); }); - it('dies on a gap in its events, and on a patch that leaves a state this format does not describe', () => { - expect(() => loadedRunOf({ snapshot: null, tail: [startedRun, fired] })).toThrow( - new StreamGap({ expected: 2, found: 3 }), + it('read a state of an older format strictly before upcasting it', () => { + expect(() => stateInCurrentFormat(1, { ...withoutInputs, inputs: 4 }, twoFormats)).toThrow( + 'A state of format 1 counts its inputs as applied', ); - expect(() => - evolveRun(newRun, applied(1, 10, [{ op: 'replace', path: '/status', value: 'paused' }]).event), - ).toThrow(/Expected "new" \| "running" \| "ended"/u); }); }); diff --git a/packages/workflow-engine/src/run-log/run-fold.ts b/packages/workflow-engine/src/run-log/run-fold.ts index c1cb6efdd..1f8492bc3 100644 --- a/packages/workflow-engine/src/run-log/run-fold.ts +++ b/packages/workflow-engine/src/run-log/run-fold.ts @@ -1,67 +1,109 @@ import { Data, Schema } from 'effect'; import { newRun, RunStateSchema, type RunState } from '../machine/run-state.ts'; -import type { PositionedEvent, RunEvent } from './run-event.ts'; +import { eventBytesOf, type PositionedEvent, type RunEvent } from './run-event.ts'; import type { StoredRun } from './run-store.ts'; import type { SinceSnapshot } from './snapshot.ts'; +import { stateFormats, type OlderFormat, type StateFormats } from './state-format.ts'; import { applyStatePatch } from './state-patch.ts'; export interface LoadedRun { readonly state: RunState; readonly version: number; - readonly historyBytes: number; readonly sinceSnapshot: SinceSnapshot; } -export class StreamGap extends Data.TaggedError('stream_gap')<{ - readonly expected: number; - readonly found: number; -}> {} +export class UnreadableRun extends Data.TaggedError('unreadable_run')<{ readonly detail: string }> {} -interface FoldStart { - readonly state: RunState; +interface Folding { + readonly format: number; + readonly state: unknown; +} + +interface FoldStart extends Folding { readonly version: number; readonly historyBytes: number; readonly snapshotBytes: number; } -const decodeState = Schema.decodeUnknownSync(RunStateSchema); +const decodeState = Schema.decodeUnknownSync(RunStateSchema, { onExcessProperty: 'error' }); -export function evolveRun(state: RunState, event: RunEvent): RunState { - return decodeState(applyStatePatch(state, event.patch)); +function olderFormatOf(formats: StateFormats, format: number): OlderFormat { + const older = formats.older.find((candidate: OlderFormat) => candidate.format === format); + if (older === undefined) { + throw new UnreadableRun({ detail: `This code reads no state of format ${format}` }); + } + return older; +} + +function initialStateOf(formats: StateFormats, format: number): unknown { + return format === formats.current ? newRun : olderFormatOf(formats, format).initial; +} + +function upcastTo(formats: StateFormats, folding: Folding, format: number): Folding { + if (format > formats.current) { + throw new UnreadableRun({ + detail: `State format ${format} is newer than ${formats.current}, the newest this code reads`, + }); + } + if (format < folding.format) { + throw new UnreadableRun({ + detail: `State format ${format} follows format ${folding.format}; formats never go back`, + }); + } + if (format === folding.format) { + return folding; + } + const older = olderFormatOf(formats, folding.format); + return upcastTo(formats, { format: folding.format + 1, state: older.upcast(older.read(folding.state)) }, format); } -function startOf({ snapshot }: StoredRun): FoldStart { - return snapshot === null - ? { state: newRun, version: 0, historyBytes: 0, snapshotBytes: 0 } - : { - state: snapshot.snapshot.state, - version: snapshot.snapshot.version, - historyBytes: snapshot.snapshot.historyBytes, - snapshotBytes: snapshot.bytes, - }; +function startOf({ snapshot, tail }: StoredRun, formats: StateFormats): FoldStart { + if (snapshot === null) { + const format = Math.min(tail[0]?.event.format ?? formats.current, formats.current); + return { format, state: initialStateOf(formats, format), version: 0, historyBytes: 0, snapshotBytes: 0 }; + } + return { ...snapshot.snapshot, snapshotBytes: snapshot.bytes }; } function requireContiguous(tail: readonly PositionedEvent[], after: number): void { for (const [index, { version }] of tail.entries()) { if (version !== after + index + 1) { - throw new StreamGap({ expected: after + index + 1, found: version }); + throw new UnreadableRun({ detail: `Event ${after + index + 1} is missing; the next event read is ${version}` }); } } } -export function loadedRunOf(stored: StoredRun): LoadedRun { - const start = startOf(stored); +function requireHistoryBytes(state: RunState, expected: number): void { + if (state.historyBytes !== expected) { + throw new UnreadableRun({ + detail: `The state counts ${state.historyBytes} bytes of history, and its events take ${expected}`, + }); + } +} + +export function stateInCurrentFormat(format: number, state: unknown, formats: StateFormats = stateFormats): RunState { + return decodeState(upcastTo(formats, { format, state }, formats.current).state); +} + +export function evolveRun(state: RunState, event: RunEvent): RunState { + const { state: patched } = upcastTo(stateFormats, { format: stateFormats.current, state }, event.format); + return decodeState(applyStatePatch(patched, event.patch)); +} + +export function loadedRunOf(stored: StoredRun, formats: StateFormats = stateFormats): LoadedRun { + const start = startOf(stored, formats); requireContiguous(stored.tail, start.version); - const folded = stored.tail.reduce( - (document: unknown, { event }: PositionedEvent) => applyStatePatch(document, event.patch), - start.state, - ); - const tailBytes = stored.tail.reduce((sum, { bytes }: PositionedEvent) => sum + bytes, 0); + const folded = stored.tail.reduce((folding: Folding, { event }: PositionedEvent): Folding => { + const upcast = upcastTo(formats, folding, event.format); + return { format: upcast.format, state: applyStatePatch(upcast.state, event.patch) }; + }, start); + const state = stateInCurrentFormat(folded.format, folded.state, formats); + const tailBytes = stored.tail.reduce((sum, { event }: PositionedEvent) => sum + eventBytesOf(event), 0); + requireHistoryBytes(state, start.historyBytes + tailBytes); return { - state: decodeState(folded), + state, version: start.version + stored.tail.length, - historyBytes: start.historyBytes + tailBytes, - sinceSnapshot: { inputs: stored.tail.length, bytes: tailBytes, snapshotBytes: start.snapshotBytes }, + sinceSnapshot: { bytes: tailBytes, snapshotBytes: start.snapshotBytes }, }; } diff --git a/packages/workflow-engine/src/run-log/run-store.ts b/packages/workflow-engine/src/run-log/run-store.ts index c38cdbf14..bccb5d15f 100644 --- a/packages/workflow-engine/src/run-log/run-store.ts +++ b/packages/workflow-engine/src/run-log/run-store.ts @@ -14,18 +14,13 @@ export interface StoredRun { readonly tail: readonly PositionedEvent[]; } -export interface AppendedEvent { - readonly version: number; - readonly bytes: number; -} - export interface RunStore { readonly load: (executionId: string) => Effect.Effect; readonly append: ( executionId: string, event: RunEvent, expectedVersion: number, - ) => Effect.Effect; + ) => Effect.Effect; readonly eventsAfter: (executionId: string, version: number) => Effect.Effect; readonly saveSnapshot: (snapshot: Snapshot) => Effect.Effect; } diff --git a/packages/workflow-engine/src/run-log/snapshot.test.ts b/packages/workflow-engine/src/run-log/snapshot.test.ts index bb0eabfd3..9afa86d11 100644 --- a/packages/workflow-engine/src/run-log/snapshot.test.ts +++ b/packages/workflow-engine/src/run-log/snapshot.test.ts @@ -6,10 +6,12 @@ import { mostSnapshotChunkBytes, snapshotChunks, snapshotFromChunks, + snapshotOf, stateFormat, + stateInCurrentFormat, type Snapshot, } from '../index.ts'; -import { executionId, runningState } from '../testing/runs.ts'; +import { runningState } from '../testing/runs.ts'; const highSurrogate = /[\uD800-\uDBFF]$/u; @@ -18,37 +20,31 @@ const lowSurrogate = /^[\uDC00-\uDFFF]/u; const utf8 = new TextEncoder(); function snapshotHolding(text: string): Snapshot { - const values = { ...runningState.machine.values, 0: { value: text, bytes: text.length, holders: 1 } }; - return { - format: stateFormat, - executionId, - version: 4000, - historyBytes: 9_000_000, - state: { ...runningState, machine: { ...runningState.machine, values } }, - }; + const values = { ...runningState.machine.values, 0: { value: text, bytes: text.length } }; + return snapshotOf({ ...runningState, machine: { ...runningState.machine, values } }, 4000); } describe('a snapshot', () => { - it('is due after 1,000 inputs, or after as many bytes of events as the last snapshot took, and at least 1 MiB', () => { + it('is due once the events since the last snapshot take as many bytes as it did, and at least 1 MiB', () => { expect([ - isSnapshotDue({ inputs: 999, bytes: 1_048_575, snapshotBytes: 0 }), - isSnapshotDue({ inputs: 1000, bytes: 0, snapshotBytes: 0 }), - isSnapshotDue({ inputs: 1, bytes: 1_048_576, snapshotBytes: 900_000 }), - isSnapshotDue({ inputs: 1, bytes: 2_000_000, snapshotBytes: 3_000_000 }), - isSnapshotDue({ inputs: 1, bytes: 3_000_000, snapshotBytes: 3_000_000 }), - ]).toEqual([false, true, true, false, true]); + isSnapshotDue({ bytes: 1_048_575, snapshotBytes: 0 }), + isSnapshotDue({ bytes: 1_048_576, snapshotBytes: 900_000 }), + isSnapshotDue({ bytes: 2_000_000, snapshotBytes: 3_000_000 }), + isSnapshotDue({ bytes: 3_000_000, snapshotBytes: 3_000_000 }), + ]).toEqual([false, true, false, true]); }); - it('round-trips a running state through its chunks, with the bytes of history it covers', () => { - const snapshot: Snapshot = { + it('names its state format and the bytes of history it covers, and round-trips through its chunks', () => { + const snapshot = snapshotOf(runningState, 1000); + + expect(snapshot).toMatchObject({ format: stateFormat, - executionId, + executionId: runningState.executionId, version: 1000, - historyBytes: 2_100_000, - state: runningState, - }; - + historyBytes: 9000, + }); expect(snapshotFromChunks(snapshotChunks(snapshot))).toEqual(Result.succeed(snapshot)); + expect(stateInCurrentFormat(snapshot.format, snapshot.state)).toEqual(runningState); }); it('is stored in chunks of at most 1 MiB of UTF-8, well under the 2 MB a row takes, that never split a character', () => { @@ -65,11 +61,8 @@ describe('a snapshot', () => { expect(snapshotFromChunks(chunks)).toEqual(Result.succeed(snapshot)); }); - it('is refused when its chunks do not make a snapshot of this state format', () => { - const snapshot: Snapshot = { format: stateFormat, executionId, version: 1, historyBytes: 10, state: runningState }; - const text = snapshotChunks(snapshot) - .join('') - .replace(`"format":${stateFormat}`, `"format":${stateFormat + 1}`); + it('is refused when its chunks do not make a snapshot', () => { + const text = snapshotChunks(snapshotOf(runningState, 1)).join('').replace(`"format":${stateFormat}`, '"format":0'); expect(Result.isFailure(snapshotFromChunks([text]))).toBe(true); }); diff --git a/packages/workflow-engine/src/run-log/snapshot.ts b/packages/workflow-engine/src/run-log/snapshot.ts index c55932202..42b4b2acb 100644 --- a/packages/workflow-engine/src/run-log/snapshot.ts +++ b/packages/workflow-engine/src/run-log/snapshot.ts @@ -1,9 +1,7 @@ import { Schema } from 'effect'; -import { RunStateSchema } from '../machine/run-state.ts'; -import { StateFormatSchema } from './state-format.ts'; - -export const snapshotEveryInputs = 1000; +import { RunStateSchema, type RunState } from '../machine/run-state.ts'; +import { stateFormat, StateFormatSchema } from './state-format.ts'; export const snapshotEveryBytes = 1_048_576; @@ -14,13 +12,12 @@ export const SnapshotSchema = Schema.Struct({ executionId: Schema.NonEmptyString, version: Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)), historyBytes: Schema.Int.check(Schema.isGreaterThanOrEqualTo(0)), - state: RunStateSchema, + state: Schema.Json, }); export type Snapshot = typeof SnapshotSchema.Type; export interface SinceSnapshot { - readonly inputs: number; readonly bytes: number; readonly snapshotBytes: number; } @@ -31,38 +28,35 @@ const encodeSnapshot = Schema.encodeSync(SnapshotTextSchema); const decodeSnapshot = Schema.decodeUnknownResult(SnapshotTextSchema); -export function isSnapshotDue({ inputs, bytes, snapshotBytes }: SinceSnapshot): boolean { - return inputs >= snapshotEveryInputs || bytes >= Math.max(snapshotEveryBytes, snapshotBytes); +const encodeState = Schema.encodeSync(Schema.toCodecJson(RunStateSchema)); + +const utf8 = new TextEncoder(); + +export function snapshotOf(state: RunState, version: number): Snapshot { + return { + format: stateFormat, + executionId: state.executionId, + version, + historyBytes: state.historyBytes, + state: encodeState(state), + }; } -function utf8LengthOf(codePoint: number): number { - if (codePoint < 0x80) { - return 1; - } - if (codePoint < 0x8_00) { - return 2; - } - return codePoint < 0x1_00_00 ? 3 : 4; +export function isSnapshotDue({ bytes, snapshotBytes }: SinceSnapshot): boolean { + return bytes >= Math.max(snapshotEveryBytes, snapshotBytes); } export function snapshotChunks(snapshot: Snapshot): readonly string[] { const text = encodeSnapshot(snapshot); - const chunks: string[] = []; - let start = 0; - let end = 0; - let bytes = 0; - for (const character of text) { - const length = utf8LengthOf(Number(character.codePointAt(0))); - if (bytes + length > mostSnapshotChunkBytes) { - chunks.push(text.slice(start, end)); - start = end; - bytes = 0; + const buffer = new Uint8Array(mostSnapshotChunkBytes); + const chunksFrom = (start: number): readonly string[] => { + if (start >= text.length) { + return []; } - bytes += length; - end += character.length; - } - chunks.push(text.slice(start)); - return chunks; + const { read } = utf8.encodeInto(text.slice(start), buffer); + return [text.slice(start, start + read), ...chunksFrom(start + read)]; + }; + return chunksFrom(0); } export function snapshotFromChunks(chunks: readonly string[]): ReturnType { diff --git a/packages/workflow-engine/src/run-log/state-format.ts b/packages/workflow-engine/src/run-log/state-format.ts index 6149e1507..d7359e6e7 100644 --- a/packages/workflow-engine/src/run-log/state-format.ts +++ b/packages/workflow-engine/src/run-log/state-format.ts @@ -2,4 +2,18 @@ import { Schema } from 'effect'; export const stateFormat = 1; -export const StateFormatSchema = Schema.Literal(stateFormat); +export const StateFormatSchema = Schema.Int.check(Schema.isGreaterThanOrEqualTo(1)); + +export interface OlderFormat { + readonly format: number; + readonly initial: unknown; + readonly read: (state: unknown) => unknown; + readonly upcast: (state: unknown) => unknown; +} + +export interface StateFormats { + readonly current: number; + readonly older: readonly OlderFormat[]; +} + +export const stateFormats: StateFormats = { current: stateFormat, older: [] }; diff --git a/packages/workflow-engine/src/settlement/record-store.ts b/packages/workflow-engine/src/settlement/record-store.ts index 2d35f35f8..4d58c12c0 100644 --- a/packages/workflow-engine/src/settlement/record-store.ts +++ b/packages/workflow-engine/src/settlement/record-store.ts @@ -5,12 +5,35 @@ import type { DispatchFailed, RunContext } from '../dispatch/dispatch-watermark. export type SettleReceipt = 'recorded' | 'already_recorded' | 'settled_otherwise' | 'unknown_execution'; +export type TroublingReceipt = Extract; + export interface SettleRequest { readonly executionId: string; readonly settlement: Settlement; } +export interface RunDue { + readonly executionId: string; + readonly version: number; + readonly nextDueAt: number | null; + readonly behind: boolean; +} + export interface RecordStore { readonly settle: (request: SettleRequest, run: RunContext) => Effect.Effect; - readonly liveRuns: () => Effect.Effect; + readonly noteDue: (due: RunDue, run: RunContext) => Effect.Effect; + readonly dueRuns: (before: number) => Effect.Effect; +} + +export interface UnsettledReport { + readonly run: RunContext; + readonly receipt: TroublingReceipt; +} + +export interface RunReporter { + readonly unsettled: (report: UnsettledReport) => Effect.Effect; +} + +export function isTroubling(receipt: SettleReceipt): receipt is TroublingReceipt { + return receipt === 'settled_otherwise' || receipt === 'unknown_execution'; } diff --git a/packages/workflow-engine/src/testing/run-store.ts b/packages/workflow-engine/src/testing/run-store.ts new file mode 100644 index 000000000..a81832548 --- /dev/null +++ b/packages/workflow-engine/src/testing/run-store.ts @@ -0,0 +1,66 @@ +import { VersionConflict } from '@beonauto/ledger'; +import { Effect, Result } from 'effect'; + +import { staleReasonOf } from '../machine/admission.ts'; +import { inputTimeOf, receiptOf } from '../machine/input-receipt.ts'; +import type { RunDecider } from '../machine/run-decider.ts'; +import { newRun } from '../machine/run-state.ts'; +import { RunEventSchema, withHistoryBytes, type PositionedEvent } from '../run-log/run-event.ts'; +import { evolveRun } from '../run-log/run-fold.ts'; +import type { RunStore } from '../run-log/run-store.ts'; +import { stateFormat } from '../run-log/state-format.ts'; + +export interface MemoryRunStore extends RunStore { + readonly events: (executionId: string) => readonly PositionedEvent[]; +} + +export function memoryRunStore(conflicts = 0): MemoryRunStore { + const streams = new Map(); + const remaining = { conflicts }; + const eventsOf = (executionId: string): readonly PositionedEvent[] => streams.get(executionId) ?? []; + return { + load: (executionId) => Effect.sync(() => ({ snapshot: null, tail: eventsOf(executionId) })), + append: (executionId, event, expectedVersion) => + Effect.suspend(() => { + remaining.conflicts -= 1; + if (remaining.conflicts >= 0 || eventsOf(executionId).length !== expectedVersion) { + return Effect.fail(new VersionConflict()); + } + streams.set(executionId, [...eventsOf(executionId), { version: expectedVersion + 1, event }]); + return Effect.void; + }), + eventsAfter: (executionId, version) => Effect.sync(() => eventsOf(executionId).slice(version)), + saveSnapshot: () => Effect.void, + events: eventsOf, + }; +} + +export const countingDecider: RunDecider = { + initialState: newRun, + evolve: evolveRun, + eventSchema: RunEventSchema, + decide: (input, state) => { + if (staleReasonOf(state, input) !== undefined) { + return Result.succeed([]); + } + const at = inputTimeOf(state, input); + const counted = withHistoryBytes( + { + type: 'input_applied', + format: stateFormat, + receipt: receiptOf(input, at), + steps: [], + patch: [ + { op: 'replace', path: '/executionId', value: input.executionId }, + { op: 'replace', path: '/status', value: 'running' }, + { op: 'replace', path: '/inputs', value: state.inputs + 1 }, + { op: 'replace', path: '/lastInputAt', value: at }, + { op: 'replace', path: '/cancelRequested', value: input.kind === 'cancel_requested' }, + ], + outputs: [], + }, + state.historyBytes, + ); + return Result.succeed([counted]); + }, +}; diff --git a/packages/workflow-engine/src/testing/runs.ts b/packages/workflow-engine/src/testing/runs.ts index 9fcc6af72..782a0d4a6 100644 --- a/packages/workflow-engine/src/testing/runs.ts +++ b/packages/workflow-engine/src/testing/runs.ts @@ -1,4 +1,5 @@ import { callKeyText, type CallKey } from '../executor/call-key.ts'; +import { heldBytesOf } from '../machine/held-values.ts'; import type { Started } from '../machine/run-input.ts'; import { newRun, type RunState, type TaskFrame } from '../machine/run-state.ts'; @@ -33,7 +34,7 @@ const asking: TaskFrame = { input: ticket, variables: { attempt: approval }, timeout: `${executionId}/timers/2`, - body: { kind: 'call', key: openCall, primitive: 'inference', name: 'classify' }, + body: { kind: 'call', key: openCall, function: 'notify', arguments: ticket, label: 'notify the owner' }, }; const forking: TaskFrame = { @@ -54,7 +55,7 @@ const forking: TaskFrame = { }, }; -export const runningState: RunState = { +const running: RunState = { ...newRun, executionId, status: 'running', @@ -82,12 +83,13 @@ export const runningState: RunState = { receivedBytes: 160, overflow: null, }, - heldBytes: 16_500, + heldBytes: 0, + historyBytes: 9000, machine: { values: { - 0: { value: { seen: 1 }, bytes: 10, holders: 1 }, - [ticket]: { value: { ticket: 7 }, bytes: 12, holders: 7 }, - [approval]: { value: 1, bytes: 1, holders: 1 }, + 0: { value: { seen: 1 }, bytes: 10 }, + [ticket]: { value: { ticket: 7 }, bytes: 12 }, + [approval]: { value: 1, bytes: 1 }, }, nextValue: 3, context: 0, @@ -112,6 +114,8 @@ export const runningState: RunState = { }, }; +export const runningState: RunState = { ...running, heldBytes: heldBytesOf(running) }; + export const started: Started = { kind: 'started', executionId, diff --git a/packages/workflow-engine/src/testing/streams.ts b/packages/workflow-engine/src/testing/streams.ts new file mode 100644 index 000000000..9ee537f3f --- /dev/null +++ b/packages/workflow-engine/src/testing/streams.ts @@ -0,0 +1,61 @@ +import type { RunOutput } from '../dispatch/run-output.ts'; +import type { InputReceipt } from '../machine/input-receipt.ts'; +import { eventBytesOf, withHistoryBytes, type PositionedEvent } from '../run-log/run-event.ts'; +import { stateFormat } from '../run-log/state-format.ts'; +import type { StatePatch } from '../run-log/state-patch.ts'; +import { at, document, executionId } from './runs.ts'; + +export interface Change { + readonly receipt: InputReceipt; + readonly patch: StatePatch; + readonly outputs?: readonly RunOutput[]; +} + +const timer = `${executionId}/timers/1`; + +const timerPath = `/timers/armed/${timer.replaceAll('/', '~1')}`; + +export function streamOf(changes: readonly Change[]): readonly PositionedEvent[] { + return changes.reduce((events: readonly PositionedEvent[], { receipt, patch, outputs = [] }: Change, index) => { + const before = events.reduce((sum, { event }: PositionedEvent) => sum + eventBytesOf(event), 0); + const event = withHistoryBytes( + { type: 'input_applied', format: stateFormat, receipt, steps: [], patch, outputs }, + before, + ); + return [...events, { version: index + 1, event }]; + }, []); +} + +export const exampleStream = streamOf([ + { + receipt: { kind: 'started', key: executionId, at }, + patch: [ + { op: 'replace', path: '/executionId', value: executionId }, + { op: 'replace', path: '/status', value: 'running' }, + { op: 'replace', path: '/workflow', value: { document, input: 1 } }, + { op: 'add', path: '/machine/values/1', value: { value: { ticket: 7 }, bytes: 12 } }, + { op: 'replace', path: '/inputs', value: 1 }, + { op: 'replace', path: '/startedAt', value: at }, + { op: 'replace', path: '/lastInputAt', value: at }, + ], + }, + { + receipt: { kind: 'event_received', key: 'event-1', at: at + 10, eventType: 'com.acme.tick' }, + patch: [ + { op: 'add', path: timerPath, value: { purpose: 'wait', reference: '/do/0', dueAt: at + 1000 } }, + { op: 'replace', path: '/timers/next', value: 2 }, + { op: 'replace', path: '/inputs', value: 2 }, + { op: 'replace', path: '/lastInputAt', value: at + 10 }, + ], + outputs: [{ kind: 'arm_timer', executionId, timerId: timer, dueAt: at + 1000, purpose: 'wait' }], + }, + { + receipt: { kind: 'timer_fired', key: timer, at: at + 1000 }, + patch: [ + { op: 'remove', path: timerPath }, + { op: 'replace', path: '/machine/context', value: 1 }, + { op: 'replace', path: '/inputs', value: 3 }, + { op: 'replace', path: '/lastInputAt', value: at + 1000 }, + ], + }, +]); diff --git a/packages/workflow-engine/src/timers/timer-id.ts b/packages/workflow-engine/src/timers/timer-id.ts index 85c8d202a..633a766d9 100644 --- a/packages/workflow-engine/src/timers/timer-id.ts +++ b/packages/workflow-engine/src/timers/timer-id.ts @@ -6,6 +6,7 @@ export const TimerPurposeSchema = Schema.Literals([ 'retry_delay', 'attempt_limit', 'deadline', + 'call_deadline', 'yield', ]); From ff608b416c035d2559351def2762249cd6079e1d Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:29:42 +0100 Subject: [PATCH 16/23] docs(workflow-engine): describe the second revision of the contract The README now covers the entries for the DSL and the limits, the DSL taking its call functions from the caller, the bounded expression cache, history bytes and how they are counted, the format rule with its corpus in place of upcasters for patches, held values by reachability, the opaque call frame, the call deadline and the start receipts, where tombstones matter, due runs with the sweep's cadence and grace, troubling receipts, the API's answers for a repeated or late event, and invariants 32 to 35. Invariant 14 no longer offers one transaction, which the ledger's port does not. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/README.md | 181 +++++++++++++++++------------ 1 file changed, 107 insertions(+), 74 deletions(-) diff --git a/packages/workflow-engine/README.md b/packages/workflow-engine/README.md index 57fdad383..3d12bbdf4 100644 --- a/packages/workflow-engine/README.md +++ b/packages/workflow-engine/README.md @@ -1,34 +1,49 @@ # @beonauto/workflow-engine -The core of the workflow engine that runs on the ledger: the contract between the machine that runs a workflow and the adapters that store, time and execute for it. It knows workflows, the inputs a run takes and an executor that performs calls. It does not know brains, prompts, models or specs: whatever an adapter needs to know about a run, such as who started it, it passes as opaque `attributes` and gets back with every output. +The contract of the workflow machine that runs on the ledger: the DSL it runs, the state it keeps, the inputs it takes and the ports to the adapters that store, time and execute for it. The machine's state is shaped by the workflow DSL, so this package is the workflow machine's contract, not a generic runtime. It knows workflows, their DSL and an executor that performs calls. It does not know brains, prompts, models, specs or primitives: the functions a workflow may call come from the caller, and whatever an adapter needs to know about a run, such as who started it, it passes as opaque `attributes` and gets back with every output. -The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. This package holds the contract: the types, the ports, the idempotency keys, the dispatch watermark, the fold of a run's log and the invariants below. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, unchanged. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. +The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, with the DSL it imports from here. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. + +## Entries + +- `@beonauto/workflow-engine`: the contract, from `src/index.ts`. +- `@beonauto/workflow-engine/dsl/`: one module of the DSL, such as `dsl/json` or `dsl/policy`. The orchestration primitive imports the DSL this way, so its bundled Temporal workflow code takes neither Effect nor the ledger from the main entry. +- `@beonauto/workflow-engine/limits`: the limits, plain numbers, for the same reason. ## How a run moves -1. An adapter submits an input for an execution, with `at`, the time on its own clock. The machine never reads a clock, and it takes `max(at, lastInputAt)` as the time of the input, so time in a run never goes back however the adapters' clocks drift. -2. Holding the run's serialisation, the engine loads the run: `RunStore.load` gives the latest snapshot and the events after it, and `loadedRunOf` folds them. -3. `staleReasonOf(state, input)` says whether the input can still change the run. An input to a run whose `started` has not arrived is `not_started`: the engine answers `{ outcome: 'not_started' }`, which an adapter answers as `not_found` so the caller tries again, as the API does today. Any other reason is `stale`. Both append nothing. -4. Otherwise the machine decides, and the decision is one event, appended with the version the engine read as the expected version: the ledger's load-decide-append loop, which loads and decides again after a version conflict, up to three more times, and then fails with `Conflict`. The engine answers `{ outcome: 'applied' }`. +1. An adapter submits an input for an execution, with `at`, the time on its own clock. The machine never reads a clock. It takes `max(at, lastInputAt)` as the time of the input, and for a fired timer at least the time it was due, so time in a run never goes back however the adapters' clocks drift (`inputTimeOf`). +2. Holding the run's serialisation, the engine runs the ledger's own load-decide-append loop, `decisionLoop` from `@beonauto/ledger`, with a load of its own: `RunStore.load` gives the latest snapshot and the events after it, and `loadedRunOf` folds them (`runLoopOf`). +3. `staleReasonOf(state, input)` says whether the input can still change the run. An input to a run whose `started` has not arrived is `not_started`: the engine answers `{ outcome: 'not_started' }`, which an adapter answers as `not_found` so the caller tries again, as the API does today. Any other reason is `stale`. Neither appends anything (`submissionOf`). +4. Otherwise the machine decides, and the decision is one event, appended with the version the loop read as the expected version. After a version conflict the loop loads and decides again, up to three more times, and then fails with the ledger's `Conflict`. The engine answers `{ outcome: 'applied' }`. 5. When a snapshot is due, the engine saves one. -6. It dispatches the outputs of every event above the run's dispatch watermark, in the order of the stream, and stops at the first output that fails. The watermark moves to the last event whose outputs were all dispatched. -7. `wake(executionId)` does step 6 again. `sweep()` walks the live runs, wakes each, and has `Timers.sweep` arm again any of the run's armed timers the timer store lost. +6. It dispatches the outputs of every event above the run's dispatch watermark, in the order of the stream, and stops at the first output that fails. The watermark moves to the last event whose outputs were all dispatched. An event that arms or cancels a timer also notes the run's next due time in the record (`runDueOf`). +7. `wake(executionId)` does step 6 again. `sweep(before)` wakes the runs that are overdue (below). ## Layers and ports -| Layer | Folder | What it holds | Port | -| ------------- | ------------------- | -------------------------------------------------------------- | --------------------------------------- | -| machine | `src/machine` | inputs, state, admission, the clock clamp, limits, the decider | none: pure | -| run log | `src/run-log` | events, state patches, state formats, the fold, snapshots | `RunStore` (one Emmett stream per run) | -| timers | `src/timers` | timer ids and what each timer is for | `Timers` | -| inbox | `src/inbox` | the external events a run receives, and their limits | none: events arrive as `event_received` | -| executor | `src/executor` | call keys and call results | `Executor` | -| dispatch | `src/dispatch` | outputs, the watermark, the order of a dispatch | `DispatchWatermark` | -| serialisation | `src/serialisation` | one input at a time for each run | `RunSerialiser` | -| settlement | `src/settlement` | the record store's settlements and its live runs | `RecordStore` (the brain's ledger) | -| engine | `src/engine` | the ports together and the engine's own interface | `WorkflowEngine` | +| Layer | Folder | What it holds | Port | +| ------------- | ------------------- | ---------------------------------------------------------------------- | --------------------------------------- | +| DSL | `src/dsl` | JSON, durations, jq expressions with their work budget, tasks, policy | none: pure | +| machine | `src/machine` | inputs, state, held values, admission, the clock, limits, the decider | none: pure | +| run log | `src/run-log` | events, state patches, state formats, the fold, snapshots | `RunStore` (one Emmett stream per run) | +| timers | `src/timers` | timer ids and what each timer is for | `Timers` | +| inbox | `src/inbox` | the external events a run receives | none: events arrive as `event_received` | +| executor | `src/executor` | call keys | `Executor` | +| dispatch | `src/dispatch` | outputs, the watermark, the order of a dispatch, a run's next due time | `DispatchWatermark` | +| serialisation | `src/serialisation` | one input at a time for each run | `RunSerialiser` | +| settlement | `src/settlement` | settle receipts, due times, troubling receipts | `RecordStore`, `RunReporter` | +| engine | `src/engine` | the loop on the ledger, the ports together, the engine's interface | `WorkflowEngine` | + +Every port answers with an Effect. None of them is a clock: time comes in with the inputs, and the sweep is given the time before which a run is overdue. + +The settlement and call-result vocabulary lives once, in `@beonauto/operations` (`SettlementSchema`, `CallResultSchema`, `invalidArguments`), beside its `Outcome`. The limits and `DslError` live once, here; the interpreter imports them. + +## The DSL -Every port answers with an Effect. None of them is a clock: time comes in with the inputs. +`policyOf(functions)` gives the policy a document is checked against. The caller gives the functions a workflow may call, each with the checks of its arguments, and the words that explain where a workflow reaches the world and how it starts; the orchestration primitive gives `execute_spec` (`primitives/orchestration/src/document/workflow-functions.ts`). The DSL itself names no function. + +Compiled expressions are kept in one cache for the process, of at most 262,144 characters of expression source, letting go of the expression used longest ago: a compiled expression measured 22 to 34 bytes of heap per character of source, so the cache holds at most about 9 MiB, and compiling one again took 10 to 150 µs. ## Inputs @@ -40,11 +55,13 @@ Every port answers with an Effect. None of them is a clock: time comes in with t | `event_received` | the event, with an `id` of 1 to 256 characters and a `type` | the run has received an event with that id | | `cancel_requested` | nothing more | a cancel was requested before | -Every input carries the execution id and `at`. Any input but `started` is `not_started` for a run that has not started and `run_ended` for one that has ended. Two inputs are never taken as stale, because they mean an adapter routed wrongly: an input for another execution, and a second `started` with another document. `staleReasonOf` dies on them with `RunMismatch`. +Every input carries the execution id and `at`. Any input but `started` is `not_started` for a run that has not started and `run_ended` for one that has ended. Three inputs are never taken as stale, because they mean an adapter routed wrongly: an input for another execution, and a second `started` with another document or another input. `staleReasonOf` dies on them with `RunMismatch`. + +For `send_execution_event`, the API answers an event the run received before (`event_received_before`) with success, since the event is delivered, and an event for a run that has ended (`run_ended`) with `not_found`, as today. A repeated event appends nothing, so it no longer counts toward the 1,024 events and 4 MiB a run takes over its life; under Temporal it did. ## The run's log -A run's stream is a state-transition log, not classic event sourcing. Each event records the change an input made to the run's state, as a patch, rather than a domain fact for the fold to interpret. Replaying a run therefore applies patches and evaluates nothing: no expression, no retry arithmetic, no version of the machine's code. A run started under one version of the machine loads under the next. What the events give up in meaning they get back in two fields written for people reading the log. +A run's stream is a state-transition log, not classic event sourcing. Each event records the change an input made to the run's state, as a patch, rather than a domain fact for the fold to interpret. Replaying a run applies patches and evaluates nothing: no expression, no retry arithmetic, no version of the machine's code. What the events give up in meaning they get back in two fields written for people reading the log. Each event, `input_applied`, holds: @@ -56,33 +73,45 @@ Each event, `input_applied`, holds: An input that is stale or not started appends nothing; logging it is the adapter's job. -## Held data and the size of a patch +### History bytes -Values a run holds live once, in the state's value table, `machine.values`: the run's input, the result of a call, the data of an event, the output of an expression. Each value has an id that counts up within the run and is never reused, its size in bytes of UTF-8 JSON, and the number of holders. Frames, list cursors, variables, branches and the context refer to values by id. Passing a value on, as a call's output becomes the cursor's data, the next task's input and the context, adds a holder, not a copy; a value leaves the table when its last holder lets go. +The state counts the bytes of its own history, `historyBytes`: the UTF-8 bytes of every event in the stream as JSON (`eventBytesOf`), the event that sets it included. The machine builds an event, measures it, and adds `replace /historyBytes` with the count before it plus the bytes of the event that carries that operation (`withHistoryBytes`, which solves for the few digits the number adds). So `decide` knows the history before it appends, and bounds it as it bounds inputs. A load checks the count: a state whose `historyBytes` is not the snapshot's count plus the bytes of the events after it is `UnreadableRun`. -The data a run holds is the sizes of the values in the table, plus 4 KiB for each open frame, plus the document. That bounds the state, and so every snapshot, at 4 MiB of held data. It also bounds a patch: an input adds each value it made once, as `add /machine/values/`, and otherwise moves ids, so a call that answers with 1 MiB gives a patch of about 1 MiB however many places the answer lands. +## State formats -A transformation can still make more data than it was given. The machine measures each event before the append; one over 1.5 MiB ends the run with a `raised` runtime error, status 500, the ending today's limits give, recorded as a small final event: the outcome, the frames gone, timers and calls cancelled and the execution settled. +Every event and every snapshot names its state format; `stateFormat` is 1. A change to the state's schema is a new format, and: -Frames never store the document: they name tasks by reference, a JSON Pointer into `workflow.document`. +- each event is folded under its own format, and a state that crosses to a newer format is read strictly under the old one and upcast by that format's upcaster (`OlderFormat.read`, `OlderFormat.upcast`) before the next event applies; +- formats never go back within a stream, and a format newer than the code is refused, both when the run loads (`UnreadableRun`); +- `packages/workflow-engine/corpus/format-.json` holds a committed stream and snapshot of every format, which must load to the state it recorded (`src/run-log/corpus.test.ts`). A new format adds its corpus and keeps every older one loading. -## State formats +Patches are never rewritten: a patch applies only to the format it was written for. `evolve` applies a patch strictly, `add` to a member that exists or `replace` and `remove` of one that does not die with `PatchFailed`, and the result must decode as the state with no member the format does not describe (`onExcessProperty: 'error'`), so a skew between a log and the code that reads it is caught when the run loads, never folded into a wrong state. -Every event and every snapshot names its state format, `stateFormat`, today 1. A change to the state's schema is a new format, and a new format ships only with an upcaster for snapshots and one for patches. `evolve` applies a patch strictly: an `add` to a member that exists, or a `replace` or `remove` of one that does not, dies with `PatchFailed`, and the result must decode as the state, so a skew between a log and the code that reads it is caught when the run loads, never folded into a wrong state. +## Held data and the size of a patch + +Values a run holds live once, in the state's value table, `machine.values`: the run's input, the result of a call, the arguments of a call, the data of an event, the output of an expression. Each value has an id that counts up within the run and is never reused, and its size in bytes of UTF-8 JSON. Frames, list cursors, variables, branches, the context and the workflow's input refer to values by id, so passing a value on moves an id, not a copy. + +After each decision the machine sweeps the table: a value stays while the frames, the context or the workflow's input reach it, and the others leave with `remove /machine/values/` (`withReachableValuesOnly`). The data a run holds, `heldBytes`, is the bytes of the values reached, plus 4 KiB for each frame, plus the document (`heldBytesOf`); a property test over 400 generated states checks it against the values an independent walk finds. That bounds the state, and every snapshot, at 4 MiB of held data. It also bounds a patch: an input adds each value it made once and otherwise moves ids, so a call that answers with 1 MiB gives a patch of about 1 MiB however many places the answer lands. + +A transformation can still make more data than it was given. The machine measures each event before the append; one over 1.5 MiB ends the run with a `raised` runtime error, status 500, the ending today's limits give, recorded as a small final event. + +Frames never store the document: they name tasks by reference, a JSON Pointer into `workflow.document`. A call frame holds the opaque `function` the call names, its `arguments` as a value id, and a `label` for people reading the log; the machine knows nothing of what the arguments mean. ## Outputs and receipts -| Output | Carries | Idempotent by | Port | Receipts | -| -------------- | ------------------------------------------------------------- | ------------- | -------------------- | ------------------------------------------------------------------------ | -| `arm_timer` | timer id, due time, purpose | timer id | `Timers.arm` | `armed`, `already_armed`, `refused_after_cancel` | -| `cancel_timer` | timer id | timer id | `Timers.cancel` | `cancelled`, `already_fired`, `tombstoned` | -| `start_call` | call key, the function, its arguments, the longest it may run | call key | `Executor.start` | `started`, `already_started`, `refused_after_cancel` | -| `cancel_call` | call key | call key | `Executor.cancel` | `cancelled`, `already_answered`, `tombstoned` | -| `settle` | execution id and settlement | execution id | `RecordStore.settle` | `recorded`, `already_recorded`, `settled_otherwise`, `unknown_execution` | +| Output | Carries | Idempotent by | Port | Receipts | +| -------------- | ------------------------------------------------------------- | ------------- | -------------------- | ------------------------------------------------------------------------------- | +| `arm_timer` | timer id, due time, purpose | timer id | `Timers.arm` | `armed`, `already_armed`, `refused_after_cancel` | +| `cancel_timer` | timer id | timer id | `Timers.cancel` | `cancelled`, `already_fired`, `tombstoned` | +| `start_call` | call key, the function, its arguments, the longest it may run | call key | `Executor.start` | `started`, `started_again`, `running`, `answered_again`, `refused_after_cancel` | +| `cancel_call` | call key | call key | `Executor.cancel` | `cancelled`, `already_answered`, `tombstoned` | +| `settle` | execution id and settlement | execution id | `RecordStore.settle` | `recorded`, `already_recorded`, `settled_otherwise`, `unknown_execution` | -A cancel for a key the timer store or the executor has never seen is recorded as a tombstone, so a start of that key that arrives later is refused. Answers come back as inputs: a fired timer as `timer_fired`, a finished call as `call_answered`. Every receipt is final; only a failure to answer, `DispatchFailed`, leaves an output to be dispatched again. +Answers come back as inputs: a fired timer as `timer_fired`, a finished call as `call_answered`. Every receipt is final; only a failure to answer, `DispatchFailed`, leaves an output to be dispatched again. A settle receipt of `settled_otherwise` or `unknown_execution` is troubling (`isTroubling`): the engine hands it to `RunReporter.unsettled`, which the adapter logs, as `reportUnsettled` does today. -The settlement vocabulary is the record store's: `succeeded` with an output, `rejected` with `invalid_input` or `unavailable` and a detail, or `failed`. +A call is idempotent by its key in this sense: it is answered at most once. A start for a key the executor has answered delivers the answer again (`answered_again`); a start for a key it is running leaves it running (`running`); a start for a key that is neither, because the host that ran it died, starts it again (`started_again`), as the Node spike's `ensureJob` restarted a call. And every open call is answered eventually: each `start_call` comes with an `arm_timer` of purpose `call_deadline`, due at the input's time plus `longestCallMs`, and when it fires with the call still open the task fails with a `communication` error, status 503, as a call that cannot be reached does today. A dead executor host therefore holds a run for at most `longestCallMs`, not until its 30-day deadline. + +A cancel for a key the timer store or the executor has never seen is recorded as a tombstone, so a start of that key that arrives later is refused. Dispatch takes outputs in the order of the stream, so a start always reaches the port before its cancel; tombstones only matter to an executor that takes work asynchronously, such as from a queue, where a cancel can overtake its start. The machine checks only the size of a call's arguments, at most 264 KiB as JSON. The executor checks what they mean, such as a primitive, a name and an input, and that a workflow does not call another workflow, and answers `rejected` with reason `invalid_arguments` and a detail when they are wrong. The machine maps a rejection's reason to the task's error through the table the interpreter uses today (`primitives/orchestration/src/interpreter/call-task.ts`), with `invalid_arguments` as a `validation` error, status 400, carrying the executor's detail. @@ -110,77 +139,81 @@ One input at a time for each run is the adapter's job. On Cloudflare it is the r On Node the timers table goes through the ledger's own SQLite driver or lives in a separate file: written through a second SQLite library to the ledger's file, committed cancels were lost (`spikes/node/results/lost-write-repeat.json` on branch `spike/engine-node`). -## Live runs and the sweep +## Due runs and the sweep + +The record keeps, for each live run, its next due time, the earliest of its armed timers, and whether its dispatch fell behind (`RunDue`). The engine writes it while dispatching an event that arms or cancels a timer, idempotent by the event's version, and when a dispatch stops at a failure. A running run always has a timer armed, its deadline, so it is always due at some time. -The live runs are the executions the record store has recorded `started` with `finishesLater` and not yet settled, what `awaitsSettlement` in `@beonauto/specs` answers: `RecordStore.liveRuns()`. On Node that is one query over the one file. `sweep()` walks them; for each it calls `wake`, and `Timers.sweep` with the run's armed timers, which arms again any the timer store lost, such as an alarm of an evicted Durable Object that gave up after its retries. +`sweep(before)` asks `RecordStore.dueRuns(before)` for the runs due before that time or behind, and only those: for each it calls `wake`, and `Timers.sweep` with the run's armed timers, which arms again any the timer store lost. It folds no other run. The adapter sweeps every minute, the most often a Cloudflare cron trigger runs, and passes a time one minute ago, so a run is swept once its timer is a minute late: alarms fired 5 ms late at p99, and the one alarm due while `wrangler dev` was stopped fired 15.6 s late when it was restarted (`spikes/cloudflare/results/timers.json`); Node's timers fired 3.7 ms late at p99 (`spikes/node/results/timers-precision.json`). ## State and snapshots A run's state is plain JSON: no `Map`, `Set`, `Date`, `undefined`, class or function, so `JSON.parse(JSON.stringify(state))` is the state. -A snapshot is `{ format, executionId, version, historyBytes, state }`, the state folded from events 1 to `version`, written only once event `version` is durable; the run store keeps only the latest. A snapshot is due after 1,000 inputs, or once the events since the last snapshot take as many bytes as that snapshot did, and at least 1 MiB, so writing snapshots never costs more bytes than the history it covers. +A snapshot is `{ format, executionId, version, historyBytes, state }`, the state folded from events 1 to `version`, written only once event `version` is durable; the run store keeps only the latest. A snapshot is due once the events since the last one take as many bytes as that snapshot did, and at least 1 MiB, so writing snapshots never costs more bytes than the history they cover. -A snapshot holds at most about 5.3 MiB: the held data (4 MiB: the values, the document and the frames), the events waiting in the inbox (1 MiB), the ids of the events received (1,024 of at most 256 characters), and the timers, calls and run counters, a few dozen bytes for each frame. It is stored in chunks of at most 1 MiB of UTF-8, cut between characters, never inside one. D1 and Durable Object SQLite both take rows of at most 2 MB; a 1 MiB chunk leaves room for the row's other columns, and an event, at most 1.5 MiB, fits in a row too. +A snapshot holds at most about 5.3 MiB: the held data (4 MiB: the values, the document and the frames), the events waiting in the inbox (1 MiB), the ids of the events received (1,024 of at most 256 characters), and the timers, calls and run counters, a few dozen bytes for each frame. It is stored in chunks of at most 1 MiB of UTF-8, cut by `TextEncoder.encodeInto` at a code point, never inside one; a 5.2 MB snapshot took 4 ms to encode and chunk. D1 and Durable Object SQLite both take rows of at most 2 MB; a 1 MiB chunk leaves room for the row's other columns, and an event, at most 1.5 MiB, fits in a row too. ## Limits -| Limit | Value | When it is reached | -| --------------------------- | ------------ | ------------------------------------------------------------------- | -| data a run holds | 4 MiB | the run ends, raised, as today | -| one event | 1.5 MiB | the run ends, raised, in a small final event | -| arguments of a call | 264 KiB | the task raises a validation error | -| tasks in one input | 100 | the machine arms a timer due at once and goes on when it fires | -| expression work | 8,000,000 | for one expression, as today | -| work in one input | 16,000,000 | as above: a timer due at once, then the rest | -| tasks without waiting | 10,000 | the run ends, raised, as today | -| events waiting in the inbox | 64, 1 MiB | the run ends, raised, as today | -| events a run receives | 1,024, 4 MiB | the run ends, raised, as today; it also bounds the ids kept | -| inputs a run takes | 100,000 | checked in `decide` from `state.inputs`: the run ends, raised | -| history | 512 MiB | checked by the engine from the bytes appended: the run ends, raised | - -## On the ledger's loop - -The engine decides and appends the way the ledger does, with the ledger's own pieces from `@beonauto/ledger`. The run store's append fails with the ledger's `VersionConflict`, and the engine retries it with `retriedOnVersionConflict`, which fails with `Conflict` after three more attempts. The SQLite run store, in the adapters' step, reads a run's tail with `EventStore.read(stream, after)` and appends with `eventAppenderOf`. +| Limit | Value | When it is reached | +| --------------------------- | --------------- | -------------------------------------------------------------------------- | +| data a run holds | 4 MiB | the run ends, raised, as today | +| one event | 1.5 MiB | the run ends, raised, in a small final event | +| arguments of a call | 264 KiB | the task raises a validation error | +| tasks in one input | 100 | the machine arms a timer due at once and goes on when it fires | +| expression work | 8,000,000 | for one expression, as today | +| work in one input | 16,000,000 | as above: a timer due at once, then the rest | +| tasks without waiting | 10,000 | the run ends, raised, as today | +| events waiting in the inbox | 64, 1 MiB | the run ends, raised, as today | +| events a run receives | 1,024, 4 MiB | the run ends, raised, as today; it also bounds the ids kept | +| inputs a run takes | 100,000 | checked in `decide` from `state.inputs`: the run ends, raised | +| history | 512 MiB | checked in `decide` from `state.historyBytes`: the run ends, raised | +| a call | `longestCallMs` | its `call_deadline` timer fires: the task fails with a communication error | + +The run that reaches the inputs or history bound ends in one more small event, so a stream holds at most 100,001 events and 512 MiB plus that event. ## Invariants Each sentence is something a reviewer can check against the code or a test. **[engine]** marks what this package and the machine guarantee; **[adapter]** marks an obligation of every adapter. 1. **[engine]** A run has exactly one stream, and the stream's version is the number of inputs the run applied. -2. **[engine]** `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source, no locale and no storage, and the same state and input give the same events (`src/engine/portability.test.ts`). -3. **[engine]** Time in a run is only ever an input's clamped `at`; random draws come from the seed in `started` and the number of draws in the state. +2. **[engine]** `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source, no locale and no storage, and the same state and input give the same events (`src/engine/portability.test.ts`, over the machine, the run log, the DSL and jq). +3. **[engine]** Time in a run is only ever an input's time from `inputTimeOf`; random draws come from the seed in `started` and the number of draws in the state. 4. **[engine]** `evolve(state, event)` applies the event's patch strictly and does nothing else (`src/run-log/run-fold.test.ts`, `src/run-log/state-patch.test.ts`). -5. **[engine]** An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version. -6. **[engine]** A stale input appends nothing, and neither does an input to a run that has not started, nor one that would change nothing. +5. **[engine]** An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version (`src/engine/run-loop.test.ts`). +6. **[engine]** A stale input appends nothing, and neither does an input to a run that has not started. 7. **[engine]** A late answer, a duplicate answer, a second delivery of an event and the fire of a cancelled timer are stale inputs (`src/machine/admission.test.ts`). 8. **[engine]** The deduplication state is bounded: armed timers and open calls are what is outstanding, and a run keeps at most 1,024 event ids. 9. **[engine]** Timer ids and value ids are never reused within a run, and the run counter of a task reference only counts up. 10. **[engine]** The outputs of a run are exactly the `outputs` of its events, and an output is dispatched only after the event that holds it is appended. -11. **[adapter]** Every output is idempotent by its key, so dispatching it twice has the effect of dispatching it once, and a cancel of a key never seen leaves a tombstone that refuses a later start. +11. **[adapter]** A timer is armed at most once and fires at least once until cancelled; a call is answered at most once, a start of a call neither answered nor running starts it again, and a cancel of a key never seen leaves a tombstone that refuses a later start. 12. **[engine]** The watermark never goes down, and every output of every event at or below it has been dispatched at least once (`src/dispatch/dispatch-watermark.test.ts`). 13. **[engine]** A run's outcome is in its stream before the record store is asked to record it, and the `settle` output is dispatched again until the record store answers. -14. **[adapter]** The run log and the record store may be different stores: two writes, each idempotent by execution id, the second retried. An adapter whose two stores are one database may make them one transaction. -15. **[engine]** Loading a run from its latest snapshot and the events after it gives the same state as folding its whole stream (`src/run-log/run-fold.test.ts`). +14. **[adapter]** The run log and the record store are two writes, each idempotent by execution id, the second retried; `EventStore.append` writes one stream, so neither adapter makes them one transaction. +15. **[engine]** Loading a run from its latest snapshot and the events after it gives the same state as folding its whole stream (`src/run-log/run-fold.test.ts`, `src/run-log/corpus.test.ts`). 16. **[adapter]** Only the latest snapshot of a run is kept, in chunks of at most 1 MiB of UTF-8 (`src/run-log/snapshot.test.ts` for the chunks). -17. **[engine]** The state of a run is plain JSON, and the data it holds, measured by size, stays at or under 4 MiB. -18. **[engine]** No event is larger than 1.5 MiB as JSON: the machine measures each event before the append and ends the run instead (`src/run-log/run-event.test.ts`); and no input runs more than 100 tasks. +17. **[engine]** The state of a run is plain JSON, and the data it holds stays at or under 4 MiB. +18. **[engine]** No event is larger than 1.5 MiB as JSON: the machine measures each event before the append and ends the run instead (`src/run-log/run-event.test.ts` for the measure); and no input runs more than 100 tasks. 19. **[adapter]** Inputs of one run are applied one at a time; the machine relies only on the expected version of each append. -20. **[engine]** Nothing in this package uses a Node-only API, a dynamic import or code generation, or imports Temporal (`src/engine/portability.test.ts`). -21. **[engine]** No module of the engine keeps a cache that grows with the history of a run: what an input costs in memory is bounded by the input and the state (`src/engine/portability.test.ts`). +20. **[engine]** Nothing in this package, nor jq, uses a Node-only API, a dynamic import, code generation or a host timer, or imports Temporal (`src/engine/portability.test.ts`). +21. **[engine]** No module of the engine keeps a cache that grows with the history of a run, and the one cache of the process, compiled expressions, is bounded (`src/engine/portability.test.ts`, `src/dsl/bounded-cache.test.ts`). 22. **[engine]** No two events of a stream have the same receipt kind and key. -23. **[engine]** `lastInputAt` never decreases (`src/machine/input-receipt.test.ts`). +23. **[engine]** `lastInputAt` never decreases, and a fired timer's input is never earlier than the time it was due (`src/machine/input-receipt.test.ts`). 24. **[engine]** A snapshot at version v is the fold of events 1 to v, and is written only after event v is durable; **[adapter]** the run store writes it only then. 25. **[engine]** A dispatch takes outputs in the order of the stream and stops at the first that fails (`src/dispatch/dispatch-watermark.test.ts`). 26. **[engine]** Every armed timer and every open call has exactly one `arm_timer` or `start_call` and at most one cancel in the stream. 27. **[engine]** An ended run has no armed timers and no open calls. 28. **[engine]** `settle` appears once in a stream, in its last event. -29. **[engine]** A run takes at most 100,000 inputs and 512 MiB of history; the input that would go past either ends the run. -30. **[engine]** Every event and every snapshot names its state format, and one of another format is refused when it is read (`src/run-log/run-event.test.ts`, `src/run-log/snapshot.test.ts`). +29. **[engine]** A run takes at most 100,000 inputs and 512 MiB of history, both checked in `decide` from the state; the input that would go past either ends the run. +30. **[engine]** Every event and every snapshot names its state format; formats never go back within a stream, a format newer than the code is refused, and a committed corpus of every format loads (`src/run-log/run-fold.test.ts`, `src/run-log/corpus.test.ts`). 31. **[engine]** Every `arm_timer` is due at or after the time of the input that armed it. +32. **[engine]** `state.historyBytes` is the bytes of the stream's events as JSON, and a load dies when it is not (`src/run-log/run-event.test.ts`, `src/run-log/run-fold.test.ts`). +33. **[engine]** Every open call has an armed `call_deadline` timer due no later than its start's time plus `longestCallMs`, so every open call is answered. +34. **[engine]** After a decision the value table holds exactly the values the frames, the context and the workflow's input reach, and `heldBytes` is their bytes, 4 KiB a frame and the document (`src/machine/held-values.test.ts`). +35. **[engine]** A troubling settle receipt is reported, never dropped. ## Open design points - `evolve` decodes the whole state after each patch, so a load costs one decode of the state whatever the tail, and an applied input one more. The machine's step measures that against the 9 ms a fold from a snapshot every 1,000 events took in workerd (`spikes/cloudflare/results/fold.json` on branch `spike/engine-cloudflare`); checking only the patched paths is the fallback. - A workflow that calls a workflow raises a `configuration` error today and a `validation` error once the executor rejects it with `invalid_arguments`; keeping `configuration` needs a reason of its own. -- The machine needs the DSL, expressions and policy that live in the orchestration primitive today; where the machine itself lives, beside them or with this package, is decided when it is written. - What replaces Temporal's limits on a run's history is decided here as 100,000 inputs and 512 MiB, both well above what a workflow could reach on Temporal; real use may move them. From 932c0987f043aadd939dfccaf03a29bb0f72ce84 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:30:40 +0100 Subject: [PATCH 17/23] docs(global): correct and extend the workflow engine decision Two figures said more than was measured: the 15.6 s late alarm was one due while wrangler dev was stopped, firing when it restarted, and local workerd enforced no CPU limit at all, not even cpu_ms = 50 (spikes/cloudflare/results/timers.json and fold.json). Settlement is two writes, never one transaction, since the ledger's port appends to one stream. The 4 MiB a run holds, and repeated events no longer counting toward its life, move under Consequences as changes tenants see. The record now states the format rule with its corpus, the call deadline, the sweep of overdue runs, the DSL living in the engine package, the test count of the rewrite (129 interpreter, 79 DSL), a plan section, and the Cloudflare executor for long calls as an open decision due before the Cloudflare adapters. Co-Authored-By: Claude Opus 5.5 --- .../0001-workflow-engine-on-the-ledger.md | 36 +++++++++++++------ 1 file changed, 25 insertions(+), 11 deletions(-) diff --git a/docs/decisions/0001-workflow-engine-on-the-ledger.md b/docs/decisions/0001-workflow-engine-on-the-ledger.md index 0a9dc9a0a..8d569b724 100644 --- a/docs/decisions/0001-workflow-engine-on-the-ledger.md +++ b/docs/decisions/0001-workflow-engine-on-the-ledger.md @@ -16,31 +16,37 @@ We weighed three other engines. Cloudflare Workflows runs only on Cloudflare, so ## Decision -A run is a decider in Emmett's workflow shape, on the load-decide-append loop the ledger already uses, with the ledger's own conflict retry: +A run is a decider in Emmett's workflow shape, on the ledger's own load-decide-append loop and conflict retry, `decisionLoop`, with a load that folds a snapshot and its tail: - `decide(input, state)` says what an input changes and `evolve(state, event)` applies it. The inputs are a start, a timer fired, a call answered, an event received and a cancel request, each with the time it arrived, never earlier than the input before it. - One stream per run. An applied input appends one event, with the version it was decided on as the expected version. -- The stream is a state-transition log, not classic event sourcing. Each event holds the change to the state as a JSON Patch, a receipt naming the input, the steps it ran and the outputs: arm or cancel a timer, start or cancel a call, settle. Replay applies patches and evaluates nothing, so a run started under one version of the interpreter loads under the next, and the receipt and steps keep the log readable. Every event and snapshot names its state format; a new format ships with upcasters. -- Outputs are dispatched after the append, in stream order behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. A cancel that arrives before its start leaves a tombstone. +- The stream is a state-transition log, not classic event sourcing. Each event holds the change to the state as a JSON Patch, a receipt naming the input, the steps it ran and the outputs: arm or cancel a timer, start or cancel a call, settle. Replay applies patches and evaluates nothing, so a run started under one version of the interpreter loads under the next, and the receipt and steps keep the log readable. +- Every event and snapshot names its state format. Each event folds under its own format, the state is upcast where the format changes, formats never go back within a stream, and a committed corpus of every format must load as a test. +- Outputs are dispatched after the append, in stream order behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. Every started call has a deadline timer, so every call is answered. - Deduplication lives in the run's state. Emmett is the store, never the engine: we use neither its workflow handler, which folds the whole stream for every input, nor its processors. -- A snapshot follows 1,000 inputs, or as many bytes of events as the last snapshot took and at least 1 MiB; only the latest is kept, in chunks of at most 1 MiB under the 2 MB row limit of D1 and Durable Object SQLite. A run may hold 4 MiB instead of 16, take 100,000 inputs and write 512 MiB of history. +- A snapshot follows once the events since the last one take as many bytes as it did, and at least 1 MiB; only the latest is kept, in chunks of at most 1 MiB under the 2 MB row limit of D1 and Durable Object SQLite. A run may take 100,000 inputs and write 512 MiB of history, both checked in `decide`. - Four adapters sit behind small ports: run store, timers, executor and record store, with the watermark and per-run serialisation beside them. -- On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes runs that fell behind. +- On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes the runs the record says are overdue. - Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process; a second process on that file is unsupported. Timers go through the ledger's own SQLite driver or a separate file. - A PostgreSQL adapter, later, assumes no 2 MB row limit and takes a lease per run for serialisation. +- The package `@beonauto/workflow-engine` is the workflow machine's contract, since the state is shaped by the workflow DSL, and the DSL lives in it. The orchestration primitive imports the DSL from it and keeps parsing, the primitive, the event and cancel operations and, for now, the Temporal runtime. The engine knows no brains, specs or primitives: the caller names the functions a workflow may call. + +Still open, due before the Cloudflare adapters: how Cloudflare executes a call that runs longer than an alarm handler, a step of a Cloudflare Workflow or a queue to a container. ## Consequences -The main cost is rewriting the interpreter as a machine that steps from state to state instead of an async function Temporal replays. Its DSL, expressions and policy stay. +The main cost is rewriting the interpreter, 129 tests of it beside the DSL's 79, as a machine that steps from state to state instead of an async function Temporal replays. Its DSL, expressions and policy stay. + +Tenants see two changes: a run may hold 4 MiB of data instead of 16, and a repeated event no longer counts toward the events a run takes over its life. We give up Temporal's durable timers, deduplicated delivery, replay, web UI and operator tools. We must build and keep correct: - Timers that fire at least once; a fire of a timer no longer armed changes nothing. - Deduplication in state, every key bounded, or snapshots grow with the run. - The watermark: a crash between append and dispatch loses nothing. -- The sweep, for alarms that fire late after eviction or give up after their retries. +- The sweep, for alarms that fire late or give up after their retries. - Serialisation per run: an in-process lock in Node, the object's thread on Cloudflare. -- Settlement in two stores, the run's stream and the brain's record, each idempotent by execution id, the second retried. +- Settlement in two stores, the run's stream and the brain's record, each idempotent by execution id, the second retried; the ledger's port appends to one stream, so they are never one transaction. Reads over the ledger and Studio replace Temporal's UI. @@ -50,11 +56,19 @@ Ended streams are kept; a deletion policy is a later decision. Tenant data is stored once, and a workflow needs no service beyond the server. +## Plan + +- The machine, on the contract and the DSL in `@beonauto/workflow-engine`. +- A conformance suite: the same workflow probes through a fake driver, the Node adapter and the workerd adapter, seeded from the server's `src/workflow-executions` tests. +- The adapters, a `cancel_execution` operation, and the workflow SDK's validators precompiled for Cloudflare. +- The cutover: `pnpm dev` without Temporal's dev server, and the README. +- Measurements before and after: the 15 recorded histories through the driver; inputs per second, and bytes per input against Temporal's 9.26 MiB for 40,000 inputs; timer lateness at p99; heap per live run; snapshot bytes per run. + ## Evidence Branch `spike/engine-node`: -- `spikes/node/results/replay.json`: the interpreter as it runs on Temporal replays 40,000 inputs in 4.6 s and retains up to 103 MiB. +- `spikes/node/results/replay.json`: the interpreter as it runs on Temporal replays 40,000 inputs in 4.6 s and retains up to 103 MiB; its log takes 9.26 MiB. - `spikes/node/results/message-id-probe.json`: Emmett appends a message with an id it has seen as a new message. - `spikes/node/results/executor-virtual.json`, `executor-real.json`: a result delivered four times settles once, a result after its timeout is ignored, a crash after the append is recovered on wake. - `spikes/node/results/timers-precision.json`, `timers-recovery.json`: timers fire 3.7 ms late at p99 when idle; after a killed scheduler all 200 fire, none twice. @@ -63,7 +77,7 @@ Branch `spike/engine-node`: Branch `spike/engine-cloudflare`: -- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. Cloudflare documents 30 s of CPU per request by default, which local workerd does not enforce: 35 s of CPU finished under the default and 2 s under `cpu_ms = 50`. -- `spikes/cloudflare/results/timers.json`: alarms fire 5 ms late at p99; an evicted object's alarm fired 15.6 s late; a sweep re-armed one that had given up. +- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. Cloudflare documents 30 s of CPU per request by default; local workerd enforced no CPU limit at all: 35 s of CPU finished with no limit set, and 2 s with `cpu_ms = 50`. +- `spikes/cloudflare/results/timers.json`: alarms fire 5 ms late at p99; an alarm due while `wrangler dev` was stopped fired 15.6 s late, when it was restarted; a sweep re-armed one that had given up. - `spikes/cloudflare/results/settlement.json`: the record was written exactly once, or given up as intended, under every injected D1 fault and crash; D1 refuses eleven events in one append. - `spikes/cloudflare/results/portability.json`: the interpreter, the DSL policy, jq and both Cloudflare ledger drivers run in workerd; the workflow SDK's validators run once precompiled. From 261a0aabe0d289d95e5faa6b815b236160e779de Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:41:37 +0100 Subject: [PATCH 18/23] fix(workflow-engine): refuse a decision of more than one event runLoopOf appended each decided event as a separate write with expectedVersion + index, so a decider that returned two events would have lost the atomicity of an input without a sound. A decision is one event (invariant 1); the loop now dies with SplitDecision on more, and appends nothing. Co-Authored-By: Claude Opus 5.5 --- .../src/engine/run-loop.test.ts | 27 +++++++++++++++++-- .../workflow-engine/src/engine/run-loop.ts | 13 ++++++--- packages/workflow-engine/src/index.ts | 2 +- 3 files changed, 35 insertions(+), 7 deletions(-) diff --git a/packages/workflow-engine/src/engine/run-loop.test.ts b/packages/workflow-engine/src/engine/run-loop.test.ts index 81f3b265b..0157ee4cd 100644 --- a/packages/workflow-engine/src/engine/run-loop.test.ts +++ b/packages/workflow-engine/src/engine/run-loop.test.ts @@ -1,8 +1,16 @@ import { Conflict } from '@beonauto/operations'; -import { Effect, Result } from 'effect'; +import { Effect, Exit, Result } from 'effect'; import { describe, expect, it } from 'vitest'; -import { eventBytesOf, runLoopOf, snapshotOf, submissionOf, type RunInput } from '../index.ts'; +import { + eventBytesOf, + runLoopOf, + snapshotOf, + SplitDecision, + submissionOf, + type RunDecider, + type RunInput, +} from '../index.ts'; import { countingDecider, memoryRunStore } from '../testing/run-store.ts'; import { at, executionId, runningState, started } from '../testing/runs.ts'; @@ -58,6 +66,21 @@ describe('the engine on the ledger loop', () => { }); }); +describe('a decision', () => { + it('dies on a decision of more than one event and appends nothing, so an input is always one atomic append', async () => { + const store = memoryRunStore(); + const twice: RunDecider = { + ...countingDecider, + decide: (input, state) => Result.map(countingDecider.decide(input, state), (events) => [...events, ...events]), + }; + + const exit = await Effect.runPromise(Effect.exit(runLoopOf(store, twice)(executionId, started))); + + expect(exit).toEqual(Exit.die(new SplitDecision({ executionId, events: 2 }))); + expect(store.events(executionId)).toEqual([]); + }); +}); + describe('the run store the tests use', () => { it('reads back the events after a version for dispatch, and takes a snapshot', async () => { const store = memoryRunStore(); diff --git a/packages/workflow-engine/src/engine/run-loop.ts b/packages/workflow-engine/src/engine/run-loop.ts index 400324d0e..4c495c9c7 100644 --- a/packages/workflow-engine/src/engine/run-loop.ts +++ b/packages/workflow-engine/src/engine/run-loop.ts @@ -1,5 +1,5 @@ import { decisionLoop, type Decided, type DecisionLoop } from '@beonauto/ledger'; -import { Effect } from 'effect'; +import { Data, Effect } from 'effect'; import { outcomeOf, staleReasonOf } from '../machine/admission.ts'; import type { RunDecider } from '../machine/run-decider.ts'; @@ -12,6 +12,11 @@ import type { Submission } from './workflow-engine.ts'; export type RunDecision = Decided; +export class SplitDecision extends Data.TaggedError('split_decision')<{ + readonly executionId: string; + readonly events: number; +}> {} + export function runLoopOf( runStore: RunStore, decider: RunDecider, @@ -19,9 +24,9 @@ export function runLoopOf( return decisionLoop( (executionId: string) => Effect.map(runStore.load(executionId), (stored) => loadedRunOf(stored)), (executionId, events, expectedVersion) => - Effect.forEach(events, (event, index) => runStore.append(executionId, event, expectedVersion + index), { - discard: true, - }), + events.length > 1 + ? Effect.die(new SplitDecision({ executionId, events: events.length })) + : Effect.forEach(events, (event) => runStore.append(executionId, event, expectedVersion), { discard: true }), decider, ); } diff --git a/packages/workflow-engine/src/index.ts b/packages/workflow-engine/src/index.ts index 7b556c55b..451783284 100644 --- a/packages/workflow-engine/src/index.ts +++ b/packages/workflow-engine/src/index.ts @@ -16,7 +16,7 @@ export { type StartCall, } from './dispatch/run-output.ts'; export { changesTimers, nextDueAtOf, runDueOf } from './dispatch/run-due.ts'; -export { runLoopOf, submissionOf, type RunDecision } from './engine/run-loop.ts'; +export { SplitDecision, runLoopOf, submissionOf, type RunDecision } from './engine/run-loop.ts'; export type { EnginePorts, Submission, SweepReport, Wake, WorkflowEngine } from './engine/workflow-engine.ts'; export { CallKeySchema, callKeyText, type CallKey } from './executor/call-key.ts'; export type { CallCancelReceipt, Executor, StartReceipt } from './executor/executor.ts'; From 6d5346277ec4a73640dfdd768cd2e06646ee67b3 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:43:06 +0100 Subject: [PATCH 19/23] docs(workflow-engine): mark the bundle entries transitional and name the first test The dsl/* and limits entries are there only to keep Effect and the ledger out of the Temporal workflow bundle, and go when Temporal goes. Invariant 5 names SplitDecision. The nested-workflow error kind leaves the open points for the decision record, and a section on the machine names invariant 33, every open call answered, as its first test. Co-Authored-By: Claude Opus 5.5 --- packages/workflow-engine/README.md | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/packages/workflow-engine/README.md b/packages/workflow-engine/README.md index 3d12bbdf4..afbc93b80 100644 --- a/packages/workflow-engine/README.md +++ b/packages/workflow-engine/README.md @@ -7,8 +7,10 @@ The same code runs in Node, where one server keeps every run in one SQLite file, ## Entries - `@beonauto/workflow-engine`: the contract, from `src/index.ts`. -- `@beonauto/workflow-engine/dsl/`: one module of the DSL, such as `dsl/json` or `dsl/policy`. The orchestration primitive imports the DSL this way, so its bundled Temporal workflow code takes neither Effect nor the ledger from the main entry. -- `@beonauto/workflow-engine/limits`: the limits, plain numbers, for the same reason. +- `@beonauto/workflow-engine/dsl/`: one module of the DSL, such as `dsl/json` or `dsl/policy`. +- `@beonauto/workflow-engine/limits`: the limits, plain numbers. + +The two subpaths are transitional. They are there only so that the orchestration primitive's bundled Temporal workflow code can import the DSL and the limits without taking Effect and the ledger from the main entry, and they go when Temporal goes. ## How a run moves @@ -180,7 +182,7 @@ Each sentence is something a reviewer can check against the code or a test. **[e 2. **[engine]** `decide(input, state)` is a pure function of its arguments: it reads no clock, no random source, no locale and no storage, and the same state and input give the same events (`src/engine/portability.test.ts`, over the machine, the run log, the DSL and jq). 3. **[engine]** Time in a run is only ever an input's time from `inputTimeOf`; random draws come from the seed in `started` and the number of draws in the state. 4. **[engine]** `evolve(state, event)` applies the event's patch strictly and does nothing else (`src/run-log/run-fold.test.ts`, `src/run-log/state-patch.test.ts`). -5. **[engine]** An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version (`src/engine/run-loop.test.ts`). +5. **[engine]** An applied input appends exactly one event, in one append, with the version the decision was made on as the expected version; the loop dies with `SplitDecision` on a decision of more than one event and appends nothing (`src/engine/run-loop.test.ts`). 6. **[engine]** A stale input appends nothing, and neither does an input to a run that has not started. 7. **[engine]** A late answer, a duplicate answer, a second delivery of an event and the fire of a cancelled timer are stale inputs (`src/machine/admission.test.ts`). 8. **[engine]** The deduplication state is bounded: armed timers and open calls are what is outstanding, and a run keeps at most 1,024 event ids. @@ -212,8 +214,11 @@ Each sentence is something a reviewer can check against the code or a test. **[e 34. **[engine]** After a decision the value table holds exactly the values the frames, the context and the workflow's input reach, and `heldBytes` is their bytes, 4 KiB a frame and the document (`src/machine/held-values.test.ts`). 35. **[engine]** A troubling settle receipt is reported, never dropped. +## Next: the machine + +The machine's first test is invariant 33: every open call is answered, by the executor or by its `call_deadline` timer. It is the one guarantee Temporal gave that this contract only promises until the machine keeps it. + ## Open design points - `evolve` decodes the whole state after each patch, so a load costs one decode of the state whatever the tail, and an applied input one more. The machine's step measures that against the 9 ms a fold from a snapshot every 1,000 events took in workerd (`spikes/cloudflare/results/fold.json` on branch `spike/engine-cloudflare`); checking only the patched paths is the fallback. -- A workflow that calls a workflow raises a `configuration` error today and a `validation` error once the executor rejects it with `invalid_arguments`; keeping `configuration` needs a reason of its own. - What replaces Temporal's limits on a run's history is decided here as 100,000 inputs and 512 MiB, both well above what a workflow could reach on Temporal; real use may move them. From 0d2cd6596694ae702f41a5cb1d028e7730e4759e Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:43:17 +0100 Subject: [PATCH 20/23] docs(global): record the nested-workflow error kind and the transitional entries A workflow that executes a workflow through a computed primitive name fails with a validation error once the executor rejects it, where the interpreter raises a configuration error today. The written case is still refused at create_spec as forbidden; both errors have status 400 and settle as invalid_input; the visible difference is the error type a catch.errors.with filter matches. The engine's dsl/* and limits entries are transitional and go when Temporal goes. Co-Authored-By: Claude Opus 5.5 --- docs/decisions/0001-workflow-engine-on-the-ledger.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/decisions/0001-workflow-engine-on-the-ledger.md b/docs/decisions/0001-workflow-engine-on-the-ledger.md index 8d569b724..c1189e714 100644 --- a/docs/decisions/0001-workflow-engine-on-the-ledger.md +++ b/docs/decisions/0001-workflow-engine-on-the-ledger.md @@ -29,7 +29,7 @@ A run is a decider in Emmett's workflow shape, on the ledger's own load-decide-a - On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes the runs the record says are overdue. - Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process; a second process on that file is unsupported. Timers go through the ledger's own SQLite driver or a separate file. - A PostgreSQL adapter, later, assumes no 2 MB row limit and takes a lease per run for serialisation. -- The package `@beonauto/workflow-engine` is the workflow machine's contract, since the state is shaped by the workflow DSL, and the DSL lives in it. The orchestration primitive imports the DSL from it and keeps parsing, the primitive, the event and cancel operations and, for now, the Temporal runtime. The engine knows no brains, specs or primitives: the caller names the functions a workflow may call. +- The package `@beonauto/workflow-engine` is the workflow machine's contract, since the state is shaped by the workflow DSL, and the DSL lives in it. The orchestration primitive imports the DSL from it and keeps parsing, the primitive, the event and cancel operations and, for now, the Temporal runtime. The engine knows no brains, specs or primitives: the caller names the functions a workflow may call. Its `./dsl/*` and `./limits` entries are transitional: they keep Effect and the ledger out of the Temporal workflow bundle and go when Temporal goes. Still open, due before the Cloudflare adapters: how Cloudflare executes a call that runs longer than an alarm handler, a step of a Cloudflare Workflow or a queue to a container. @@ -37,7 +37,7 @@ Still open, due before the Cloudflare adapters: how Cloudflare executes a call t The main cost is rewriting the interpreter, 129 tests of it beside the DSL's 79, as a machine that steps from state to state instead of an async function Temporal replays. Its DSL, expressions and policy stay. -Tenants see two changes: a run may hold 4 MiB of data instead of 16, and a repeated event no longer counts toward the events a run takes over its life. +Tenants see three changes. A run may hold 4 MiB of data instead of 16. A repeated event no longer counts toward the events a run takes over its life. And a workflow that executes a workflow through a primitive name it computes fails with a `validation` error, once the executor rejects the call as `invalid_arguments`, where the interpreter raises a `configuration` error today. The written case, `primitive: orchestration`, is still refused at `create_spec` as forbidden; only a computed name reaches the executor. Both errors have status 400 and settle the execution as `invalid_input`; the visible difference is the error `type`, which a `catch.errors.with` filter matches. We give up Temporal's durable timers, deduplicated delivery, replay, web UI and operator tools. We must build and keep correct: From da3a51b3d84bc090fa14cc4cd0f97d212a75bad5 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 20:48:03 +0100 Subject: [PATCH 21/23] docs(global): let a package's exports map name its entry points The layout rule named two files as the only entry points, while several packages already export subpaths that their READMEs explain. The rule now says what the repository does: the exports map names the few entry points, and a subpath needs a reason in the README. Co-Authored-By: Claude Fable 5.1 --- CLAUDE.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CLAUDE.md b/CLAUDE.md index 251b9101d..ddd9728c3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -43,7 +43,7 @@ Run one package's gate with `pnpm turbo run lint typecheck test --filter @beonau - 100% coverage per file is the gate. Never add coverage-ignore comments or coverage excludes; write the test. - Do not write comments. Make the code read like English through names and ordering. - Tests live next to the code as `*.test.ts`, test behaviour through the public interface, and prefer injected fakes over mocks. -- A package with more than about 12 source files groups them one level deep, in folders named after concepts that hold fewer than about 15 files each, with tests beside the code they test. Nothing new goes directly under `src`, and there are no barrel files: the entry points are `src/index.ts`, `src/testing/index.ts` and a package's commands (the server's `src/main.ts` and `dev.ts`, the identity package's `src/key-command.ts`). +- A package with more than about 12 source files groups them one level deep, in folders named after concepts that hold fewer than about 15 files each, with tests beside the code they test. Nothing new goes directly under `src`, and there are no barrel files: a package's entry points are the few paths its `exports` map names (`src/index.ts`, `src/testing/index.ts` and, where its README says why, a subpath) and its commands (the server's `src/main.ts` and `dev.ts`, the identity package's `src/key-command.ts`). - Commits are conventional with a scope named after a package or primitive folder (`feat(server): ...`, `docs(inference): ...`); `global`, `deps`, `ci` and `release` are the other scopes. - `pnpm check` must pass before you finish. Fixing a problem is a change, so rerun it. - When something fails, assume your change broke it. What is on `main` passed the same gate. From ed468d9ccc30e70154d5ca0843bfda71f3c27405 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 21:32:03 +0100 Subject: [PATCH 22/23] docs(global): describe the hosted runtime by its constraints, not its vendor Where Auto's cloud hosting runs the image is internal. The decision record now states once what that runtime allows, a single-threaded isolate per run woken by at-least-once alarms that are dropped after bounded retries, about 128 MB of memory, bounded CPU per wake-up, no long-lived process, no code generation, its own SQLite with 2 MB rows, and a shared database without interactive transactions that binds at most 100 parameters, and reuses it everywhere the vendor was named. The layout is one isolate per run, one for the brain's record and one for the org's registry. Every number stays; the hosted-runtime measurements are cited as kept in the private repository. The engine and ledger READMEs and two test names follow. Co-Authored-By: Claude Opus 5.5 --- .../0001-workflow-engine-on-the-ledger.md | 30 ++++++++++--------- packages/ledger/README.md | 4 +-- packages/ledger/src/portable-entry.test.ts | 2 +- packages/workflow-engine/README.md | 12 ++++---- .../src/engine/portability.test.ts | 2 +- 5 files changed, 26 insertions(+), 24 deletions(-) diff --git a/docs/decisions/0001-workflow-engine-on-the-ledger.md b/docs/decisions/0001-workflow-engine-on-the-ledger.md index c1189e714..685602a65 100644 --- a/docs/decisions/0001-workflow-engine-on-the-ledger.md +++ b/docs/decisions/0001-workflow-engine-on-the-ledger.md @@ -5,14 +5,16 @@ ## Context +auto-brain's image runs self-hosted, as one Node process with one SQLite file, and in Auto's cloud hosting. There the engine must run in a single-threaded isolate for each run, woken by alarms that fire at least once and are dropped after a bounded number of failed retries, with bounded memory (about 128 MB) and bounded CPU for each wake-up, no long-lived process and no code generation; each isolate keeps its state in its own SQLite, whose rows take at most 2 MB, and the database the isolates share has no interactive transactions and binds at most 100 parameters in one statement. This record calls that the hosted runtime. + Workflow specs run on Temporal today, the wrong place for them: -- Temporal cannot run on Cloudflare Workers, where Auto's cloud hosting runs auto-brain: there is no Temporal worker for workerd. +- Temporal cannot run in the hosted runtime: a Temporal worker is a long-lived process, and the hosted runtime has none. - Nobody asked for Temporal. People ask for workflows that wait, retry and survive a restart, and a self-hosted server must run a Temporal service beside it to get them. - Tenant data is stored twice: Temporal's history holds each workflow's document, input, the outputs of its calls and its events, unencrypted, beside the brain's ledger. -- We already have the store: the ledger keeps every brain's events with Emmett on SQLite, on the sqlite3, D1 and Durable Object drivers. +- We already have the store: the ledger keeps every brain's events with Emmett on SQLite, through the sqlite3 driver self-hosted and through the drivers for the isolate's own and the shared SQLite in the hosted runtime. -We weighed three other engines. Cloudflare Workflows runs only on Cloudflare, so self-hosted servers would need a second engine. Restate and Inngest are services of their own: a self-hosted server would run one beside it, as with Temporal, and each keeps step results in its own store, so tenant data would still be stored twice. +We weighed three other engines. The hosted runtime's own workflow service runs only there, so self-hosted servers would need a second engine. Restate and Inngest are services of their own: a self-hosted server would run one beside it, as with Temporal, and each keeps step results in its own store, so tenant data would still be stored twice. ## Decision @@ -24,14 +26,14 @@ A run is a decider in Emmett's workflow shape, on the ledger's own load-decide-a - Every event and snapshot names its state format. Each event folds under its own format, the state is upcast where the format changes, formats never go back within a stream, and a committed corpus of every format must load as a test. - Outputs are dispatched after the append, in stream order behind a watermark per run, and again on wake until they all succeed. Each is idempotent by its key: timer id, call key (execution, task reference, run) or execution id. Every started call has a deadline timer, so every call is answered. - Deduplication lives in the run's state. Emmett is the store, never the engine: we use neither its workflow handler, which folds the whole stream for every input, nor its processors. -- A snapshot follows once the events since the last one take as many bytes as it did, and at least 1 MiB; only the latest is kept, in chunks of at most 1 MiB under the 2 MB row limit of D1 and Durable Object SQLite. A run may take 100,000 inputs and write 512 MiB of history, both checked in `decide`. +- A snapshot follows once the events since the last one take as many bytes as it did, and at least 1 MiB; only the latest is kept, in chunks of at most 1 MiB under the 2 MB row limit of the hosted runtime's SQLite. A run may take 100,000 inputs and write 512 MiB of history, both checked in `decide`. - Four adapters sit behind small ports: run store, timers, executor and record store, with the watermark and per-run serialisation beside them. -- On Cloudflare, each run is one Durable Object, its stream in the object's SQLite and its timers on the object's alarm; a brain object keeps the record and an org object the registry; a cron sweep wakes the runs the record says are overdue. +- In the hosted runtime: one isolate per run, its stream in the isolate's SQLite and its timers on the isolate's alarm; one for the brain's record; one for the org's registry; a scheduled sweep wakes the runs the record says are overdue. - Self-hosted, one server keeps the ledger and every run in one SQLite file, in one process; a second process on that file is unsupported. Timers go through the ledger's own SQLite driver or a separate file. - A PostgreSQL adapter, later, assumes no 2 MB row limit and takes a lease per run for serialisation. - The package `@beonauto/workflow-engine` is the workflow machine's contract, since the state is shaped by the workflow DSL, and the DSL lives in it. The orchestration primitive imports the DSL from it and keeps parsing, the primitive, the event and cancel operations and, for now, the Temporal runtime. The engine knows no brains, specs or primitives: the caller names the functions a workflow may call. Its `./dsl/*` and `./limits` entries are transitional: they keep Effect and the ledger out of the Temporal workflow bundle and go when Temporal goes. -Still open, due before the Cloudflare adapters: how Cloudflare executes a call that runs longer than an alarm handler, a step of a Cloudflare Workflow or a queue to a container. +Still open, due before the hosted adapters: how the hosted runtime executes a call that runs longer than the CPU one wake-up allows, as a step of a durable workflow service outside the isolate or through a queue to a container. ## Consequences @@ -45,7 +47,7 @@ We give up Temporal's durable timers, deduplicated delivery, replay, web UI and - Deduplication in state, every key bounded, or snapshots grow with the run. - The watermark: a crash between append and dispatch loses nothing. - The sweep, for alarms that fire late or give up after their retries. -- Serialisation per run: an in-process lock in Node, the object's thread on Cloudflare. +- Serialisation per run: an in-process lock in Node, the isolate's single thread in the hosted runtime. - Settlement in two stores, the run's stream and the brain's record, each idempotent by execution id, the second retried; the ledger's port appends to one stream, so they are never one transaction. Reads over the ledger and Studio replace Temporal's UI. @@ -59,8 +61,8 @@ Tenant data is stored once, and a workflow needs no service beyond the server. ## Plan - The machine, on the contract and the DSL in `@beonauto/workflow-engine`. -- A conformance suite: the same workflow probes through a fake driver, the Node adapter and the workerd adapter, seeded from the server's `src/workflow-executions` tests. -- The adapters, a `cancel_execution` operation, and the workflow SDK's validators precompiled for Cloudflare. +- A conformance suite: the same workflow probes through a fake driver, the Node adapter and the hosted adapter, seeded from the server's `src/workflow-executions` tests. +- The adapters, a `cancel_execution` operation, and the workflow SDK's validators precompiled, since the hosted runtime allows no code generation. - The cutover: `pnpm dev` without Temporal's dev server, and the README. - Measurements before and after: the 15 recorded histories through the driver; inputs per second, and bytes per input against Temporal's 9.26 MiB for 40,000 inputs; timer lateness at p99; heap per live run; snapshot bytes per run. @@ -75,9 +77,9 @@ Branch `spike/engine-node`: - `spikes/node/results/timers-two-processes.json`: two processes double-fire 227 of 300 timers unless each claims a timer first. - `spikes/node/results/lost-write-repeat.json`: timers written through a second SQLite library to the ledger's file lost committed cancels in three runs of three. -Branch `spike/engine-cloudflare`: +The measurements in the hosted runtime are kept in the private repository: -- `spikes/cloudflare/results/fold.json`, `heap.json`: folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. Cloudflare documents 30 s of CPU per request by default; local workerd enforced no CPU limit at all: 35 s of CPU finished with no limit set, and 2 s with `cpu_ms = 50`. -- `spikes/cloudflare/results/timers.json`: alarms fire 5 ms late at p99; an alarm due while `wrangler dev` was stopped fired 15.6 s late, when it was restarted; a sweep re-armed one that had given up. -- `spikes/cloudflare/results/settlement.json`: the record was written exactly once, or given up as intended, under every injected D1 fault and crash; D1 refuses eleven events in one append. -- `spikes/cloudflare/results/portability.json`: the interpreter, the DSL policy, jq and both Cloudflare ledger drivers run in workerd; the workflow SDK's validators run once precompiled. +- Folding 40,000 events cold takes 239 ms and holds 66 MiB; from a snapshot every 1,000 events, 9 ms and 1.6 MiB. The hosted runtime documents 30 s of CPU for each wake-up by default; its local runtime enforced no CPU limit at all: 35 s of CPU finished with no limit set, and 2 s with a limit of 50 ms. +- Alarms fire 5 ms late at p99; an alarm due while the local runtime was stopped fired 15.6 s late, when it was restarted; a sweep re-armed one that had given up. +- The record was written exactly once, or given up as intended, under every injected fault of the shared database and every crash; the shared database refuses eleven events in one append. +- The interpreter, the DSL policy, jq and both hosted ledger drivers run in the isolate; the workflow SDK's validators run once precompiled. diff --git a/packages/ledger/README.md b/packages/ledger/README.md index ecc1d7823..b777c8695 100644 --- a/packages/ledger/README.md +++ b/packages/ledger/README.md @@ -44,9 +44,9 @@ Each SQLite connection may cache up to 8 MiB of pages and maps none of the file ## Portability -The same ledger will run on Cloudflare D1, so it follows these rules: +The same ledger must also run on hosted SQLite databases that bind at most 100 parameters in one statement and offer no interactive transactions, so it follows these rules: - It uses only the event store's own operations: read a stream, or its tail after a version, append with an expected version, migrate, close. It writes no SQL and registers no projections or consumers. - It does not rely on transactions or rollback: each command makes at most one append. -- An append carries at most eight events. Emmett binds ten parameters for each event it inserts, and D1 accepts at most 100 bound parameters in one query. +- An append carries at most eight events. Emmett binds ten parameters for each event it inserts, and such a database binds at most 100 in one statement. - Only `src/open-event-store.ts` knows which SQLite driver is in use and that the database is a file. diff --git a/packages/ledger/src/portable-entry.test.ts b/packages/ledger/src/portable-entry.test.ts index 5d2c6579b..28bb965a2 100644 --- a/packages/ledger/src/portable-entry.test.ts +++ b/packages/ledger/src/portable-entry.test.ts @@ -32,7 +32,7 @@ function packagesReachableFrom(entry: string): ReadonlySet { } describe('the ledger entry points', () => { - it('keep the main entry free of native and Node-only modules, so it loads on Cloudflare Workers', () => { + it('keep the main entry free of native and Node-only modules', () => { const packages = [...packagesReachableFrom('index.ts')]; expect(packages.filter((name) => name.startsWith('node:'))).toEqual([]); diff --git a/packages/workflow-engine/README.md b/packages/workflow-engine/README.md index afbc93b80..6669099ec 100644 --- a/packages/workflow-engine/README.md +++ b/packages/workflow-engine/README.md @@ -2,7 +2,7 @@ The contract of the workflow machine that runs on the ledger: the DSL it runs, the state it keeps, the inputs it takes and the ports to the adapters that store, time and execute for it. The machine's state is shaped by the workflow DSL, so this package is the workflow machine's contract, not a generic runtime. It knows workflows, their DSL and an executor that performs calls. It does not know brains, prompts, models, specs or primitives: the functions a workflow may call come from the caller, and whatever an adapter needs to know about a run, such as who started it, it passes as opaque `attributes` and gets back with every output. -The same code runs in Node, where one server keeps every run in one SQLite file, and in workerd, where each run is a Durable Object. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, with the DSL it imports from here. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. +The same code runs in Node, where one server keeps every run in one SQLite file, and in Auto's cloud hosting, the hosted runtime, where each run is an isolate of its own. The decision record states what the hosted runtime allows: a single-threaded isolate per run, woken by alarms that fire at least once and are dropped after a bounded number of failed retries, with about 128 MB of memory and bounded CPU per wake-up, no long-lived process and no code generation, and its own SQLite with rows of at most 2 MB; the shared database has no interactive transactions. The machine that decides an input is the next step, and the adapters the one after; until then the orchestration primitive runs workflows on Temporal, with the DSL it imports from here. [The decision record](../../docs/decisions/0001-workflow-engine-on-the-ledger.md) says why. ## Entries @@ -129,7 +129,7 @@ The machine checks only the size of a call's arguments, at most 264 KiB as JSON. A run keeps one counter of runs for each task reference, `state.runs`, so the third time a task runs its run is 3, and a call it starts has that run in its key. An adapter that gives a call its own execution, as a call to a spec has, derives that execution's id from the call key, so a call dispatched twice is one execution. -The event store deduplicates nothing: Emmett appends a message with an id it has seen before as a new message, on SQLite and on D1. Deduplication lives in the run's state. +The event store deduplicates nothing: Emmett appends a message with an id it has seen before as a new message, on the self-hosted SQLite and on the hosted runtime's shared database. Deduplication lives in the run's state. ## The dispatch watermark @@ -137,7 +137,7 @@ The watermark of a run is a stream version. Every output of every event at or be ## Serialisation -One input at a time for each run is the adapter's job. On Cloudflare it is the run's Durable Object, whose single thread takes one request at a time. On Node it is one process for each SQLite file, holding a lock per run in memory; a second process on the same file is outside the contract, and nothing claims or leases a run. A PostgreSQL adapter, later, will take a lease per run. Whatever slips past, the expected version of the append catches. +One input at a time for each run is the adapter's job. In the hosted runtime it is the run's isolate, whose single thread takes one request at a time. On Node it is one process for each SQLite file, holding a lock per run in memory; a second process on the same file is outside the contract, and nothing claims or leases a run. A PostgreSQL adapter, later, will take a lease per run. Whatever slips past, the expected version of the append catches. On Node the timers table goes through the ledger's own SQLite driver or lives in a separate file: written through a second SQLite library to the ledger's file, committed cancels were lost (`spikes/node/results/lost-write-repeat.json` on branch `spike/engine-node`). @@ -145,7 +145,7 @@ On Node the timers table goes through the ledger's own SQLite driver or lives in The record keeps, for each live run, its next due time, the earliest of its armed timers, and whether its dispatch fell behind (`RunDue`). The engine writes it while dispatching an event that arms or cancels a timer, idempotent by the event's version, and when a dispatch stops at a failure. A running run always has a timer armed, its deadline, so it is always due at some time. -`sweep(before)` asks `RecordStore.dueRuns(before)` for the runs due before that time or behind, and only those: for each it calls `wake`, and `Timers.sweep` with the run's armed timers, which arms again any the timer store lost. It folds no other run. The adapter sweeps every minute, the most often a Cloudflare cron trigger runs, and passes a time one minute ago, so a run is swept once its timer is a minute late: alarms fired 5 ms late at p99, and the one alarm due while `wrangler dev` was stopped fired 15.6 s late when it was restarted (`spikes/cloudflare/results/timers.json`); Node's timers fired 3.7 ms late at p99 (`spikes/node/results/timers-precision.json`). +`sweep(before)` asks `RecordStore.dueRuns(before)` for the runs due before that time or behind, and only those: for each it calls `wake`, and `Timers.sweep` with the run's armed timers, which arms again any the timer store lost. It folds no other run. The adapter sweeps every minute and passes a time one minute ago, so a run is swept once its timer is a minute late: alarms in the hosted runtime fired 5 ms late at p99, and the one alarm due while its local runtime was stopped fired 15.6 s late when it was restarted (measurements kept in the private repository); Node's timers fired 3.7 ms late at p99 (`spikes/node/results/timers-precision.json` on branch `spike/engine-node`). ## State and snapshots @@ -153,7 +153,7 @@ A run's state is plain JSON: no `Map`, `Set`, `Date`, `undefined`, class or func A snapshot is `{ format, executionId, version, historyBytes, state }`, the state folded from events 1 to `version`, written only once event `version` is durable; the run store keeps only the latest. A snapshot is due once the events since the last one take as many bytes as that snapshot did, and at least 1 MiB, so writing snapshots never costs more bytes than the history they cover. -A snapshot holds at most about 5.3 MiB: the held data (4 MiB: the values, the document and the frames), the events waiting in the inbox (1 MiB), the ids of the events received (1,024 of at most 256 characters), and the timers, calls and run counters, a few dozen bytes for each frame. It is stored in chunks of at most 1 MiB of UTF-8, cut by `TextEncoder.encodeInto` at a code point, never inside one; a 5.2 MB snapshot took 4 ms to encode and chunk. D1 and Durable Object SQLite both take rows of at most 2 MB; a 1 MiB chunk leaves room for the row's other columns, and an event, at most 1.5 MiB, fits in a row too. +A snapshot holds at most about 5.3 MiB: the held data (4 MiB: the values, the document and the frames), the events waiting in the inbox (1 MiB), the ids of the events received (1,024 of at most 256 characters), and the timers, calls and run counters, a few dozen bytes for each frame. It is stored in chunks of at most 1 MiB of UTF-8, cut by `TextEncoder.encodeInto` at a code point, never inside one; a 5.2 MB snapshot took 4 ms to encode and chunk. The hosted runtime's SQLite takes rows of at most 2 MB; a 1 MiB chunk leaves room for the row's other columns, and an event, at most 1.5 MiB, fits in a row too. ## Limits @@ -220,5 +220,5 @@ The machine's first test is invariant 33: every open call is answered, by the ex ## Open design points -- `evolve` decodes the whole state after each patch, so a load costs one decode of the state whatever the tail, and an applied input one more. The machine's step measures that against the 9 ms a fold from a snapshot every 1,000 events took in workerd (`spikes/cloudflare/results/fold.json` on branch `spike/engine-cloudflare`); checking only the patched paths is the fallback. +- `evolve` decodes the whole state after each patch, so a load costs one decode of the state whatever the tail, and an applied input one more. The machine's step measures that against the 9 ms a fold from a snapshot every 1,000 events took in the hosted runtime (measurements kept in the private repository); checking only the patched paths is the fallback. - What replaces Temporal's limits on a run's history is decided here as 100,000 inputs and 512 MiB, both well above what a workflow could reach on Temporal; real use may move them. diff --git a/packages/workflow-engine/src/engine/portability.test.ts b/packages/workflow-engine/src/engine/portability.test.ts index 96b2de24c..4abc82a64 100644 --- a/packages/workflow-engine/src/engine/portability.test.ts +++ b/packages/workflow-engine/src/engine/portability.test.ts @@ -87,7 +87,7 @@ const everySource = sourcesUnder('.'); const machineAndRunLog = [...sourcesUnder('machine'), ...sourcesUnder('run-log'), ...sourcesUnder('dsl')]; describe('the engine core', () => { - it('uses no Node-only API, no dynamic import, no code generation and no Temporal, so it runs in workerd as in Node', () => { + it('uses no Node-only API, no dynamic import, no code generation and no Temporal', () => { expect(everySource.length).toBeGreaterThan(20); expect(findingsIn(everySource, nodeOnly)).toEqual([]); }); From 323e07b11cc61bbd029af04785033e62eb4707d2 Mon Sep 17 00:00:00 2001 From: Rami Date: Sun, 4 Oct 2026 21:33:23 +0100 Subject: [PATCH 23/23] docs(global): say where the image runs without naming the host Where Auto's cloud hosting runs is internal to the private repo. This repository says only that the image runs self-hosted and in Auto's cloud hosting. Co-Authored-By: Claude Fable 5.1 --- CLAUDE.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CLAUDE.md b/CLAUDE.md index ddd9728c3..53de2498f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2,7 +2,7 @@ ## What this is -`auto-brain` is the runtime for business brains, a pnpm monorepo. It is source-available under the Elastic License 2.0 (see `LICENSING.md`); never call it open source. Its container image runs self-hosted and in Auto's cloud hosting on Cloudflare Containers. Auto Studio (the management plane, with governance and observability) and the Cloudflare side live in the private `on.auto` repo, not here. +`auto-brain` is the runtime for business brains, a pnpm monorepo. It is source-available under the Elastic License 2.0 (see `LICENSING.md`); never call it open source. Its container image runs self-hosted and in Auto's cloud hosting. Auto Studio (the management plane, with governance and observability) and the hosting side live in the private `on.auto` repo, not here. - `packages/server`: the Node.js server (`@beonauto/server`), plus everything that packages it into a container (`Dockerfile`, `Dockerfile.dockerignore`) - `packages/api`: the API (`@beonauto/api`), a Hono app that answers every request, with problem documents, the `Origin` and `Host` checks, authentication and the operation routes