Clinker engine internals · ← Retraction Protocol

The retraction
loop

When an Aggregate's group_by leaves out a correlation-key field, its output rows no longer belong to one key group, so a failure after it cannot simply dead-letter "the group". The engine instead holds the Aggregate's downstream back, runs it at commit time, traces each failure to the source rows behind it, takes those rows out of the Aggregate, and runs again until nothing new fails. This page replays that loop, phase by phase, on three of the engine's own test pipelines.

01

When it runsRelaxed aggregates and the region held back for commit

Relaxed or strict

An Aggregate is relaxed when its group_by omits any field of its input's correlation-key set (group_by_omits_any_ck_field); time-windowed aggregates stay strict in the planner, though the run-time check does not exclude them yet (#1374). A relaxed Aggregate stamps a synthetic $ck.aggregate.<name> column holding the group index, and its output key set becomes (input key ∩ group_by) ∪ that column, so everything downstream is keyed by group. With no relaxed Aggregate the run takes the fast path: no loop, all retraction counters stay 0.

sum, count, avg, weighted_avg, collect and any retract by subtracting cached values (lineage mode). min and max keep each group's raw contributions and re-fold what survives (buffer mode). --explain prints a === Retraction === block with each relaxed Aggregate's mode and its deferred region.

The deferred region

From each relaxed Aggregate (the producer) the planner walks downstream; every node reached is a region member until a Sink, which becomes a region output. On the forward pass the producer runs, keeps its aggregator state for retraction, and parks its output; members and outputs do not run at all. Sinks outside a region write into correlation buffer cells keyed by the row's $ck.* values, and per-record failures are held in those cells. Rows a member reads from outside the region (another input of a Combine, a Route branch, a Cull port, a composition body's input) are parked once on the forward pass, charged against the memory limit and spillable, and every iteration of the loop reads them again in arrival order. Everything is decided at commit.

02

The loopDetect → recompute → dispatch → re-detect → expand or stop → flush

Test pipeline

Aggregator state

Correlation buffer

Retract set

Counters

Result

03

Known gapsPaths that do not behave as designed today

Degrade loses the group

When an Aggregate cannot be corrected (no retained state, a failed retract, a failed re-emit), its output slot is drained and it is added to a degrade list that nothing reads. The group's rows are lost rather than dead-lettered, and a later node that needed the drained output can stop the run with an internal error (#1288).

No spilling

A relaxed Aggregate whose state spills to disk cannot finalize in place; the run stops with an internal "spill failed" error rather than a budget diagnostic or the degrade path (#1288).

Windows before the Aggregate

Windows after the Aggregate are rebuilt from the re-emitted rows on every iteration. Windows before it are never re-evaluated, yet the planner's buffer-mode mark (which --explain counts) lifts the E150 check for them (#1376).

correlation_fanout_policy

all and primary are accepted on a Sink or the pipeline but behave like any: the one place collateral sparing is decided passes "full match" and "primary match" as always true, and the Combine-level field is never read (#1375).

The iteration cap

The loop is bounded by plan nodes (composition-body nodes included) + source rows + 1 iterations, since each pass must add at least one source row. Reaching the cap is a panic!, not a typed error (#1378).