Clinker promises bounded memory for finite batch jobs. No single mechanism delivers that. A budget, a ledger, an arbitrator, spill files, producer pauses and a memory-aware scheduler share the job. This page takes each one apart and gives you a control to try it. The simulator in section 05 runs them all at once.
An author sets one value, pipeline.memory.limit (default 512M; the --memory-limit flag overrides it). The arbitrator turns it into three watermarks:
resume_threshold, which must fall inside (0, 0.80); anything else is E324.E310 report only when that frees nothing. Three refusals run no round: a request made off the walk (a Source reading ahead, a writer, a join's own matching thread), one request larger than the whole limit, and Cull's per-group decisions.The 20% gap between soft and hard is the spike allowance. Spill reacts at batch boundaries, not on every allocation, so short overshoots have to fit somewhere.
Each stage that holds memory (Source ingest channels, hash Aggregates, sort buffers, grace-hash partitions, IEJoin arrays, Reshape and Cull group buffers, window arenas, and every inter-stage node_buffers slot) registers a MemoryConsumer with the run's single MemoryArbitrator.
A consumer holds a ConsumerHandle, a few atomics: live bytes, peak bytes, a spill-requested flag, a paused flag, and an active flag. The arbitrator never pushes numbers in. It pulls current_usage() from each consumer on every poll, so a grace-hash join with partitions on disk reports only what is still in memory.
A registration lasts as long as the state it describes. A Source's consumer is removed when its stream ends (its reader's end of input, or the walk stopping), and a buffer slot's when its last reader finishes. The registry therefore lists only live memory. The rows a finished Source read stay charged in its name until the steps holding them drop or spill them; the ledger keeps them as a finished Source's (retired_source), and an E310 report shows them as memory not held by any one node.
Newer storage (CSV decoding, prepared output, owned record storage) is stricter. An AllocationLease reserves the full Layout before anything is allocated. Growing a buffer reserves the new block while the old one is still charged. If the reservation is refused, the old contents and their charge stay as they were. Moving a grant between owners never leaves a moment where it is uncharged.
Allocate a vector, then double it. During growth, the old and new blocks are both charged at once, and this overlap is what most often hits the budget.
Every consumer has two fixed traits. Its spill priority (lower spills first) says how cheaply it can spill. Its can_back_pressure flag says whether its producer can be paused instead. Only Sources can be paused.
The memory.backpressure setting picks a policy:
| knob | policy object | picks |
|---|---|---|
pause (default) | BackPressurePreferred → Priority | the first pausable consumer, or else the lowest priority, most reclaimable on ties |
spill | Priority | lowest priority, most reclaimable on ties; never pauses |
both | BackPressurePreferred → LargestFirst | the first pausable consumer, or else whoever a spill would free the most from |
Both policies rank on reclaimable_bytes — what a spill of that consumer would free now, not what it holds charged. A consumer that reports 0, such as the output staging grant or a window arena, is never a candidate, so charged-only state never shadows one that can actually give memory back.
reconcile_backpressure handles pause and resume, reading the charged total. It skips any Source the walk thread is reading from right now (pausing that one would deadlock the run).
poll_arbitration asks the policy for a victim. If the victim cannot be paused, it sets the victim's spill flag, and the stage spills at its next batch boundary. If the victim is pausable, step ② does nothing, because step ① already handled pausing.So under pause, while any Source is registered, step ② picks that Source and nobody gets a spill request from the arbitrator. Spill still happens, through the stages' own thresholds, the slot admission check and the reclaim passes described in section 04.
Tick consumers to register them and set the bytes each holds. Each column runs the real selection code for one policy. Registration order matters for first pausable.
crossbeam channel is full, and streaming handoffs work the same way. This happens in the transport layer, and the arbitrator isn't involved.Condvar. Below the resume watermark it starts again. Between the two thresholds nothing changes (hysteresis).postcard frames, LZ4-compressed when the per-schema compress setting chooses to. Cumulative disk use is checked against storage.spill.disk_cap_bytes (E320).E310 report only when a pass freed nothing and a final pass freed nothing either. A request off the walk, one larger than the whole limit, and Cull's per-group decisions are refused at once, with no round. Probe loops also check every 10K output records.try_spill sets a flag and does no I/O. The stage reads the flag with take_spill_request() at its next boundary and spills there. Buffer slots are served in a sweep before the next node is dispatched.node_buffers slot is published and should_spill() returns true, the producer writes those rows straight to a spill file.spill_reclaimable(over), which runs one reclaim pass over the walk-owned victims a spill would free bytes from, in the policy's order. over is how far the larger of process RSS and the charged total (current_pressure()) sits above the soft limit. Only then does it resume the Source and mark it active, so the resumed Source doesn't push memory straight back over the soft limit.peak_rss only ever rises (fetch_max), so once the run has crossed the soft limit every later spill poll counts as under pressure; spilling too often is safe. Pause and resume read the charged total instead, which can fall, so a reader is never left parked and memory the process holds for other reasons never pauses one.| consumer | prio | pausable |
|---|---|---|
| node_buffers slot | 0 | no |
| output staging (writer) | 0 | no charged-only |
| grace-hash Combine | 10 | no |
| Reshape · Cull | 15 | no |
| sort buffer · IEJoin build | 20 | no |
| sort-merge Combine | 25 | no |
| hash Aggregate · inline-hash Combine | 30 | no |
| Source ingest | N/A | yes pause |
| streaming Aggregate | N/A | no |
| credential registry · scan materialization · window arena | last | no last resort |
The IEJoin sort buffer, the sort-merge Combine and the inline-hash Combine report nothing a spill would free now, so neither a poll nor a reclaim pass elects them: the two sort kernels spill on thresholds of their own, and the inline-hash table never spills. Their rows give the order they take once a pass can reach them.
Order follows reload cost. Re-reading a buffer slot is one postcard round-trip. A grace partition costs a little more. Reshape and Cull must re-split groups when they reload. A sort needs an external merge. A hash Aggregate costs the most to bring back.
Two Sources start ingesting together. The walk thread drains orders into a buffer slot, aggregates it, then drains refs into a grace-hash build, then probes and writes. Each tick runs the spill poll with its pause/resume reconcile, flag servicing, the walk's requests (which reclaim before they refuse), the join's own limit check once it builds its tables and as it probes, and the Source reads, which are refused at once when they do not fit, following the code described above. Rates, sizes and the RSS overhead model are made up for teaching, so read the shapes of the curves, not the numbers.
pause once no Source is registered; from then on it picks the biggest holder whatever its priority. Section 03 shows the difference directly. Oversized group shows the one shape spill can't fix: its request runs a reclaim pass and then a final pass, both free nothing, and only then is it refused. Join check reclaims shows where the grace-hash join checks the hard limit. It reads refs with no limit check, spilling its largest partition when the spill poll trips. Once refs is read, it builds a table for each partition still in memory, and the tables' index does not fit. Its check runs a reclaim pass, which spills the Aggregate, the index then fits, and the run completes.The walk runs one node to completion before starting the next. When several nodes are runnable together (independent chains feeding a later Combine), next_runnable picks one by minimising this tuple:
( !fits, −predicted_freed, −subtree_reclaim, stable_index )
predicted_peak fits within soft_limit − charged. An unknown peak (0B) always counts as fitting.--explain). Without them every term is 0, so the order is plain topological. Scheduling never changes output; it only affects how high memory peaks.A request did not fit. On the walk the engine first ran reclaim passes, and they freed nothing; off the walk, for one request larger than the whole limit, and for Cull's per-group decisions, no round ran. Common causes: one indivisible unit (a Reshape or Cull group, one aggregate row, a range-join block pair) is larger than the budget, or what fills the limit cannot be written to disk. The report names the node that asked and what the memory was for, the charged total beside private memory, the five largest holders with the state each is in, what the reclaim round asked and freed, a pasteable limit floor, and one remedy.
Under pause or both, a limit below the process's baseline RSS is rejected at startup. Otherwise the run would pause forever. spill skips this check and spills hard instead.
Cumulative spill passed storage.spill.disk_cap_bytes. That's your own limit; the volume may still have space. The file that went over the cap is deleted on abort.
The disk itself ran out. Also: tmpfs /tmp is RAM, so spilling there frees nothing. Point storage.spill.dir at a real disk.
The value has to fall strictly between 0 and the 0.80 soft limit, otherwise the hysteresis band doesn't exist.