Apache DataFusion: Nested Schema Nullability Adaptation in Aggregation
Engine-level query execution correctness fix resolving nested schema nullability mismatches across aggregation boundaries in Apache DataFusion without plan disruption.

PR #24394
merged upstream
2326917
merge commit
#24069
issue resolved
+1,093 / -3
1 commit · 4 files
01 · Engineering problem
The failure mode
When evaluating nested structures and dictionaries, input RecordBatches frequently arrive with tighter nullability than the planned schema due to runtime data distribution. In aggregation pipelines, schema divergence triggered strict schema invariant panics or failed physical expression evaluations.
02 · Technical response
What changed
- 01Engine-level query execution correctness fix resolving nested schema nullability mismatches across aggregation boundaries in Apache DataFusion without plan disruption.
- 02Input batches with stricter nested nullability are adapted cleanly to planned aggregate schemas.
- 03Preserves physical expression invariants and avoids plan disruption across aggregation boundaries.
- 04Added extensive regression test matrix covering nested structs, dictionaries, and mixed nullability states.
03 · Verified outcome
Evidence, not implication
- ✓Zero plan disruption: Input batches with tighter nullability adapt dynamically to planned aggregate schemas.
- ✓Nested schema safety: Struct, list, and dictionary nullability variants execute cleanly without schema invariant failures.
- ✓Comprehensive regression suite added with +1,093 lines of rigorous end-to-end and physical plan tests.
- ✓Merged upstream into apache/datafusion main branch (August 24, 2026).
04 · Public record
Source evidence
Independent engineering and open-source work by Patrick Ribbsaeter / Ribbsaeter Systems. References to maintainers, repositories, or companies identify the public technical context only. No employment, client, vendor, or partnership relationship is implied unless explicitly stated.