Work

Query enginesUpstream2026

Faster, bounded query planning in Apache DataFusion

Removed two planner hot spots: repeated name resolution on wide queries and runaway regex compilation.

11.9×faster normalization on wide queries

  1. Parse

    • SQL to logical plan

  2. Resolve

    • Column normalization

      Context computed once, reused

  3. Optimize

    • Constant folding

      Expensive regexes deferred

  4. Execute

    • Physical plan

Where the two fixes sit in the planning pipeline.

Context

Apache DataFusion is the Rust query engine underneath a growing number of databases and analytics products. Planning time is invisible until a query is wide or adversarial, and then it dominates.

The problem

Two separate paths made planning far slower than it needed to be. Every unqualified column repeated the same plan traversal, so wide projections grew superlinearly. Separately, a pathological constant regex could spend the engine's full compilation budget while the query was still being planned.

My role

Profiled both paths, designed the fixes, and took them through maintainer review.

Approach

  • Computed the column-normalization context once per expression batch and reused it, keeping the existing fast path for already-qualified columns and literals.
  • Added a backward-compatible extension point that lets functions defer expensive constant evaluation from planning to execution.
  • Applied a planning-time limit to regex functions so ordinary patterns still fold while pathological ones keep their runtime behavior.

Outcome

  • Normalizing 2,000 unqualified expressions dropped from 74.7 ms to 6.3 ms in independent benchmarks published by a DataFusion contributor.
  • An issue-shaped planning benchmark went from 5.37 s to 0.17 s.
  • No public API changes; the new hook is available to other functions.
Next storyNew public extension APIs for Apache Maven 4