HeadlinesBriefing HeadlinesBriefing.com

Readyset SQL Rewriting: Dataflow Query Pipeline

Hacker News •
×

Most databases re-execute queries from scratch each time, making read latency proportional to query complexity and data size. Readyset takes a different approach: it compiles each query into a dataflow graph—a network of operators (joins, filters, aggregations, projections) that continuously maintains the query's result as data changes. When a row is inserted, updated, or deleted upstream, the change propagates through the graph, and the cached result is incrementally updated. Reads become lookups into a pre-computed materialized view, not full query executions.

This architecture imposes structural constraints on SQL. Binary joins require equality predicates; range predicates or expression-based join keys are not supported. Correlated subqueries must be rewritten into equivalent joins, as the dataflow has no per-outer-row execution. Derived tables are supported but compile into fully materialized intermediate nodes; the rewrite pipeline aggressively inlines them where safe.

Supported join types include INNER, LEFT OUTER, and CROSS; RIGHT and FULL OUTER are not supported due to the difficulty of tracking absent matches. Aggregations require explicit GROUP BY column references and at least one aggregate-derived projection.

The tradeoff is that queries must be expressed in a form the dataflow engine can compile—more restrictive than standard SQL—but the payoff is incremental maintenance and faster reads.

Source: Hacker News · Summarized by HeadlinesBriefing