Today we are shipping Polars 2.0. In the earlier announcement post we went through the rationale of the version bump. This post we will discuss what features 2.0 brings. Even though we didn't intend to make it a big feature release, it still packs a lot to get enthousiastic about.
Our initial version of out-of-core (spill-to-disk) support is enabled, a lot of very core performance improvements, first class SQL support, which together with the performance improvements has Polars leading Data Fusion and Duck DB in TPC-H and TPC-DS1 benchmarks, a new Map dtype, and stricter Polars on dtypes and explicitness, leading to faster feedback, and faster AI iteration.
Polars 2.0 will be the marking point where we will treat SQL as first class citizen. Polars SQL coverage has increased dramatically over last few months. In Polars 2.0, we want to enable that to more workloads, including SQL. To make this performant, we shipped many improvements to our optimizer and engine.
Calling collect on a Lazy Frame will now default to the streaming engine, leading to massive memory and performance improvements on most queries. Out-of-core (spill to disk) is now enabled by default. It starts spilling at ~80% of RAM.
Source: Hacker News · Summarized by HeadlinesBriefing