TL;DR
Get networking and server gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Polars says version 2.0 makes its streaming engine the default for LazyFrame collection and enables initial out-of-core processing that can spill supported operations to disk. The release also expands SQL support and adds a Map data type; performance claims are based on benchmarks published by the Polars team, not an independent comparison.
Polars has released version 2.0, making its streaming engine the default when users collect a LazyFrame and enabling initial out-of-core processing that can spill supported work to disk. The release also elevates SQL support and adds a Map data type, but changes observable row-order behavior for some operations unless users opt to maintain order.
The Polars team says calling collect on a LazyFrame now uses the streaming engine by default, which it says can reduce memory use and improve performance on many queries. The change affects operations including joins, group-bys and unpivots: the engine does not guarantee their row order by default. Users who need that behavior can request maintain_order=True where supported.
Version 2.0 also enables spill-to-disk support by default. According to the release post, spilling begins at about 80% of available RAM, with a default disk budget of 64 GB. The currently supported operations include sorts, window functions and many expressions. The team says support for joins and group-bys is planned, but those operations are not yet listed as supported by the initial release.
The release adds a native Map dtype for Arrow MapType data, representing key-value collections in a form similar to a dictionary. It also includes optimizer and engine changes, including join reordering, improved common-subplan elimination and dynamic predicates or bloom filters. Polars describes stricter handling of data types and explicitness as another change intended to provide faster feedback during development.
Streaming Changes Query Behavior
The shift to streaming as the default can affect both memory use and the assumptions applications make about results. For workloads that exceed available RAM, disk spilling may allow some supported queries to finish instead of relying entirely on memory. But users should check order-sensitive downstream code: the release explicitly warns that some operations do not preserve observable row order by default.
For teams choosing a data-processing engine, Polars is positioning version 2.0 for a broader range of workloads, including SQL. Its published benchmark results may inform evaluations, but they are vendor-run comparisons under specified hardware and test conditions. They should not be treated as a guarantee that Polars will be faster for every dataset, query or deployment.
high performance data processing laptop
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
SQL Benchmarks and Test Limits
Polars says it tested its SQL engine on queries derived from TPC-H and TPC-DS, comparing it with DuckDB 1.5.6, DuckDB 2.0 alpha and DataFusion 54.0.0. Tests ran on two AWS machine types: a c7a.4xlarge with 16 virtual CPUs and 32 GB of RAM, and a c7a.metal with 192 virtual CPUs and 384 GB. The team reports that each query ran five times in a hot setting and that it compared the best run for each query.
These results have qualifications. Polars says all queries were completed by Polars and both DuckDB versions, while DataFusion timed out on TPC-DS query 72, timed out once on query 67 and ran out of memory on TPC-H query 18 on the smaller machine. Those queries were excluded from the results for every engine. The release post says default Polars was fastest on all but one benchmark, but also reports overhead when scaling to 192 threads that hurt small-data queries; Polars limited to 32 cores was competitive or faster across the benchmarks described. The team says it has diagnosed that issue and hopes to address it in a later release.
Polars also provides a public repository for its benchmark setup and encourages readers to replicate the results. The benchmark uses generated Parquet data, specific cloud hardware and defined run conditions, so the reported ranking applies to that test rather than every production workload.
As an affiliate, we earn on qualifying purchases.
Workloads Still Need Testing
The release post does not give a calendar date in the supplied material, nor does it establish how much memory or runtime improvement users should expect across real-world workloads. The benchmark findings come from Polars’ own test setup; independent replication and results on other datasets, storage systems and hardware may differ.
The team describes the spill threshold as approximate and says it may need tuning. Its post does not specify when out-of-core support for joins and group-bys will arrive. It also says the 192-thread overhead has been diagnosed, but provides no confirmed fix date. Users should test the new defaults against their own queries, especially where row ordering matters.
As an affiliate, we earn on qualifying purchases.
More Spill Support Planned
Polars says it plans to extend out-of-core processing to joins and group-bys, broadening the range of workloads that can spill intermediate data to disk. The release post does not give a schedule for that work.
The team also hopes to address the overhead it observed when scaling to 192 threads in a subsequent release. In the meantime, users can inspect the published benchmark repository, replicate the tests and validate version 2.0 against their own performance, memory and row-order requirements.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main change in Polars 2.0?
LazyFrame collection now defaults to the streaming engine. Polars also enables initial spill-to-disk support and expands its SQL and data-type features.
Does Polars 2.0 preserve row order?
Not by default for some operations, including joins, group-bys and unpivots, according to the release post. Users who need observable order can use maintain_order=True where supported.
Which operations can spill to disk in this release?
The release lists sorts, window functions and many expressions as supported. The Polars team says support for joins and group-bys is planned, but not yet available in the initial out-of-core implementation.
Did Polars independently verify its benchmark claims?
The results in the release post are from tests conducted by the Polars team. It shared a benchmark repository and invited replication, but the supplied material does not report independent verification.
When will out-of-core joins and group-bys arrive?
Polars says those capabilities are on its roadmap but does not provide a release date in the version 2.0 post.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
