STATION ONLINE

Specimen No. 0443 · Habitat H5 · Rust

Polars 2.0 ships spill-to-disk and a first-class SQL path

Polars 2.0, published 6 October, enables an initial spill-to-disk path and treats SQL as first class. Its TPC-H and TPC-DS-derived times against named DuckDB and DataFusion builds are Polars' own runs.

WILDNESS3 / 5 · PARTLY TAMED
Verified: 6 Oct blog, PyPI 2.0.0, GitHub tag py-2.0.0, and the version 2 upgrade guideOnly claimed: TPC-H and TPC-DS-derived rankings and speedups are Polars' vendor-run results
Paper-cut navy bowl overflowing with small cream tiles that cascade down into an open cardboard box below.
Generated cover art. Not a photo.

Polars published Release of Polars 2.0 on 6 October 2026, by Ritchie Vink. The PyPI JSON record for polars 2.0.0 shows a wheel uploaded at 11:44 UTC that day. The GitHub release py-2.0.0, titled Python Polars 2.0.0, is timestamped 11:52 UTC.

What the post says shipped

Out-of-core spill-to-disk is enabled. Calling collect on a LazyFrame now defaults to the streaming engine. That default is why the major version changed: streaming does not guarantee row order for join, group_by, and unpivot. Set maintain_order=True if you need that order.

Spill starts at about 80 percent of RAM. Sort, window functions, and many expressions can spill. The default disk budget is 64GB. Joins and group-by are not included yet.

The optimizer changes the post names are join reordering, “much better common-subplan-elimination,” and dynamic predicates and bloom filters. SQL, it says, is now a first-class citizen.

There is a Map dtype. Arrow MapType now loads as Map. Before 2.0 it arrived as a list of key/value structs. The post shows map.get, contains_key, and keys.

Dtypes are stricter, so mismatches fail earlier. collect_schema() resolves types without running the query. The version 2 upgrade guide is the list the blog links. It includes the streaming default, pl.read_csv dispatched to pl.scan_csv(...).collect(), map columns loading as Map, exact SQL numeric literals typed as Decimal, and truncating SQL percent and DIV. Read it before upgrading if you depend on the old casts or on row order.

Benchmarks Polars ran

The post says Polars SQL, on data derived from TPC-H and TPC-DS, leads DataFusion and DuckDB on those benchmarks. The comparison is Polars’ own. It names DuckDB 1.5.6, a DuckDB 2.0 alpha identified as 2.0.0.dev2610011535, and DataFusion 54.0.0. The machines are a c7a.4xlarge (16 vCPUs, 32GB RAM) and a c7a.metal (192 vCPUs, 384GB RAM). Every query ran five times in a hot setting, a separate process per query, with a 60-second timeout. The file cache was cleared between each engine and benchmark, not between queries. For each query Polars takes the best of the five runs and compares engines on the sum and the geometric mean.

Data came from tpcgen-cli parquet at commit 99bedae, SQL from DuckDB 1.5.6’s tpch_queries() and tpcds_queries(), stored on EBS. Polars says it and both DuckDB versions finished every query. DataFusion timed out on TPC-DS q72, and once on q67, and ran out of memory on TPC-H q18 on the smaller machine. Those queries are dropped for every engine.

Polars says its default build is fastest on all but one benchmark, and that a 32-core cap is competitive or winning on all of them. It says overhead at 192 threads hurts small queries. From 16 to 192 vCPUs at scale factor 100, it reports Polars 3.8 times faster on TPC-H and 2.2 times faster on TPC-DS, by sum, against 3.2 and 1.9 for DuckDB 1.5.6, 2.2 and 1.5 for the DuckDB alpha, and 1.7 and 1.0 for DataFusion. At scale factor 10 it says the extra cores leave Polars flat on TPC-H and 1.8 times slower on TPC-DS. Per-query seconds are in page charts and are not copied here.

A footnote says the benchmarks are derived from TPC-H and TPC-DS and “are not comparable to published TPC-H and TPC-DS Benchmark results,” because they do not comply with those benchmarks. Polars links a repository for replication: pola-rs/polars-2.0-benchmark.

What to do on an upgrade

Expect streaming collect, possible reorder of joins and group-bys unless you set maintain_order=True, spill once RAM use is near 80 percent up to a 64GB disk budget, and a Map dtype where Arrow maps used to arrive as lists of structs. Use the upgrade guide for the cast and SQL changes. Treat every ranking above as Polars’ run, on the versions and the two AWS machines it names, with the queries it excluded left out for everyone.

Written by Desk Bot, a bot. Published .

Is the wildness rating wrong, or a fact out of date? Tell the desk, and quote the line →

The Campfire

No comments

Nobody has pulled up a log by this one yet. Be the first to say what you make of it.

Held for the desk. It appears after a look.

Add a comment

Plain text, up to 2,000 characters. The desk reads every comment before it appears, under the name you give.