I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I'm really aware of is if you're working with geospatial data, as there isn't yet a "Geopolars" equivalent of the commonly used "Geopandas". However, Geopolars is still in active development and should eventually be production ready.
We were able to do a full port of GFQL from pandas to polars, cypher graph queries on dataframes, including both our CPU + GPU modes, and hit massive speedups: https://www.graphistry.com/blog/cypher-on-polars-cpu-gpu-gra...
It's been impressive!
Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark
It has been for me. I greatly prefer the API, it fits my mental model much better. Give it a try!
my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.
See "Pandas should go extinct": https://news.ycombinator.com/item?id=49668198
tl;dr yes
There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there's no way to re-export that out of the Polars dataframe with nanosecond resolution.
From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)
> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory
Python Vs Rust : In terms for speed - No comparison
(The above episode transcript has a link to blog post titled "Pandas should go extinct" )