Dev.to
7/22/2026

The original title is "One Histogram, Four Engines: Fixing the Statistics Silo Problem"
Original: One Histogram, Four Engines: Fixing the Statistics Silo Problem
Short summary
DuckDB, DataFusion, Polars, and Postgres each compute and store table statistics independently, so histograms built in ELT pipelines are invisible to other query engines. samkhya serializes sketches like HyperLogLog and equi-depth histograms into versioned Iceberg Puffin sidecars that any adapter can load unchanged. It also clamps corrections with a provable pessimistic bound so stale or wrong sketches can never make a query plan worse than default estimates.
- •Statistics are engine-local cache, not shared data assets — causing redundant computation and drift across engines
- •samkhya serializes sketches into portable Iceberg Puffin sidecars loadable by any adapter
- •Pessimistic bounding ensures stale or incorrect sketches never degrade optimizer plans
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



