Repository navigation
Issue #8794 : Write Iceberg PATH tables at their path on the Spark engine - #8795
kotwal-itpro wants to merge 2 commits into
Conversation
…ark engine An Iceberg PATH table was addressed as hop_iceberg.`<uri>` in a single Hadoop catalog whose warehouse is in java.io.tmpdir. Iceberg only treats an identifier as a location when Spark builds it from load(path), so the catalog took the whole URI as a table name and stored the table under its own warehouse instead of at the path. Each folder that holds PATH tables now gets its own Hadoop catalog, hop_iceberg_<hash of the folder>, with that folder as warehouse, and a table is addressed as hop_iceberg_<hash>.`<name>`. The Hadoop catalog keeps such a table at <warehouse>/<name>, which is the table path. Reads, writes, merges, time travel and maintenance procedures all go through it. Generated-by: Claude Opus 5.5
|
Built the branch and ran the I also checked Iceberg PATH maintenance, which the tests don't cover: in a fresh session, Optional leftovers: the old Javadoc block above |
…ATH tables Address review: put the Javadoc of both ensureIcebergPathCatalog methods back on the right method, describe PATH identifiers as hop_iceberg_<hash>.`table` in the class and resolver Javadoc, drop the hop_iceberg hint from the PATH read and write errors, and update the troubleshooting table in the lakehouse guide. Generated-by: Claude Opus 5.5
|
Thanks @bamaer, and for checking maintenance on a real path too, that's the part the tests don't cover. I pushed a8b0877 with the leftovers: the Javadoc is back on the right |
Fixes #8794.
On the native Spark engine, Iceberg tables in PATH mode were never written at their path. They ended up under
${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/, in a sub-folder named after the whole table URI. Reads looked in the same place, so a write followed by a read worked, which is why the existing tests didn't notice.Cause
A PATH table was addressed as
hop_iceberg.`file:///data/orders`in one Hadoop catalog whose warehouse is in the temp folder. Iceberg'sSparkCatalogonly treats an identifier as a location when it is aPathIdentifier, which Spark creates forspark.read.format("iceberg").load(path), not for a quoted name in SQL orwriteTo(). So the Hadoop catalog saw a table calledfile:///data/ordersin the default namespace and stored it at<warehouse>/file:/data/orders.Fix
Each folder that holds PATH tables gets its own Hadoop catalog,
hop_iceberg_<hash of the folder>, with that folder as its warehouse, and the table is addressed by its folder name:hop_iceberg_1a2b3c4d5e6f.`orders`. A Hadoop catalog stores a table of the default namespace at<warehouse>/<name>, which is exactly the table path. The catalog is registered on the session the first time a table in that folder is used.Because every PATH caller goes through
SparkLakeTableSupport.icebergPathSqlIdentifier()andresolveLakeTargetSqlId(), this covers reads (including time travel), writes in every save mode, merges, and the maintenance procedures (CALL hop_iceberg_<hash>.system.expire_snapshots(table => 'orders')etc.).I looked at reading with
load(path)instead, but writes have no path equivalent that can also create a table, and merge and maintenance need a SQL identifier anyway, so a catalog per folder keeps one code path for everything.Upgrading
Tables written with 2.19 are still in the temp folder. The lakehouse guide now has a note on copying such a table to its path. They aren't moved automatically: the temp folder may already have been cleaned, and silently moving data around felt wrong. Happy to add an automatic fallback (read from the old location when the table isn't at the path) if you think it's worth it.
Testing
SparkLakeTableIcebergPathTestnow also checks that the table's metadata is at the table path and that nothing is written to the old warehouse. That assertion fails on main and passes with this change.hop-engines-sparkpass.-Pfull) and ranintegration-tests/iceberg(round trip, overwrite and time travel, merge) withhop-run: all pass, the tables are at/tmp/hop-it-iceberg/..., andhop-iceberg-path-catalog-warehouseis no longer created.The docs that described
hop_iceberg.uri(lakehouse guide, output, merge, maintenance, catalog pages) are updated.