Skip to content

Issue #8794 : Write Iceberg PATH tables at their path on the Spark engine - #8795

Open
kotwal-itpro wants to merge 2 commits into
apache:mainfrom
kotwal-itpro:fix-spark-iceberg-path
Open

kotwal-itpro wants to merge 2 commits into
apache:mainfrom
kotwal-itpro:fix-spark-iceberg-path

Conversation

@kotwal-itpro

Copy link
Copy Markdown
Contributor

Fixes #8794.

On the native Spark engine, Iceberg tables in PATH mode were never written at their path. They ended up under ${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/, in a sub-folder named after the whole table URI. Reads looked in the same place, so a write followed by a read worked, which is why the existing tests didn't notice.

Cause

A PATH table was addressed as hop_iceberg.`file:///data/orders` in one Hadoop catalog whose warehouse is in the temp folder. Iceberg's SparkCatalog only treats an identifier as a location when it is a PathIdentifier, which Spark creates for spark.read.format("iceberg").load(path), not for a quoted name in SQL or writeTo(). So the Hadoop catalog saw a table called file:///data/orders in the default namespace and stored it at <warehouse>/file:/data/orders.

Fix

Each folder that holds PATH tables gets its own Hadoop catalog, hop_iceberg_<hash of the folder>, with that folder as its warehouse, and the table is addressed by its folder name: hop_iceberg_1a2b3c4d5e6f.`orders`. A Hadoop catalog stores a table of the default namespace at <warehouse>/<name>, which is exactly the table path. The catalog is registered on the session the first time a table in that folder is used.

Because every PATH caller goes through SparkLakeTableSupport.icebergPathSqlIdentifier() and resolveLakeTargetSqlId(), this covers reads (including time travel), writes in every save mode, merges, and the maintenance procedures (CALL hop_iceberg_<hash>.system.expire_snapshots(table => 'orders') etc.).

I looked at reading with load(path) instead, but writes have no path equivalent that can also create a table, and merge and maintenance need a SQL identifier anyway, so a catalog per folder keeps one code path for everything.

Upgrading

Tables written with 2.19 are still in the temp folder. The lakehouse guide now has a note on copying such a table to its path. They aren't moved automatically: the temp folder may already have been cleaned, and silently moving data around felt wrong. Happy to add an automatic fallback (read from the old location when the table isn't at the path) if you think it's worth it.

Testing

  • SparkLakeTableIcebergPathTest now also checks that the table's metadata is at the table path and that nothing is written to the old warehouse. That assertion fails on main and passes with this change.
  • New unit tests for the folder/table mapping: the parent folder is the warehouse, tables in one folder share a catalog, other folders get another, and a path without a parent folder is rejected.
  • All 195 tests of hop-engines-spark pass.
  • Built the full client (-Pfull) and ran integration-tests/iceberg (round trip, overwrite and time travel, merge) with hop-run: all pass, the tables are at /tmp/hop-it-iceberg/..., and hop-iceberg-path-catalog-warehouse is no longer created.

The docs that described hop_iceberg.uri (lakehouse guide, output, merge, maintenance, catalog pages) are updated.

…ark engine

An Iceberg PATH table was addressed as hop_iceberg.`<uri>` in a single
Hadoop catalog whose warehouse is in java.io.tmpdir. Iceberg only treats
an identifier as a location when Spark builds it from load(path), so the
catalog took the whole URI as a table name and stored the table under
its own warehouse instead of at the path.

Each folder that holds PATH tables now gets its own Hadoop catalog,
hop_iceberg_<hash of the folder>, with that folder as warehouse, and a
table is addressed as hop_iceberg_<hash>.`<name>`. The Hadoop catalog
keeps such a table at <warehouse>/<name>, which is the table path.
Reads, writes, merges, time travel and maintenance procedures all go
through it.

Generated-by: Claude Opus 5.5
@bamaer

bamaer commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Built the branch and ran the hop-engines-spark tests: all pass. The new assertion in SparkLakeTableIcebergPathTest fails against main's SparkLakeTableSupport, so it covers the bug.

I also checked Iceberg PATH maintenance, which the tests don't cover: in a fresh session, REWRITE_MANIFESTS and EXPIRE_SNAPSHOTS hit the table at its path. Looks good to me.

Optional leftovers: the old Javadoc block above ensureIcebergPathCatalog(SparkSession, String), the hop_iceberg mention in the readIcebergPath error message, and lakehouse.adoc:295.

…ATH tables

Address review: put the Javadoc of both ensureIcebergPathCatalog methods
back on the right method, describe PATH identifiers as
hop_iceberg_<hash>.`table` in the class and resolver Javadoc, drop the
hop_iceberg hint from the PATH read and write errors, and update the
troubleshooting table in the lakehouse guide.

Generated-by: Claude Opus 5.5
@kotwal-itpro

Copy link
Copy Markdown
Contributor Author

Thanks @bamaer, and for checking maintenance on a real path too, that's the part the tests don't cover.

I pushed a8b0877 with the leftovers: the Javadoc is back on the right ensureIcebergPathCatalog method (the old one now says it's for two-part TABLE identifiers), the class and resolver Javadoc describe hop_iceberg_<hash>.\table`, the PATH read and write errors no longer mention hop_iceberg, and the troubleshooting row at lakehouse.adoc:295is updated. Thehop_iceberghint in the TABLE-mode error stays, since two-part TABLE identifiers still use that catalog.hop-engines-spark` tests still pass (195).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Native Spark engine writes Iceberg PATH tables under java.io.tmpdir instead of the table path

2 participants