Skip to content

[Bug]: Native Spark engine writes Iceberg PATH tables under java.io.tmpdir instead of the table path #8794

Description

@kotwal-itpro

Apache Hop version?

2.19.0 and current main (2.20.0-SNAPSHOT)

Java version?

21

Operating system

macOS

What happened?

What happens

With Spark lake table output in PATH mode, format iceberg, the table is not written at the table path. A pipeline that writes to file:///tmp/hop-it-iceberg/0001-table leaves nothing in /tmp/hop-it-iceberg/. The table ends up here instead:

${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/file:/tmp/hop-it-iceberg/0001-table/metadata/v1.metadata.json

Spark lake table input in PATH mode reads from the same place, so a write followed by a read works, and the existing integration-tests/iceberg scenarios pass. But anything else that looks at the path (another engine, another tool, a second Hop server) doesn't find the table, and the data sits in a temp folder that the OS may clean up.

Why

SparkLakeTableSupport.icebergPathSqlIdentifier() turns the path into hop_iceberg.`file:///tmp/.../0001-table`, and hop_iceberg is registered as a Hadoop catalog with its warehouse in java.io.tmpdir (LakeSessionPlan.defaultIcebergPathWarehouse()). Iceberg's SparkCatalog only treats an identifier as a location when it is a PathIdentifier, which Spark creates for spark.read.format("iceberg").load(path) but not for a quoted name in SQL or in writeTo(...). So the Hadoop catalog sees a table whose name is the whole URI, in the default namespace, and puts it at <warehouse>/<name>.

How to reproduce

  1. Run integration-tests/iceberg/main-0001-input-output.hwf (or any pipeline with Spark lake table output, PATH mode, iceberg, on the native Spark engine).
  2. Look at the table path: there is no metadata/ folder.
  3. Look under ${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/: the table is there, in a folder called file:.

Possible fixes

  • Reads: use the path API, spark.read().format("iceberg").option("snapshot-id", ...)/("as-of-timestamp", ...).load(path), which goes through PathIdentifier.
  • Writes: give the Hadoop catalog the table's parent folder as its warehouse, so warehouse/namespace/table resolves to the real path: for example a catalog per parent folder (hop_iceberg_<hash>), with the parent's last segment as the namespace and the folder name as the table. Or create the table at its location with HadoopTables first and then write with the path API.

I found this while adding the local engine read/write for #5524: an integration test that writes a PATH table with the Spark engine and reads it with the local engine fails because the table isn't at the path. I have a fix with tests ready and will open a PR for it, @mattcasters.

Issue Priority

Priority: 2

Issue Component

Component: Transforms

Activity

  1. added 3 commits that reference this issue on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions