Apache Hop version?
2.19.0 and current main (2.20.0-SNAPSHOT)
Java version?
21
Operating system
macOS
What happened?
What happens
With Spark lake table output in PATH mode, format iceberg, the table is not written at the table path. A pipeline that writes to file:///tmp/hop-it-iceberg/0001-table leaves nothing in /tmp/hop-it-iceberg/. The table ends up here instead:
${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/file:/tmp/hop-it-iceberg/0001-table/metadata/v1.metadata.json
Spark lake table input in PATH mode reads from the same place, so a write followed by a read works, and the existing integration-tests/iceberg scenarios pass. But anything else that looks at the path (another engine, another tool, a second Hop server) doesn't find the table, and the data sits in a temp folder that the OS may clean up.
Why
SparkLakeTableSupport.icebergPathSqlIdentifier() turns the path into hop_iceberg.`file:///tmp/.../0001-table`, and hop_iceberg is registered as a Hadoop catalog with its warehouse in java.io.tmpdir (LakeSessionPlan.defaultIcebergPathWarehouse()). Iceberg's SparkCatalog only treats an identifier as a location when it is a PathIdentifier, which Spark creates for spark.read.format("iceberg").load(path) but not for a quoted name in SQL or in writeTo(...). So the Hadoop catalog sees a table whose name is the whole URI, in the default namespace, and puts it at <warehouse>/<name>.
How to reproduce
- Run
integration-tests/iceberg/main-0001-input-output.hwf (or any pipeline with Spark lake table output, PATH mode, iceberg, on the native Spark engine).
- Look at the table path: there is no
metadata/ folder.
- Look under
${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/: the table is there, in a folder called file:.
Possible fixes
- Reads: use the path API,
spark.read().format("iceberg").option("snapshot-id", ...)/("as-of-timestamp", ...).load(path), which goes through PathIdentifier.
- Writes: give the Hadoop catalog the table's parent folder as its warehouse, so
warehouse/namespace/table resolves to the real path: for example a catalog per parent folder (hop_iceberg_<hash>), with the parent's last segment as the namespace and the folder name as the table. Or create the table at its location with HadoopTables first and then write with the path API.
I found this while adding the local engine read/write for #5524: an integration test that writes a PATH table with the Spark engine and reads it with the local engine fails because the table isn't at the path. I have a fix with tests ready and will open a PR for it, @mattcasters.
Issue Priority
Priority: 2
Issue Component
Component: Transforms
Apache Hop version?
2.19.0 and current main (2.20.0-SNAPSHOT)
Java version?
21
Operating system
macOS
What happened?
What happens
With Spark lake table output in PATH mode, format
iceberg, the table is not written at the table path. A pipeline that writes tofile:///tmp/hop-it-iceberg/0001-tableleaves nothing in/tmp/hop-it-iceberg/. The table ends up here instead:Spark lake table input in PATH mode reads from the same place, so a write followed by a read works, and the existing
integration-tests/icebergscenarios pass. But anything else that looks at the path (another engine, another tool, a second Hop server) doesn't find the table, and the data sits in a temp folder that the OS may clean up.Why
SparkLakeTableSupport.icebergPathSqlIdentifier()turns the path intohop_iceberg.`file:///tmp/.../0001-table`, andhop_icebergis registered as a Hadoop catalog with its warehouse injava.io.tmpdir(LakeSessionPlan.defaultIcebergPathWarehouse()). Iceberg'sSparkCatalogonly treats an identifier as a location when it is aPathIdentifier, which Spark creates forspark.read.format("iceberg").load(path)but not for a quoted name in SQL or inwriteTo(...). So the Hadoop catalog sees a table whose name is the whole URI, in the default namespace, and puts it at<warehouse>/<name>.How to reproduce
integration-tests/iceberg/main-0001-input-output.hwf(or any pipeline with Spark lake table output, PATH mode, iceberg, on the native Spark engine).metadata/folder.${java.io.tmpdir}/hop-iceberg-path-catalog-warehouse/: the table is there, in a folder calledfile:.Possible fixes
spark.read().format("iceberg").option("snapshot-id", ...)/("as-of-timestamp", ...).load(path), which goes throughPathIdentifier.warehouse/namespace/tableresolves to the real path: for example a catalog per parent folder (hop_iceberg_<hash>), with the parent's last segment as the namespace and the folder name as the table. Or create the table at its location withHadoopTablesfirst and then write with the path API.I found this while adding the local engine read/write for #5524: an integration test that writes a PATH table with the Spark engine and reads it with the local engine fails because the table isn't at the path. I have a fix with tests ready and will open a PR for it, @mattcasters.
Issue Priority
Priority: 2
Issue Component
Component: Transforms