Skip to content

Issue #5524 : Move the lake table transforms into a lakehouse plugin - #8792

Open
kotwal-itpro wants to merge 1 commit into
apache:mainfrom
kotwal-itpro:lakehouse-plugin-move
Open

kotwal-itpro wants to merge 1 commit into
apache:mainfrom
kotwal-itpro:lakehouse-plugin-move

Conversation

@kotwal-itpro

@kotwal-itpro kotwal-itpro commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

First step for #5524, following the plan agreed in the issue with @hansva and @mattcasters.

This moves the lake table transforms (input, output, merge, maintenance) and the catalog metadata type out of the Spark engine plugin into a new plugin, plugins/tech/lakehouse. Nothing changes for users yet. The point is to give the local engine a place to implement the same transforms in the next PRs, so we end up with one set of transforms that runs on both engines.

What moves

From (Spark plugin) To (hop-tech-lakehouse)
SparkLakeTable{Input,Output,Merge,Maintenance}{,Meta,Dialog,Data} org.apache.hop.lakehouse.transforms.LakeTable*
SparkCatalog, SparkCatalogEditor, SparkCatalogTemplate org.apache.hop.lakehouse.metadata.LakeCatalog*
Messages and icons for the above same, under org/apache/hop/lakehouse/...
Meta injection tests and SparkCatalogTemplateTest moved with the classes

The Spark plugin keeps everything Spark-specific: the SparkLakeTable*Handler classes, SparkLakeTableSupport, the SQL builders, SparkCatalogApplier and LakeSessionPlan. It reaches the moved classes through dependencies.xml (../../tech/lakehouse), the same way it already uses memgroupby, mergejoin and sort.

Compatibility

Existing pipelines and metadata load unchanged:

  • Transform plugin IDs keep their values (SparkLakeTableInput, SparkLakeTableOutput, SparkLakeTableMerge, SparkLakeTableMaintenance). They now live in LakehouseConst, and SparkConst refers to them.
  • The catalog keeps the metadata key spark-catalog, with SparkCatalog still as a legacy key, so existing objects stay in the same folder.
  • XML and metadata property keys are unchanged, and so are the i18n keys. Display names still say "Spark lake table ..." for now. I'd rather rename them in the PR that makes them work on the local engine, so this one stays a pure move.
  • The Spark handlers already rebuild each meta from the transform XML (new LakeTableInputMeta() + loadXml), so the Spark plugin having its own copy of the classes in its class loader causes no problems. LakeSessionPlan.loadMeta had a fallback that cast the in-memory meta when that XML load failed; a cast can't work across class loaders, so it now copies the meta through its serialized form instead (LakeSessionPlanCopyTest).

Two small things I had to decide:

  • SparkField is also used by the Spark file input and SQL transforms, so it stays where it is. The lakehouse plugin gets its own LakeField with the same serialized keys, and SparkLakeTableSupport.toSparkFields converts at the one place the Spark side needs it.
  • The merge action and maintenance operation constants moved onto LakeTableMergeMeta and LakeTableMaintenanceMeta. SparkMergeSqlBuilder and SparkMaintenanceSqlBuilder refer to those, so the values are defined once.

The new plugin is added to assemblies/plugins and assemblies/debug.

Testing

  • hop-tech-lakehouse: 15 tests (meta injection for all four transforms, catalog templates).
  • hop-engines-spark: all 177 tests pass, including the Iceberg and Delta path, table mode, time travel, merge and maintenance tests.
  • apache-rat:check and spotless:apply clean on both modules.

Next

  1. Local engine read support for Iceberg.
  2. Local engine write support for Iceberg.
  3. Integration tests for the local engine, docs and samples.

addresses #5524

@kotwal-itpro
kotwal-itpro force-pushed the lakehouse-plugin-move branch from d479791 to e145e72 Compare October 7, 2026 17:43
…lugin

The Spark lake table input, output, merge and maintenance transforms and
the Spark catalog metadata type move from the Spark engine plugin into a
new plugin, plugins/tech/lakehouse, so the local engine can implement
them later. The Spark plugin keeps its handlers and SQL builders and
loads the moved classes through dependencies.xml.

No behavior change: plugin IDs, the spark-catalog metadata key, the
serialized field keys and the i18n keys are unchanged, so existing
pipelines and metadata load as before.

Generated-by: Claude Opus 5.5
@kotwal-itpro
kotwal-itpro force-pushed the lakehouse-plugin-move branch from e145e72 to a06d5bd Compare October 7, 2026 19:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant