Repository navigation
Issue #5524 : Move the lake table transforms into a lakehouse plugin - #8792
Open
kotwal-itpro wants to merge 1 commit into
Open
kotwal-itpro wants to merge 1 commit into
kotwal-itpro wants to merge 1 commit into
Conversation
kotwal-itpro
force-pushed
the
lakehouse-plugin-move
branch
from
October 7, 2026 17:43
d479791 to
e145e72
Compare
…lugin The Spark lake table input, output, merge and maintenance transforms and the Spark catalog metadata type move from the Spark engine plugin into a new plugin, plugins/tech/lakehouse, so the local engine can implement them later. The Spark plugin keeps its handlers and SQL builders and loads the moved classes through dependencies.xml. No behavior change: plugin IDs, the spark-catalog metadata key, the serialized field keys and the i18n keys are unchanged, so existing pipelines and metadata load as before. Generated-by: Claude Opus 5.5
kotwal-itpro
force-pushed
the
lakehouse-plugin-move
branch
from
October 7, 2026 19:47
e145e72 to
a06d5bd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First step for #5524, following the plan agreed in the issue with @hansva and @mattcasters.
This moves the lake table transforms (input, output, merge, maintenance) and the catalog metadata type out of the Spark engine plugin into a new plugin,
plugins/tech/lakehouse. Nothing changes for users yet. The point is to give the local engine a place to implement the same transforms in the next PRs, so we end up with one set of transforms that runs on both engines.What moves
hop-tech-lakehouse)SparkLakeTable{Input,Output,Merge,Maintenance}{,Meta,Dialog,Data}org.apache.hop.lakehouse.transforms.LakeTable*SparkCatalog,SparkCatalogEditor,SparkCatalogTemplateorg.apache.hop.lakehouse.metadata.LakeCatalog*org/apache/hop/lakehouse/...SparkCatalogTemplateTestThe Spark plugin keeps everything Spark-specific: the
SparkLakeTable*Handlerclasses,SparkLakeTableSupport, the SQL builders,SparkCatalogApplierandLakeSessionPlan. It reaches the moved classes throughdependencies.xml(../../tech/lakehouse), the same way it already uses memgroupby, mergejoin and sort.Compatibility
Existing pipelines and metadata load unchanged:
SparkLakeTableInput,SparkLakeTableOutput,SparkLakeTableMerge,SparkLakeTableMaintenance). They now live inLakehouseConst, andSparkConstrefers to them.spark-catalog, withSparkCatalogstill as a legacy key, so existing objects stay in the same folder.new LakeTableInputMeta()+loadXml), so the Spark plugin having its own copy of the classes in its class loader causes no problems.LakeSessionPlan.loadMetahad a fallback that cast the in-memory meta when that XML load failed; a cast can't work across class loaders, so it now copies the meta through its serialized form instead (LakeSessionPlanCopyTest).Two small things I had to decide:
SparkFieldis also used by the Spark file input and SQL transforms, so it stays where it is. The lakehouse plugin gets its ownLakeFieldwith the same serialized keys, andSparkLakeTableSupport.toSparkFieldsconverts at the one place the Spark side needs it.LakeTableMergeMetaandLakeTableMaintenanceMeta.SparkMergeSqlBuilderandSparkMaintenanceSqlBuilderrefer to those, so the values are defined once.The new plugin is added to
assemblies/pluginsandassemblies/debug.Testing
hop-tech-lakehouse: 15 tests (meta injection for all four transforms, catalog templates).hop-engines-spark: all 177 tests pass, including the Iceberg and Delta path, table mode, time travel, merge and maintenance tests.apache-rat:checkandspotless:applyclean on both modules.Next
addresses #5524