Skip to content

[Feature Request]: Pipeline-level field type defaults, so type and length don't have to be re-entered in every transform #8766

Description

@sfkeller

What would you like to happen?

Description

As a pipeline developer, I want to define the data type, length and precision of a field once per pipeline, so I can avoid re-entering the same definition in every transform that creates or redefines that field.

Current situation

Fields that only pass through a transform keep their metadata. But every transform that creates or redefines a field (Calculator, Formula, JavaScript, Add constants, Concat fields, Value mapper, Regex evaluation, and others) needs its own type, length and precision in its dialog.

In a pipeline with around 20 transforms, a field such as CUSTOMER_NAME that should be String(80) has to be configured in almost every one of them. For each transform I must open the dialog, go to the field list and enter the type again. This is repetitive and error-prone. Inconsistent lengths can silently lead to truncation or mismatches further downstream.

Proposal

Add an optional pipeline-level field definition (a "field schema" or "default field types") that maps a field name to a type, length and precision. When a transform creates or redefines a field and no explicit type is set, it would take the definition from this mapping. Possible refinements:

  • Pre-fill the type, length and precision in the transform dialog from the mapping, so the user can still see and override them.
  • Offer a "learn from first occurrence" option, so the first explicit definition of a field becomes the default for later transforms in the same pipeline (this is how e.g. FME (Safe Inc.) handles attributes).
  • Keep explicit settings in a transform as the highest priority, so existing pipelines behave exactly as before.

Goal

Define a field's type once and reuse it consistently across the pipeline, with fewer clicks, less duplicated configuration and fewer type mismatches.

Current workarounds (and their limits)

  • A Select values transform with a Meta-data tab right after the source. This only helps for pass-through fields, not for fields created later.
  • Overwriting existing fields instead of creating new ones. This is not always possible.
  • Variables such as ${NAME_LEN} in the length column. This still requires touching every transform once.
  • Metadata injection. This is heavy for a simple pipeline.

Additional information

  • Hop version: 2.18.1 (latest release at time of writing)
  • Java version: Java 17 (the minimum Hop requires)
  • OS: Windows 11 Enterprise
  • Happy to discuss the design in the dev chat or on the mailing list first, and to test a prototype.

Issue Priority

Priority: 3

Issue Component

Component: Hop Gui

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions