Skip to content

Add samples for all transforms to the samples project, fixes #2237 - #8786

Open
bamaer wants to merge 2 commits into
apache:mainfrom
bamaer:2237
Open

bamaer wants to merge 2 commits into
apache:mainfrom
bamaer:2237

Conversation

@bamaer

@bamaer bamaer commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Please add a meaningful description for your change here


Thank you for your contribution! Follow this checklist to help us incorporate your contribution quickly and easily:

  • Run mvn clean install apache-rat:check to make sure basic checks pass. A more thorough check will be performed on your pull request automatically.
  • If you have a group of commits related to the same change, please squash your commits into one and force push your branch using git rebase -i.
  • Mention the appropriate issue in your description (for example: addresses #123), if applicable.

To make clear that you license your contribution under the Apache License Version 2.0, January 2004
you have to acknowledge this by using the following check-box.

</transform>
<transform>
<name>database impact input</name>
<type>DbImpactInput</type>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] Database impact input only analyses a pipeline when filename-field names a field that holds a pipeline or workflow path. This transform never sets that property, and the only incoming value is the literal sample-value, so DbImpactInput skips the row and writes nothing. The note says the sample collects impact, but a run produces no impact rows.

Suggestion: Add a filename field, for example ${PROJECT_HOME}/transforms/table-exists.hpl, and set <filename-field> to that field.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. A Get variables transform now builds ${PROJECT_HOME}/transforms/tableoutput-basic.hpl into a filename field, and filename-field points at that field. I didn't use a data grid: DbImpactInput doesn't resolve variables in the filename value, so the literal ${PROJECT_HOME} path failed. The sample now returns 5 impact rows.

<method>none</method>
<schema_name/>
</partitioning>
<server_field>serverName</server_field>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] GetServerStatus only reads server_field after an input row arrives. This pipeline has no upstream transform, so processRow() gets a null row and finishes before it looks up serverName. The note says a reachable Hop server is required, but nothing is queried, and the sample is not on the integration-test ignore list.

Suggestion: Feed a server name from a data grid, and either ship a matching Hop-server metadata object or add get-server-status.hpl to the ignore list so a missing server does not fail read-samples-build-hop-run.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. A data grid feeds serverName = local-hop-server, and the sample ships metadata/server/local-hop-server.json (localhost:8080, cluster/cluster). When no server is running, the sample still finishes cleanly with available = N and the reason in errorMessage, so it stays off the ignore list.

<group_time/>
<parameters/>
<inherit_all_vars>Y</inherit_all_vars>
<execution_result_target_transform/>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] The note says the child rows are passed back as a data stream, but execution_result_target_transform is empty and there is no output-rows source transform. PipelineExecutor only copies child rows and execution statistics when those targets are set, so pipeline-executor-child.hpl runs and its child_value row is discarded.

Suggestion: Hop the executor to a downstream transform, set the execution-result target, and set the output-rows source to child data. Otherwise change the note so it does not claim the child rows are returned.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. The child now ends with Copy rows to result. In the parent, the execution results go to execution results and the result rows (child_value) go to child rows, both Dummy transforms. "Output rows source" isn't a parent setting: result_rows_target_transform only receives what the child copies to the result.

<item>requires Beam run configuration and a Hive metastore</item>
</line>
<line>
<item>structured-extract-basic.hpl</item>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nit] structured-extract-basic.hpl is already in this ignore list ("requires an AI provider named ollama-local"). This second row only repeats the exclusion with a slightly different reason.

Suggestion: Drop the duplicate row.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the duplicate row.

…ne executor samples, drop a duplicate ignore row
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants