Skip to content

[Feature] Native pivot / group-by-column dataset transform for long-format records #21742

Description

@namannitr

What problem does this feature solve?

ECharts dataset is strongest when the source is already a pivot table (one row per category, one column per series). A large share of real data — SQL result sets, CSV exports, analytics event logs — arrives in long / tidy / record form instead:

dataset: {
  dimensions: ['Country', 'Year', 'Population'],
  source: [
    ['Brazil', 2011, 18203],
    ['Brazil', 2012, 19325],
    ['Indonesia', 2011, 23489],
    ['USA', 2011, 29034]
  ]
}

There is still no first-class way to turn that into N series (one per Year, or one per Country) without either:

  • pre-pivoting outside ECharts, or
  • one filter transform per distinct group (O(groups × rows)), which does not scale.

This was originally described in #16083 (closed stale; still requested in 2025–2026). @100pah agreed a pivot transformer belongs in the transform pipeline, possibly piped after aggregate.

Built-in filter / sort cannot express this. Third-party echarts-simple-transform only aggregates one groupBy dimension and does not reshape long → wide. Draft PR #16903 migrates aggregate/id into core but does not add pivot / “split into series by column”.

We hit this in production SQL → option pipelines: one query returns tidy rows, several cartesian/polar series need to share that table, and encode cannot invent series from a grouping column.

Related: #15306 (built-in aggregate; still open), #16903 (draft migration, stale).

This is not an alternative to an existing API — it is the missing reshape that makes dataset + encode usable for the format databases actually emit.

What does the proposed API look like?

Keep it a single-upstream dataset.transform, consistent with filter/sort and with @100pah’s note on #16083 (standalone, or piped after aggregate).

1. Split long data into multiple series (the #16083 bar case)

dataset: [
  { id: 'raw', source: /* Country, Year, Population */ },
  {
    fromDatasetId: 'raw',
    transform: {
      type: 'pivot', // or 'group'
      config: {
        row: 'Country',
        column: 'Year',
        value: 'Population',
        fill: 0
      }
    }
  }
]
// result dimensions: Country, 2011, 2012, ...
series: [
  { type: 'bar', encode: { x: 'Population', y: 'Country' } } // plus one series per year via encode, or
]

A companion option for “one series per distinct value of column” (so callers do not have to list years) would close the remaining gap:

series: {
  type: 'bar',
  datasetId: 'raw',
  encode: { x: 'Population', y: 'Country' },
  // proposed — not currently valid
  groupBy: 'Year'
}

series.groupBy is the better UX for bar/line/scatter; dataset.transform: pivot is the better fit for the existing transform plugin model and for heatmap/treemap that want a wide table. Both can share one implementation.

2. Pipe after aggregate (#15306)

Long data with duplicates first:

dataset: [
  { source: rows },
  { transform: { type: 'aggregate', config: { groupBy: ['Country', 'Year'], output: [{ from: 'Population', method: 'sum' }] } } },
  { transform: { type: 'pivot', config: { row: 'Country', column: 'Year', value: 'Population' } } }
]

3. What this should not try to do

Cross-dataset joins (sankey nodes+edges, heatmap cells+axis catalogs) stay outside this transform — a pivot only sees one upstream dataset, which is the right scope for core. Multi-source composition belongs in application code.

Prior art

Happy to iterate on names (pivot vs spread vs group) and on whether series-level groupBy is in scope for a first RFC. I am not opening a core PR until there is maintainer agreement on the shape — per the contributing wiki this is RFC-sized, not a drive-by patch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enThis issue is in EnglishpendingWe are not sure about whether this is a bug/new feature.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions