diff --git a/r/R/dataset.R b/r/R/dataset.R index 4ccf338d267e..1cb136dacf57 100644 --- a/r/R/dataset.R +++ b/r/R/dataset.R @@ -64,6 +64,23 @@ #' column types, as described above. If neither are provided, no partitioning #' information will be taken from the file paths. #' +#' @section Adding the source filename as a column: +#' +#' Partitioning only recovers information encoded in directory names. If you +#' need to know which file each row came from, call [add_filename()] inside a +#' `dplyr` query on the dataset: +#' +#' ```r +#' open_dataset("nyc-taxi") |> +#' mutate(file = add_filename()) |> +#' collect() +#' ``` +#' +#' This is useful, for example, when you have opened a subdirectory of a +#' partitioned dataset directly (so the partition columns above that directory +#' are not inferred) and want to recover the partition values from the path. +#' See [add_filename()] for details and limitations. +#' #' @param sources One of: #' * a string path or URI to a directory containing data files #' * a [FileSystem] that references a directory containing data files @@ -125,7 +142,7 @@ #' or call [`$NewScan()`][Scanner] to construct a query directly. #' @export #' @seealso \href{https://arrow.apache.org/docs/r/articles/dataset.html}{ -#' datasets article} +#' datasets article}, [add_filename()] #' @include arrow-object.R #' @examplesIf arrow_with_dataset() & arrow_with_parquet() #' # Set up directory for examples diff --git a/r/R/dplyr-funcs-augmented.R b/r/R/dplyr-funcs-augmented.R index 97b924b5e119..0d47d7d729c7 100644 --- a/r/R/dplyr-funcs-augmented.R +++ b/r/R/dplyr-funcs-augmented.R @@ -18,27 +18,42 @@ #' Add the data filename as a column #' #' This function only exists inside `arrow` `dplyr` queries, and it only is -#' valid when querying on a `FileSystemDataset`. +#' valid when querying on a `FileSystemDataset`, such as one created by +#' [open_dataset()]. Use it inside `mutate()` to add a column holding the path +#' of the file each row was read from. #' -#' To use filenames generated by this function in subsequent pipeline steps, you -#' must either call \code{\link[dplyr:compute]{compute()}} or -#' \code{\link[dplyr:collect]{collect()}} first. See Examples. +#' The filename column can be used in later `select()`, `arrange()` and +#' `group_by()` steps of the same query. However, it can't be used in +#' `filter()`, and some functions (such as `substr()`) are not supported on it. +#' In these cases, call \code{\link[dplyr:compute]{compute()}} or +#' \code{\link[dplyr:collect]{collect()}} first. [add_filename()] must also be +#' called before any aggregation or join. See Examples. #' #' @return A `FieldRef` \code{\link{Expression}} that refers to the filename #' augmented column. #' +#' @seealso [open_dataset()] +#' #' @examples \dontrun{ -#' open_dataset("nyc-taxi") |> mutate( -#' file = -#' add_filename() -#' ) +#' open_dataset("nyc-taxi") |> +#' mutate(file = add_filename()) |> +#' collect() +#' +#' # Simple expressions on the new column work in a later mutate(), for +#' # example to recover a partition value from the path +#' open_dataset("nyc-taxi/year=2015") |> +#' mutate(file = add_filename()) |> +#' mutate(year_from_path = sub(".*year=([0-9]{4}).*", "\\1", file)) |> +#' collect() #' -#' # To use a verb like mutate() with add_filename() we need to first call -#' # compute() +#' # To filter() on the new column, or use functions such as substr() that +#' # need to know its type, call compute() or collect() first #' open_dataset("nyc-taxi") |> #' mutate(file = add_filename()) |> #' compute() |> -#' mutate(filename_length = nchar(file)) +#' filter(endsWith(file, "part-0.parquet")) |> +#' mutate(file_start = substr(file, 1, 10)) |> +#' collect() #' } #' #' @keywords internal diff --git a/r/R/util.R b/r/R/util.R index dc48c5fb8ff8..233b0cb6869d 100644 --- a/r/R/util.R +++ b/r/R/util.R @@ -230,8 +230,8 @@ handle_augmented_field_misuse <- function(msg, call) { i = paste( "`add_filename()` or use of the `__filename` augmented field can only", "be used with Dataset objects, can only be added before doing", - "an aggregation or a join, and cannot be referenced in subsequent", - "pipeline steps until either compute() or collect() is called." + "an aggregation or a join, and cannot be used in filter() until", + "either compute() or collect() is called." ) ) abort(msg, call = call) diff --git a/r/_pkgdown.yml b/r/_pkgdown.yml index 5aa96dfff41b..a9bc60ae49e7 100644 --- a/r/_pkgdown.yml +++ b/r/_pkgdown.yml @@ -244,6 +244,8 @@ reference: Functionality for computing values on Arrow data objects. contents: - acero + - add_filename + - cast - arrow-functions - arrow-verbs - arrow-dplyr diff --git a/r/man/add_filename.Rd b/r/man/add_filename.Rd index 29457e25d382..22f659d8ab39 100644 --- a/r/man/add_filename.Rd +++ b/r/man/add_filename.Rd @@ -12,27 +12,43 @@ augmented column. } \description{ This function only exists inside \code{arrow} \code{dplyr} queries, and it only is -valid when querying on a \code{FileSystemDataset}. +valid when querying on a \code{FileSystemDataset}, such as one created by +\code{\link[=open_dataset]{open_dataset()}}. Use it inside \code{mutate()} to add a column holding the path +of the file each row was read from. } \details{ -To use filenames generated by this function in subsequent pipeline steps, you -must either call \code{\link[dplyr:compute]{compute()}} or -\code{\link[dplyr:collect]{collect()}} first. See Examples. +The filename column can be used in later \code{select()}, \code{arrange()} and +\code{group_by()} steps of the same query. However, it can't be used in +\code{filter()}, and some functions (such as \code{substr()}) are not supported on it. +In these cases, call \code{\link[dplyr:compute]{compute()}} or +\code{\link[dplyr:collect]{collect()}} first. \code{\link[=add_filename]{add_filename()}} must also be +called before any aggregation or join. See Examples. } \examples{ \dontrun{ -open_dataset("nyc-taxi") |> mutate( - file = - add_filename() -) +open_dataset("nyc-taxi") |> + mutate(file = add_filename()) |> + collect() -# To use a verb like mutate() with add_filename() we need to first call -# compute() +# Simple expressions on the new column work in a later mutate(), for +# example to recover a partition value from the path +open_dataset("nyc-taxi/year=2015") |> + mutate(file = add_filename()) |> + mutate(year_from_path = sub(".*year=([0-9]{4}).*", "\\\\1", file)) |> + collect() + +# To filter() on the new column, or use functions such as substr() that +# need to know its type, call compute() or collect() first open_dataset("nyc-taxi") |> mutate(file = add_filename()) |> compute() |> - mutate(filename_length = nchar(file)) + filter(endsWith(file, "part-0.parquet")) |> + mutate(file_start = substr(file, 1, 10)) |> + collect() } +} +\seealso{ +\code{\link[=open_dataset]{open_dataset()}} } \keyword{internal} diff --git a/r/man/open_dataset.Rd b/r/man/open_dataset.Rd index e2707212d035..41181c736eef 100644 --- a/r/man/open_dataset.Rd +++ b/r/man/open_dataset.Rd @@ -161,6 +161,24 @@ column types, as described above. If neither are provided, no partitioning information will be taken from the file paths. } +\section{Adding the source filename as a column}{ + + +Partitioning only recovers information encoded in directory names. If you +need to know which file each row came from, call \code{\link[=add_filename]{add_filename()}} inside a +\code{dplyr} query on the dataset: + +\if{html}{\out{