Site icon R-bloggers

pipeflow: v0.4.0 on CRAN

[This article was first published on R some blog, and kindly contributed to R-bloggers]. (You can report issue about the content on this page here)
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.

Pipeline setup

{pipeflow} builds a pipeline by adding one R function per step with pip_add (see also the Get started article). In this post, we’ll work with the following toy pipeline:

pip <- pip_new("my-pip") |>
    pip_add(
        step = "load",
        fun = \(n = 5) seq_len(n),
        tags = c("io", "daily")
    ) |>
    pip_add(
        step = "clean",
        fun = \(x = ~load) x * 2,
        tags = c("io", "core")
    ) |>
    pip_add(
        step = "fit",
        fun = \(x = ~clean) sum(x),
        tags = c("model", "core")
    ) |>
    pip_add(
        step = "report",
        fun = \(x = ~fit) paste("result:", x),
        tags = "report"
    )
< aside>

Each step must have a unique name, and function parameters can refer to other steps — for example, x = ~load takes the output of the load step. We also assign tags to each step, which we will use to filter steps later.

Before we run it, let’s have a look:

pip
<pipeflow> my-pip (4 steps)
---------------------------
     step params depends state       tags
1:   load      n           new   io,daily
2:  clean      x    load   new    io,core
3:    fit      x   clean   new model,core
4: report      x     fit   new     report
---------------------------
<ready> last run: never

We see one row per step, showing its parameters (params), the steps it depends on (depends), its current state (all new before the first run), and our tags. Once we run the pipeline, …

pip_run(pip)
info [2026-10-03 09:36:44.793 UTC]: Starting run of pipeflow 'my-pip'
info [2026-10-03 09:36:44.794 UTC]: Step 1/4 load
info [2026-10-03 09:36:44.795 UTC]: Step 2/4 clean
info [2026-10-03 09:36:44.797 UTC]: Step 3/4 fit
info [2026-10-03 09:36:44.799 UTC]: Step 4/4 report
info [2026-10-03 09:36:44.803 UTC]: Finished run of pipeflow 'my-pip'
pip
<pipeflow> my-pip (4 steps)
---------------------------
     step params depends state            out       tags
1:   load      n          done      1,2,3,4,5   io,daily
2:  clean      x    load  done  2, 4, 6, 8,10    io,core
3:    fit      x   clean  done             30 model,core
4: report      x     fit  done     result: 30     report
---------------------------
<ready> last run: 2026-10-03 11:36:44

… there is a new out column with the results, which can also be accessed directly:

pip[["report", "out"]]
[1] "result: 30"

Pipeline views in action

Real pipelines get long, and often you only care about a subset of steps: a topic, a stage, or the steps that produce the desired outputs. Views (new in v0.4.0) let you work on such a subset without copying anything. They reference the underlying pipeline, so every operation applied to it — running it, updating parameters, collecting output, … — writes through to the original pipeline, restricted to the steps the view covers.

Creating views

pip_view() returns a view containing only the steps that match the given filters:

pip_view(pip, tags = "core")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44
pip_view(pip, tags = "core", state = "done")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44
pip_view(pip, step = c("clean", "fit"))
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44

By default steps must match all filters (logical AND), while the values within one filter are alternatives (OR). join = "union" keeps steps that match any filter, and fixed = FALSE treats filter values as regular expressions:

pip_view(pip, tags = "report", step = "clean", join = "union")
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
     step params depends state            out    tags
1:  clean      x    load  done  2, 4, 6, 8,10 io,core
2: report      x     fit  done     result: 30  report
------------------------------------------
<ready> last run: 2026-10-03 11:36:44
pip_view(pip, step = "^f", fixed = FALSE)
<pipeflow_view> my-pip view (1 of 4 steps)
------------------------------------------
   step params depends state out       tags
1:  fit      x   clean  done  30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44

Selecting steps with [

The extract operator (also new in v0.4.0) provides a data.table-like way of selecting steps and returns a view by default:

pip[c("load", "fit")]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
   step params depends state       out       tags
1: load      n          done 1,2,3,4,5   io,daily
2:  fit      x   clean  done        30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44

Boolean filters are evaluated in the context of the step table.

pip[tags %like% "core"]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44
pip[step %in% c("clean", "fit") & state == "done"]
<pipeflow_view> my-pip view (2 of 4 steps)
------------------------------------------
    step params depends state            out       tags
1: clean      x    load  done  2, 4, 6, 8,10    io,core
2:   fit      x   clean  done             30 model,core
------------------------------------------
<ready> last run: 2026-10-03 11:36:44

Running views

Running a view executes the steps it covers together with any upstream dependencies that are not up to date. The run log marks the view’s own steps as [view] and the dependencies pulled in as [upstream]:

pip_reset(pip)  # reset the pipeline to "new" state for demonstration
pip_run(pip_view(pip, step = "report"))
info [2026-10-03 09:36:44.883 UTC]: Starting run of pipeflow 'my-pip view'
info [2026-10-03 09:36:44.883 UTC]: Step 1/4 [upstream] load
info [2026-10-03 09:36:44.883 UTC]: Step 2/4 [upstream] clean
info [2026-10-03 09:36:44.884 UTC]: Step 3/4 [upstream] fit
info [2026-10-03 09:36:44.885 UTC]: Step 4/4 [view] report
info [2026-10-03 09:36:44.886 UTC]: Finished run of pipeflow 'my-pip view'

Afterwards, the original pipeline is up to date for the covered steps. Let’s collect all results related to the “core” steps:

pip[tags %like% "core"] |> pip_collect()
$clean
[1]  2  4  6  8 10

$fit
[1] 30

For a complete guide on composing views, see the Pipeline views vignette.

More new and updated features

Views are only one of the additions. Each of the following has its own article on the documentation site:

Under the hood, significant performance gains have been made by implementing the dependency graph in C++ (since v0.3.0) and by improving the step and parameter bookkeeping in the underlying {data.table}. Finally, by dropping lgr and jsonlite, external package dependencies have been reduced to {data.table} and Rcpp.

{pipeflow} vs {targets}

{targets} remains R’s de-facto standard for heavy-duty, reproducible pipelines: it persists results to disk, records provenance, and scales to distributed infrastructure via crew. {pipeflow} does not try to compete on that home turf but targets a different niche — the interactive session, often as the backend of a Shiny app — where a pipeline is built and modified on the fly and has to respond while a user waits. Since responsiveness is key in that use case, {pipeflow} has been optimized for low latency.

The pipeflow vs targets article compares the in-session latency of the two packages head to head. The figure below gives a flavour of the results:

It shows the scenario where one parameter has changed1 and the pipeline is rerun, with each step doing 5 ms of work. The x-axis represents the number of steps in the pipeline (its size), and the y-axis shows the time required to re-run it.2

Across the tested sizes, {pipeflow} is consistently the faster of the two — see the article for the full benchmark and methodology.

Overall, the two are best seen as complementary: reach for {targets} for large batch jobs that must be reproducible end to end, and for {pipeflow} when the pipeline is part of an interactive application.

Wrapping up

{pipeflow} 0.4.0 is on CRAN:

install.packages("pipeflow")

The full changelog lists every change, and the documentation includes the articles linked above. If you build Shiny apps or spend a lot of time exploring parameters interactively, give it a try — feedback and issues are welcome on GitHub.


  1. Imagine a user changing the plot axis in a Shiny app.↩︎

  2. That means checking which steps are affected by the parameter change and then executing those steps.↩︎

To leave a comment for the author, please follow the link and comment on their blog: R some blog.

R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
Exit mobile version