Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
A spreadsheet exported from Scopus contains papers, but it does not contain the search that found them. The exact query, retrieval date, paging choices and number of records returned by each part of the search often end up in notes or browser history. When the search has to be updated, it can be difficult to tell whether a changed result reflects the literature or the method.
I wrote scopusflow (Bernabeu, 2026) to keep those decisions with the records. A search begins as a plan, each completed part is cached and the final object can write its own search record. Matching packages are available for R (CRAN, GitHub) and Python (PyPI, GitHub).
The example below uses a small corpus bundled with the R package. A stand-in server, defined in a short script, answers the requests that would otherwise go to the Scopus API, so the code runs without an API key or any connection to Elsevier.
Development
The workflow below separates planning, retrieval, inspection and reporting. Only retrieval needs a key, so the other three stages run here against the corpus bundled with the package.
library(scopusflow)
packageVersion('scopusflow')
#> [1] '0.4.0'
Describe the Search First
The bundled example concerns research on graphene supercapacitors. scopus_query() adds the field syntax, and scopus_plan() divides the search into yearly cells. The plan can be inspected and saved before any quota is spent.
query <- scopus_query('graphene supercapacitor', .field = 'TITLE-ABS-KEY')
plan <- scopus_plan(query, years = 2015:2024, partition = 'year')
query
#> [1] "TITLE-ABS-KEY(graphene supercapacitor)"
plan
#> <scopus_plan> (10 cells, view "STANDARD", partition "year")
#> # A tibble: 10 × 6
#> cell query date year view page_size
#> * <int> <chr> <chr> <int> <chr> <int>
#> 1 1 TITLE-ABS-KEY(graphene supercapacitor) 2015 2015 STANDARD 200
#> 2 2 TITLE-ABS-KEY(graphene supercapacitor) 2016 2016 STANDARD 200
#> 3 3 TITLE-ABS-KEY(graphene supercapacitor) 2017 2017 STANDARD 200
#> 4 4 TITLE-ABS-KEY(graphene supercapacitor) 2018 2018 STANDARD 200
#> 5 5 TITLE-ABS-KEY(graphene supercapacitor) 2019 2019 STANDARD 200
#> 6 6 TITLE-ABS-KEY(graphene supercapacitor) 2020 2020 STANDARD 200
#> 7 7 TITLE-ABS-KEY(graphene supercapacitor) 2021 2021 STANDARD 200
#> 8 8 TITLE-ABS-KEY(graphene supercapacitor) 2022 2022 STANDARD 200
#> 9 9 TITLE-ABS-KEY(graphene supercapacitor) 2023 2023 STANDARD 200
#> 10 10 TITLE-ABS-KEY(graphene supercapacitor) 2024 2024 STANDARD 200
The plan already yields a draft methods description, written to PRISMA-S, the extension of the PRISMA statement for reporting literature searches (Rethlefsen et al., 2021). The draft also flags the items that the review authors will have to report themselves.
cat(format(scopus_search_report(plan), style = 'paragraph')) #> The search described here has not been run. The search expression is TITLE-ABS-KEY(graphene supercapacitor), limited to publication years 2015 to 2024. It would be partitioned into 10 cells, one per year, each retrieved through the STANDARD view in pages of 200 records, under a paging mode chosen when the search is run. The PRISMA-S items this record cannot supply, among them peer review of the strategy, grey literature and any other database searched, remain yours to report.
Resume From the Last Completed Cell
The bundled corpus now stands in for the result of that plan. The cache starts empty, so every yearly cell has to be requested.
source('stand_in.R')
stand_in <- scopus_stand_in(example_records)
options(scopusflow.api_key = 'offline-demo-key')
cache <- file.path(tempdir(), 'scopusflow-blog-cache')
unlink(cache, recursive = TRUE)
httr2::with_mocked_responses(stand_in$handler, {
records <- scopus_fetch_plan(plan, cache_dir = cache, resume = TRUE)
})
c(records = nrow(records), requests = stand_in$state$requests)
#> records requests
#> 138 10
Each year is written to its own checkpoint. Repeating the plan reloads those files and makes no further request.
requests_before <- stand_in$state$requests
httr2::with_mocked_responses(stand_in$handler, {
repeated <- scopus_fetch_plan(plan, cache_dir = cache, resume = TRUE)
})
c(records = nrow(repeated),
new_requests = stand_in$state$requests - requests_before)
#> records new_requests
#> 138 0
The practical benefit is clearest after an interruption, since the completed cells are kept. Each checkpoint also stores the query, year, view and page size it was fetched under, which prevents an incompatible plan from silently reusing it.
Inspect the Result
The records carry their plan, retrieval time and package version, along with the count for each cell. That provenance later helps to tell a change in the literature apart from a change in the method. Figure 1 counts the bundled records for each year.
httr2::with_mocked_responses(stand_in$handler, {
trend <- scopus_trend('graphene supercapacitor', years = 2015:2024,
field = 'TITLE-ABS-KEY')
})
plot_scopus_trend(trend) +
# Larger markers than the default, so that each yearly count stands out
ggplot2::geom_point(colour = '#31688E', size = 2.6) +
# plot_scopus_trend() captions its own output "Source: 'Scopus' Search API",
# which is not where these records come from: the demonstration corpus is
# offline and OpenAlex-derived, as the caption below the figure says. It is
# dropped rather than corrected, because a note belongs under the figure. The
# title goes with it, since the caption now sits directly above the plot and
# said the same thing twice.
ggplot2::labs(caption = NULL, title = NULL, y = NULL) +
ggplot2::theme(axis.text = ggplot2::element_text(size = 10.5))
Figure 1: Offline OpenAlex-Derived Demonstration Records, 2015 to 2024, From the Corpus Bundled With the Package. The counts show the workflow and say nothing about the Scopus literature on the topic.
Successive harvests can be compared by DOI. In this illustration, the baseline ends in 2023 and the later harvest adds the 2024 records.
baseline <- records[records$year <= 2023, ] changes <- scopus_diff_dois(old = baseline, new = records) changes #> <scopus_doi_diff> 14 added, 0 removed, 113 unchanged #> # A tibble: 127 × 2 #> doi status #> <chr> <fct> #> 1 10.1002/adfm.202315137 added #> 2 10.1002/asia.202400548 added #> 3 10.1002/slct.202302535 added #> 4 10.1016/j.cej.2024.148822 added #> 5 10.1016/j.diamond.2024.110842 added #> 6 10.1016/j.isci.2024.111696 added #> 7 10.1016/j.jallcom.2024.175000 added #> 8 10.1016/j.jallcom.2024.177248 added #> 9 10.1016/j.jpowsour.2024.234127 added #> 10 10.1016/j.jpowsour.2024.236149 added #> # ℹ 117 more rows
Let the Result Write the Record
The search record in Table 1 is generated from the object that holds the publications. It reports what the software knows and identifies the PRISMA-S items that still require the review authors’ judgement, such as citation searching or peer review of the search strategy. As this run never reached Scopus, the summary below leaves out the database, platform and retrieval date that a live report would include.
| Detail | Value |
|---|---|
| Search expression | TITLE-ABS-KEY(graphene supercapacitor) |
| Years | 2015 to 2024 |
| Records in bundled corpus | 138 |
| Records with a DOI | 127 |
| PRISMA-S checklist items | 16 |
The same records can then leave R as a DOI list, BibTeX, RIS or a table for bibliometrix. Saving the native object preserves the attached provenance for a future update.
Discussion
The result that matters in this example is the provenance. A rerunnable search records enough of the method to explain why a later DOI set differs, although database coverage, query design and screening remain matters of judgement. Those decisions belong in the review protocol and should be reported with the generated search record. For a review, I would keep the plan under version control, save the native result object and archive the search record alongside the screening data. The R reference and the Python documentation cover cursor paging, de-duplication, topic comparisons and authentication for live searches.
Limits
The offline corpus only demonstrates the workflow, so its counts are not evidence about the contents of Scopus or about graphene research. Elsevier’s terms do not allow Scopus records to be redistributed, which is why the example uses 138 records retrieved from OpenAlex on 22 July 2026 and reshaped to the package’s return format. A live search still depends on Elsevier’s coverage, indexing and quota, and abstract retrieval draws on a separate allowance.
References
Bernabeu, P. (2026). scopusflow: A reproducible workflow layer for Scopus bibliographic searches (Version 0.4.0) [Computer software]. CRAN. https://doi.org/10.32614/CRAN.package.scopusflow
Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., Ayala, A. P., Moher, D., Page, M. J., Koffel, J. B., & PRISMA-S Group. (2021). PRISMA-S: An extension to the PRISMA statement for reporting literature searches in systematic reviews. Systematic Reviews, 10(1), Article 39. https://doi.org/10.1186/s13643-020-01542-z
R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
