Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
A verbal theory can sound precise while leaving its structure uncertain. Its constructs may overlap, a prediction may not follow from any proposition and an amendment may protect the theory without adding a new risk. These are properties of the theory’s specification, yet ordinary prose gives software nothing to inspect.
I wrote theoryforge (Bernabeu, 2026) to store that specification as a small YAML or JSON document. The document records constructs, propositions, predictions, alternatives, versions and provenance under a published schema. The package can then check the cross-references between those parts, derive implications from the causal graph and produce a dossier for review. Its R and Python implementations share that format.
The example is a theory of modality switching in grounded conceptual processing. People are slower to verify a conceptual property when the preceding trial involved a different perceptual modality than when it involved the same one (Pecher et al., 2003). The authors took this switching cost as evidence that conceptual processing draws on perceptual simulation, as the theory of perceptual symbol systems proposes (Barsalou, 1999).
Development
This section develops the theory in four stages: validate the document, read the causal graph, derive a discriminating test and preserve the resulting dossier. Every stage works on the same YAML file, and none of them asks the reader to take its structure on trust.
library(theoryforge)
packageVersion('theoryforge')
#> [1] '0.6.0'
switching <- tf_read(tf_example_path('modality-switching.theory.yaml'))
c(id = switching$id,
maturity = switching$maturity,
constructs = length(switching$constructs),
propositions = length(switching$propositions),
predictions = length(switching$predictions))
#> id maturity constructs
#> "modality-switching-2026" "developing" "5"
#> propositions predictions
#> "4" "4"
Treat the Theory as a Checked Document
Structural validation checks required fields and allowed values. Full validation also follows identifiers across the document, so a prediction cannot claim to derive from a proposition that does not exist.
tf_validate(switching, full = TRUE) check <- tf_check(switching) c(score = check$aggregate_score, gate = check$gate, blockers_failed = check$n_blockers_failed) #> score gate blockers_failed #> "85.1" "pass" "0"
The checklist has 12 items, each tied to a source in the methodological literature. They cover features that make a theory appraisable, such as whether it states what it prohibits, whether its predictions derive from stated propositions, and how precise and risky those predictions are (e.g., Meehl, 1990). The score measures how completely the theory is specified. A well-specified theory can still be wrong, while a promising informal account can score poorly because it is not yet explicit enough to test.
Read the Graph
The theory’s propositions form a causal graph. Because the source is structured data, the same graph can be rendered for a reader, as in Figure 1, or exported for another program.
flowchart TB c_sensorimotor_experience["Sensorimotor experience<br>with a concept"] c_modality_activation["Modality-specific<br>perceptual activation"] c_switch_cost["Cost of switching modality<br>between consecutive trials"] c_conceptual_access["Ease of conceptual access"] c_lexical_familiarity["Lexical familiarity with<br>the word form"] c_sensorimotor_experience -->|increases| c_modality_activation c_modality_activation -->|increases| c_switch_cost c_modality_activation -->|increases| c_conceptual_access c_lexical_familiarity -->|increases| c_conceptual_access
Figure 1: The Constructs of the Theory and the Arrows It Asserts Between Them, as theoryforge Reads Them From the File.
An acyclic causal graph also implies conditional independencies. Each one names an association that should be absent if the graph is correct, either outright or once certain constructs are held constant. Together, these claims define what the theory forbids in observed data.
The graph contains four arrows and entails six testable independencies. These six form a basis set, so every other independence the graph implies follows from them. The most discriminating one is that sensorimotor experience and lexical familiarity should be uncorrelated, because the only path between them is blocked by a collider at conceptual access. An observed correlation between the two, with nothing conditioned on, would count against the graph.
Give the Claims a Discriminating Test
To test these claims, I simulated two data sets of 4,000 observations each. One follows the theory’s graph, and the other follows a rival account in which sensorimotor experience also contributes to lexical familiarity, with every other path unchanged. In each data set, I assessed all six independencies with the linear conditional-independence tests in the dagitty package (Textor et al., 2016). Table 1 shows how the key claim fares in both.
| World | Correlation | p |
|---|---|---|
| Data generated from the theory | −0.009 | 0.56 |
| Data generated from the rival | 0.445 | <1e−99 |
All six implied independencies are compatible with the data generated from the theory. In the data generated from the rival, only the claim shown in the table fails. There, sensorimotor experience and lexical familiarity are substantially correlated, while the estimates for the other five claims stay close to zero. The failure therefore locates the part of the graph that conflicts with the data, a diagnosis that no single global score could give.
Preserve the Result With the Theory
The dossier gathers the rigour checklist, the severity of each prediction, the theory’s provenance and a preregistration of its hypotheses into one Markdown record. The R and Python packages produce the same dossier for the same theory, so a checksum of that record identifies the audited state of the theory in either language. Table 2 repeats the score and gate and shows the first characters of the SHA-256 checksum.
| Detail | Value |
|---|---|
| Rigour score | 85.1/100 |
| Gate | pass |
| SHA-256 prefix | 4c07a9235e37 |
Discussion
Formalisation changes what can be criticised. A reader can go beyond asking whether a theory sounds plausible and inspect which constructs are linked, which observations would count against the graph and how an amendment changes those commitments.
The package documentation covers structural equation model compilation, severity scoring, preregistration, amendment appraisal, version differences and deposits to the Open Science Framework. Those tools build on the step shown here, which is to write the theory in a form that states what it claims and what it forbids.
Limits
Formalisation moves judgement into the open without removing it. Researchers still decide which constructs and arrows belong in the graph, how the constructs are measured, and whether an observed departure matters scientifically. The theory document records where each of those judgements enters. The checklist screens the structure and wording of a specification, so it cannot tell whether the constructs themselves are well chosen. Theories with feedback loops also need methods beyond the conditional-independence analysis used here, because theoryforge derives implications only from acyclic graphs.
References
Barsalou, L. W. (1999). Perceptual symbol systems. Behavioral and Brain Sciences, 22(4), 577–660. https://doi.org/10.1017/S0140525X99002149
Bernabeu, P. (2026). theoryforge: Systematic theory development (Version 0.6.0) [Computer software]. CRAN. https://doi.org/10.32614/CRAN.package.theoryforge
Meehl, P. E. (1990). Appraising and amending theories: The strategy of Lakatosian defense and two principles that warrant it. Psychological Inquiry, 1(2), 108–141. https://doi.org/10.1207/s15327965pli0102_1
Pecher, D., Zeelenberg, R., & Barsalou, L. W. (2003). Verifying different-modality properties for concepts produces switching costs. Psychological Science, 14(2), 119–124. https://doi.org/10.1111/1467-9280.t01-1-01429
Textor, J., van der Zander, B., Gilthorpe, M. S., Liśkiewicz, M., & Ellison, G. T. H. (2016). Robust causal inference using directed acyclic graphs: The R package ‘dagitty’. International Journal of Epidemiology, 45(6), 1887–1894. https://doi.org/10.1093/ije/dyw341
R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
