Institutions Are a Package Deal: What a Correlation Can’t Tell You About the Rule of Law
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.
TLDR: This is the third post in a series examining relationships between the rule of law and other institutions as measured by the World Bank’s Worldwide Governance Indicators (WGI). Earlier posts demonstrated that the World Justice Project’s (WJP) measure can be proxied by the WGI Rule of Law (RoL) index, to take advantage of its wider coverage over a longer time-period.
This post builds on that analysis by asking how institutions should be conceptualized and analyzed: can the rule of law (and other institutions) sensibly be examined one at a time, or do they need to be treated as parts of an interdependent system? Using the six dimensions of the WGI as a proxy for institutional strength, results suggest dimensions are indeed bound together. Countries that score well on the rule of law also frequently score well on government effectiveness and corruption controls, year-to-year movements are positively linked across most dimensions and a common measure (or cause) accounts for the bulk of the variation between countries.
None of this will surprise anyone familiar with the literature and the analysis isn’t intended to prove institutions matter or that the WGI measure them. The point is narrower: to show why the rule of law can’t be examined in isolation and the risks of coming to the wrong conclusions when we try. It also sets up later posts in the series which will look at how the rule of law is connected with other national characteristics, such as economic growth.
Background
Institutions describe the constraints that structure political, economic and social interaction, such as how power is gained and exercised, how order is kept, and the way public resources are sourced, divided and used. Institutions shape incentives and opportunities: good institutions lead to more of the things we care about, like economic growth, peace and prosperity; while bad institutions lead to more of the thing we want less of, like poverty, inequality and conflict.
A sizable share of my work sits in economic development, so institutions are never far from mind. This series, though, started with a specific job: a client hired me to look at the links between the rule of law and economic growth, and I went looking for an accessible introduction to the topic I could share. What I found was this analysis from the Atlantic Council on why the rule of law is the key to prosperity.
Having noted how important institutions are and how difficult they are to define and measure, the authors move briskly to their conclusion that the rule of law is the single most influential factor behind long-term economic growth and societal wellbeing:

Source: Annie (Yu-Lin) Lee, Joseph Lemoine, 20 August 2025, Why the rule of law is the key to prosperity: Lessons from thirty years of data, Atlantic Council (link) [Accessed 5 August 2026]
That’s quite the claim and one of the inspirations for this post as from the looks of it, much of their evidence seems to come from the strength of a collection of pairwise correlations between the rule of law and their chosen proxies for well being (as measured by the Atlantic Council’s Freedom and Prosperity Index).
Now, I’m no big city lawyer, but it seems to me that pairwise correlations are a poor piece of evidence to provide for such a strong claim. Particularly when we’re wading into the murky waters of prosperity, institutions and the causal connection between the two. And the characteristics being analyzed are plausibly all part of the same interdependent system.
But, I’m not here to judge. Firstly, as the authors do mention that prosperity depends on the interplay of multiple institutional pillars, not just the rule of law. They’ve also tackled a complicated topic in an accessible way, which is an achievement in itself. And I’ve been there before too: as cross-country correlations are fun to explore and can provide a seemingly endless array of plausible policy interventions for making the world a better place. Also, the point of this post isn’t to criticize their analysis, but to help fill the gap I noticed when searching for resources on the topic.
Instead, this post attempts to fill a gap by presenting analysis demonstrating why the rule of law, and institutions more generally, are best conceptualized as interdependent pillars within a mutually reinforcing system. That interdependence is what makes the literature so conceptually interesting and so frustrating to analyze, since the little data available on the topic arrives bundled with multicollinearity, endogeneity, collider bias, omitted variables and measurement error.
The rule of law and institutions: a primer
To obnoxiously paraphrase research1, the rule of law is thought to influence economic development and wider prosperity through a variety of avenues, such as enabling the enforcement of property rights, easing trade between unrelated parties and providing a peaceful means for resolving disputes. However, because the rule of law and its outcomes also depend on a wider set of institutions and a tangled coalition different interests, analyzing it is no simple task.2
On one level this is because institutions and their outcomes are interdependent. For instance, maybe the rule of law does drive prosperity, but perhaps prosperous countries are better able to invest in their legal systems too. Similarly, perhaps how the rule of law drives economic outcomes changes depending on the political system, geography and urbanization.
Defining and measuring amorphous concepts like the rule of law also happens to be hard, which means researchers lean heavily on surveys asking respondents for their perception of corruption, law and order and government effectiveness (etc). However, because a person’s opinion of a country’s performance in one area is likely to be influenced by their impression of it in others, these measures are likely to agree with one another for reasons that have little to do with the institutions themselves. A respondent who has watched a corruption scandal unfold is unlikely to rate the courts generously that same year.
Finally, it’s generally accepted that institutions are slow moving (or ‘sticky’), which means they don’t change much from one year to the next and will often exert their influence on wider outcomes indirectly. As a result, some of the most influential studies exploring the connection between institutions and prosperity analyse periods of a hundred years or more, such as Acemoglu, Johnson and Robinson who compared current income with a century old proxy for institutional quality (settler mortality).
Figure: Income vs. Settler mortality
Countries with higher GDP per capita now tended to be those with stronger colonial institutions (as proxied by settler mortality)

Source: Acemoglu, Daron, Simon Johnson, and James A. Robinson. 2001. “The Colonial Origins of Comparative Development: An Empirical Investigation.” American Economic Review 91 (5): 1369–1401. DOI: 10.1257/aer.91.5.1369
The Worldwide Governance Indicators
Although there are good reasons to question whether the WGI provides a reliable and holistic measure of institutions, (or even that it measure what it claims to) it is arguably one of the more reliable and well-tested attempts at measuring the cross-country quality of institutions. I also have a personal preference for the WGI as somebody who is hired to design, build and evaluate composite indicators, because it gets a core part of index design right: transparently sharing their methodology and data, and being open to making revisions to reflect feedback.
The WGI also has the practical advantage of being intuitive enough for outsiders to understand what it’s trying to measure. With governance reflecting the traditions and institutions by which authority in a country is exercised across six dimensions:
- Voice and Accountability: perceptions of the extent to which citizens can participate in selecting their government including electoral integrity, and of accountability mechanisms for citizens—reflected in the ability to access information, governmental oversight bodies, and a robust traditional/digital media landscape; and
- Political Stability: perceptions of the extent to which political power and governance are secure from destabilization, and of the likelihood that authority will be challenged or altered through violent, coercive, or unconstitutional means. The second aspect—government capacity to formulate and implement sound policies—includes:
- Government Effectiveness: perceptions of the quality of public services, the civil service, policy formulation and implementation, and the credibility of a government’s decisions; and
- Regulatory Quality: perceptions of the government’s ability to design and implement policies and regulations that promote private sector development. The third aspect—respect for institutions that govern economic and social interactions— comprises:
- Rule of Law: perceptions of the extent to which agents respect and follow the rules of society, including contract enforcement, property rights, the police, courts, and the likelihood of crime and violence; and
- Control of Corruption: perceptions of the extent to which public power is used for private gain, including both petty and grand corruption, as well as capture of the state by elites and private interests.
Source: World Bank, 2025, “The Worldwide Governance Indicators: Revised Methodology for Measuring Governance Using Perception Data December 2025”, link.
Note: For brevity, this post uses institutions, governance and the WGI interchangeably, in full knowledge that they aren’t the same thing. The WGI is the World Bank’s conceptualization of a particular set of institutions it considers useful, interesting and/or relevant. The WGI is not meant to be an all-encompassing measure suited to every use case, it’s just the most suitable measure for this analysis. The scores are also proxies for institutional quality rather than measures of it. Proxies come with the territory when trying to measure hard-to-define things like institutions and/or governance, as direct measurement is unavailable or unsuitable, particularly for cross-country comparisons.
Figure: Worldwide Governance Indicator Dimensions
The WGI attempts to measure governance across six dimensions

Source: my exceptional design skills
FYI: Daniel Kaufmann has also publicly responded to criticism of the WGI, which has resulted in several iterations of the methodology over time.
Project Setup and Data Cleaning
Once again, the WGI data can be downloaded here. Because the code analyzes all WGI indicators, fnc_read_wgi_sheets() load data from each sheet and combine it into a single dataframe.
Code: Plotting functions
#load the packages we'll probably need
library(tidyverse)
library(readxl)
library(janitor)
library(qgraph)
library(rstatix)
library(broom)
#Note: ppcor is also required, but called directly to avoid conflicts
#Install with install.packages("ppcor") if needed.
#define path
ref_wgi_path <- "./Data/wgidataset_with_sourcedata-2025.xlsx"
# sheet code -> output column name
ref_wgi_indices <- c(
va = "wgi_voice", # Voice and Accountability
pv = "wgi_polstab", # Political Stability & Absence of Violence
ge = "wgi_goveff", # Government Effectiveness
rq = "wgi_regqual", # Regulatory Quality
rl = "wgi_rol", # Rule of Law
cc = "wgi_corrupt" # Control of Corruption
)
# Read one WGI sheet and return iso3c, year, and the single renamed estimate.
# Renaming per sheet avoids the shared 'governance_estimate...' column colliding
# on join. Fails loudly if a sheet's estimate column is named differently.
fnc_read_wgi_sheet <- function(sheet_code, col_name, path) {
read_excel(path, sheet = sheet_code) |>
clean_names() |>
rename(iso3c = economy_code,
!!col_name := governance_estimate_approx_2_5_to_2_5) |>
select(iso3c, year, all_of(col_name))
}
# read all six, then join on the common keys
dta_wgi_2025 <- ref_wgi_indices |>
imap(\(col_name, sheet_code) fnc_read_wgi_sheet(sheet_code, col_name, ref_wgi_path)) |>
reduce(full_join, by = c("iso3c", "year"))
#Add metadata (country + year-specific income bracket) taken once, from any sheet
tmp_dta_wgi_meta <- read_excel(ref_wgi_path, sheet = "rl") |>
clean_names() |>
rename(iso3c = economy_code, country = economy_name) |>
select(iso3c, year, country, income_classification)
#add meta data
dta_wgi_2025 <- dta_wgi_2025 |>
left_join(tmp_dta_wgi_meta, by = c("iso3c", "year")) |>
relocate(iso3c, country, year, income_classification)
#drop temporary objects (ref_ objects are kept so the chunk can be re-run)
rm(tmp_dta_wgi_meta)
#define index column names and labels:
ref_wgi_cols <- c("wgi_voice", "wgi_polstab", "wgi_goveff",
"wgi_regqual", "wgi_rol", "wgi_corrupt")
ref_wgi_labels <- c(wgi_voice = "Voice & accountability",
wgi_polstab = "Political stability",
wgi_goveff = "Govt effectiveness",
wgi_regqual = "Regulatory quality",
wgi_rol = "Rule of law",
wgi_corrupt = "Control of corruption")
#shortened labels
ref_wgi_short <- c(wgi_voice = "Voice & Acc.",
wgi_polstab = "Pol. stability",
wgi_goveff = "Gov. effect.",
wgi_regqual = "Reg. quality",
wgi_rol = "Rule of law",
wgi_corrupt = "Corruption")
# assumptions
ref_min_years <- 20 # minimum years of data before a country average is used
ref_pcor_cut <- 0.05 # partial correlations below this aren't drawn in the network
The rule of law as part of an institutional portfolio
To demonstrate what a statistical minefield analyzing institutions is, we’ll start by applying the same technique that I criticized earlier: pairwise correlations. With the idea being to demonstrate that different measures of institutional strength tend to agree with one another. In addition to this being what you might expect if institutions were endogenously determined, interdependent and/or WGI dimensions measured something similar. It also illustrates why you can’t rely on associations alone to cleanly establish the causal connection between a particular institution and economic and social outcomes.
For comparing the shared strength of institutions we’ll use a country’s average score for each WGI dimension from 1996 to 2024.3 Although this reduces our sample size and ignores a variety of country-specific factors that might influence institutions, it provides a simple way to reduce year-to-year noise and focus on between-country effects. It’s also unlikely we’ll lose too much information, given institutions move slowly from year-to-year. But, averaging cuts both ways too: as stripping out year-to-year noise also strips out measurement error, making correlations mechanically higher within any single year.
Reflecting that WGI dimensions might just be unrelated series that are trending together, we’ll also examine pairwise correlations for annual changes in WGI scores across dimensions. Aside from checking whether the level associations are just shared trends, dimensions moving together year-to-year would be consistent with institutions being connected to one another. For instance, if the rule of law supports efforts to fight corruption, we might expect the quality of these two institutions to move in the same direction.
However, the use of averages for comparing levels and raw values for first-differences also warrants caution when making comparisons: As while the two describe the same countries, the different units of analysis and number of observations aren’t directly comparable.
Code: calculate country averages and first differences
#Country averages (levels).
#Countries are only kept where every dimension has at least ref_min_years of
#data
dta_wgi_country <- dta_wgi_2025 |>
group_by(iso3c) |>
filter(if_all(all_of(ref_wgi_cols), \(x) sum(!is.na(x)) >= ref_min_years)) |>
summarise(across(all_of(ref_wgi_cols), \(x) mean(x, na.rm = TRUE)),
nmb_years = n(),
.groups = "drop")
#how many countries survive the coverage filter
chk_wgi_country_nmb <- nrow(dta_wgi_country)
chk_wgi_country_nmb
#first difference column names
ref_wgi_diff_cols <- paste0("d_", ref_wgi_cols)
#work out which rows sit next to a consecutive year
#(the WGI was published every other year before 2002, so we can't just lag blindly)
dta_wgi_gaps <- dta_wgi_2025 |>
arrange(iso3c, year) |>
group_by(iso3c) |>
mutate(yr_gap = year - lag(year)) |>
ungroup()
#then take the year-on-year change, but only where the gap is a single year
dta_wgi_diff <- dta_wgi_gaps |>
group_by(iso3c) |>
mutate(across(all_of(ref_wgi_cols),
\(x) if_else(yr_gap == 1, x - lag(x), NA_real_),
.names = "d_{.col}")) |>
ungroup() |>
select(iso3c, year, all_of(c(ref_wgi_cols, ref_wgi_diff_cols)))
#Complete-case frames used for BOTH the zero-order and partial correlations.
dta_wgi_lvl_cc <- dta_wgi_country |>
select(all_of(ref_wgi_cols)) |>
drop_na()
dta_wgi_dif_cc <- dta_wgi_diff |>
select(iso3c, all_of(ref_wgi_diff_cols)) |>
drop_na()
#sample sizes
chk_sample_sizes <- tibble(
series = c("WGI Levels", "WGI First Differences"),
nmb_obs = c(nrow(dta_wgi_lvl_cc), nrow(dta_wgi_dif_cc)),
nmb_country = c(nrow(dta_wgi_lvl_cc), n_distinct(dta_wgi_dif_cc$iso3c)),
unit = c("country", "country-year")
)
Note: The WGI methodology has been updated in 2025 to improve comparability of governance scores over time.
Code: Correlation analysis
{r}
#create a function for calculating correlations
fnc_cor_p <- function(dta, cols) {
dta |>
select(all_of(cols)) |>
cor_test(vars = cols) |>
rename(x = var1, y = var2) |>
filter(match(x, cols) < match(y, cols)) |>
mutate(across(c(x, y), \(v) sub("^d_", "", v))) |>
select(x, y, cor, p)
}
#calculate association(s) on the shared complete-case frames
sum_wgi_cor_level <- fnc_cor_p(dta_wgi_lvl_cc, ref_wgi_cols)
sum_wgi_cor_diff <- fnc_cor_p(dta_wgi_dif_cc, ref_wgi_diff_cols)
#Holm adjustment for testing 15 pairs at once. Note the p-values for the
#differences assume ~5,000 independent observations when they are really ~200
#countries observed repeatedly - so significance there is close to guaranteed
#and the effect sizes are what matter.
sum_wgi_cor_level <- sum_wgi_cor_level |>
mutate(sig = p.adjust(p, "holm") < 0.05,
r2 = cor^2)
sum_wgi_cor_diff <- sum_wgi_cor_diff |>
mutate(sig = p.adjust(p, "holm") < 0.05,
r2 = cor^2)
This code chunk just creates a basic bubble plot for visualizing the correlation matrix. This is perhaps more verbose than it needs to be, but it felt fair to give Claude the satisfaction of trying to apply my style guidelines.
Code: Plotting functions
#Bubble plot for a correlation matrix. value_col lets the same function handle
#zero-order (cor) and partial (pcor) tables.
fnc_plot_bubbles <- function(dta, value_col = "cor", wrap = 12) {
ref_lbl <- setNames(str_wrap(ref_wgi_labels, wrap), ref_wgi_cols)
dta <- dta |>
rename(value = all_of(value_col)) |>
mutate(x = factor(x, ref_wgi_cols),
y = factor(y, ref_wgi_cols),
lbl = ifelse(sig, sprintf("%.2f", value), sprintf("(%.2f)", value)))
ggplot(dta, aes(x, y)) +
geom_point(aes(size = abs(value), fill = value), shape = 21, colour = "white") +
geom_text(aes(label = lbl, colour = abs(value) > 0.55), size = 3, fontface = "bold") +
scale_fill_gradient2(low = "#922C40", mid = "#FFFFFF", high = "#16AF8E",
midpoint = 0, limits = c(-1, 1), name = "Correlation") +
scale_colour_manual(values = c(`TRUE` = "white", `FALSE` = "#121212"), guide = "none") +
scale_size_area(max_size = 16, limits = c(0, 1), guide = "none") +
scale_x_discrete(labels = ref_lbl) +
scale_y_discrete(labels = ref_lbl, limits = rev) +
coord_fixed() +
labs(x = NULL, y = NULL) +
theme_minimal()
}
# Convert the long pairwise tibble back into a symmetric matrix for qgraph
fnc_pcor_matrix <- function(dta_pcor, col) {
tmp_mat <- matrix(0, length(ref_wgi_cols), length(ref_wgi_cols),
dimnames = list(ref_wgi_cols, ref_wgi_cols))
tmp_mat[cbind(dta_pcor$x, dta_pcor$y)] <- dta_pcor[[col]]
tmp_mat[cbind(dta_pcor$y, dta_pcor$x)] <- dta_pcor[[col]]
tmp_mat
}
#Network of direct links. Edges below `cut` in absolute size are NOT DRAWN, so a
#missing line means "smaller than the threshold", not "zero".
fnc_plot_network <- function(dta_pcor, cut = ref_pcor_cut, layout = "spring", seed = 123) {
set.seed(seed)
mat_pcor <- fnc_pcor_matrix(dta_pcor, "pcor")
chk_sig <- fnc_pcor_matrix(dta_pcor, "sig") == 1
qgraph(mat_pcor,
layout = layout, minimum = cut, maximum = 1, esize = 9, fade = TRUE,
posCol = "#16AF8E", negCol = "#922C40",
color = "white", border.color = "#E5E7EB", border.width = 2,
labels = ref_wgi_short[ref_wgi_cols], label.cex = 0.8,
label.color = "#121212", vsize = 11, shape = "circle",
lty = ifelse(chk_sig, 1, 2),
edge.labels = ifelse(chk_sig, sprintf("%.2f", mat_pcor),
sprintf("(%.2f)", mat_pcor)),
edge.label.cex = 0.75, edge.label.bg = "white",
edge.label.color = "#121212", mar = rep(5, 4))
}
Strong institutions coincide with one another
In this section, pairwise correlations are used to test whether the strength of individual institutions tends to occur together.
It’ll come as no surprise that all of the WGI’s dimensions are strongly associated with each other, which suggests that on average a country scoring highly on one dimension probably scores highly on the others too. For the WGI’s Rule of Law measure, the pairwise associations are strong across the board, with the implied R squared statistic suggesting it accounts for somewhere between 65 and 90 percent of the variation in the other dimensions. Political Stability has the weakest associations, accounting for somewhere between 50 and 75 percent.
Figure: WGI pairwise correlations (levels)
The strength of institutions are strongly associated with one another across WGI dimensions

The network plot presents the same pairwise associations once the influence of the other dimensions is accounted for. Where partial correlations fall well below the zero-order correlations shown above, it points to dimensions carrying overlapping information. For the Rule of Law, the network diagram presents much weaker associations with other dimensions to the figure above, with moderate associations remaining for Voice and Accountability, Control of Corruption and Political Stability. Although it’s best not to take the implications of this analysis too far, one interpretation of these results that makes intuitive sense is that different institutions relate to one another differently. And sometimes this relationship might be indirect, such as government effectiveness indirectly influencing the rule of law via corruption controls.
Figure: WGI pairwise partial correlations (levels)
Pairwise correlations are lower once the influence of other WGI dimensions are accounted for

Code: Plotting pairwise correlations at levels
#produce and display the correlation matrix plt_wgi_cor_level <- fnc_plot_bubbles(sum_wgi_cor_level) plt_wgi_cor_level #network plot for partial correlations plt_wgi_net_level <- fnc_plot_network(sum_wgi_pcor_level)
Institutions are connected to one another over time (albeit, loosely)
This section explores pairwise correlations between year-to-year movements of WGI dimension to explore evidence for institutions being interdependent.
Focusing on year-to-year movements in WGI dimensions rather than levels results in smaller correlation statistics across the board, which is to be expected given taking the first difference naturally inflates the influence of noise and measurement errors. Still, the associations point to a similar picture to associations at the levels: a positive and statistically significant association between all dimensions, with Political Stability being the weakest.
Having said this, statistical significance just suggests that shared movement between series isn’t zero, not that it’s particularly interesting. Added to this, given the first differences pool values so they are considered independent (despite being from the same countries), significance is close to guaranteed, which is yet another reason they should be interpreted with caution.
The implied explanatory power is also probably what matters more and it’s generally modest. For instance, the R squared statistic for Political stability suggests it can explain between 1 to 4 percent of the year-to-year variation in other dimensions. Whereas the Rule of Law, which holds the strongest pairwise explanatory power across dimensions (setting political stability aside), only explains somewhere between 9 to 17 percent of the variation, which while being nothing to sneeze at, still leaves a lot of unexplained movement.
Figure: WGI pairwise correlations (year-to-year changes)
Year-to-year movements in institutional quality are statistically associated with one another

Note: Associations between first differences are expected to be weaker for several reasons: institutions are slow moving; interdependence between institutions may operate with lags or in a non-linear fashion; and as differencing strips out persistent components of each series while retaining measurement errors intact, the signal to noise ratio is likely to be lower. The authors of the WGI also warn against taking too much stock of year-to-year changes in scores.
The second plot once again presents how year-to-year movements are associated with each other, but after accounting for influences outside the examined pair. Once again, the Rule of Law holds the strongest association with other dimensions, but its association with Control of Corruption is much lower, suggesting that much of the pairwise association between year-to-year changes above relate to other dimension outside the pair.
Figure: WGI pairwise partial correlations (year-to-year changes)
Pairwise correlations are reduced, but in most cases remain statistically significant, once the influence of other dimensions is accounted for

Code: plot first difference associations
#correlaton plot for correlations @ first difference
plt_wgi_cor_diff <- fnc_plot_bubbles(sum_wgi_cor_diff)
plt_wgi_cor_diff
#network plot for partial correlatons @ first difference
plt_wgi_net_diff <- fnc_plot_network(sum_wgi_pcor_diff,
layout = plt_wgi_net_level$layout)
Six dimensions, one signal(?)
This section uses Principal Components Analysis (PCA) to explore whether all six dimensions of the WGI measure, or are caused by, the same thing.
A strong critique of the WGI is that dimensions more or less measure the same thing. From a conceptual standpoint this might have some weight, due to the perceived quality of one dimension influencing perceptions of another, the use of overlapping data sources and the fact that the definitions are not precisely defined. But, as somebody that works a lot with composite indices, I’d say this comes with the territory and I’m happy to leave it to the experts to argue among themselves.
However, from a statistical standpoint this might matter a lot, as it determines whether a measured link between the Rule of Law and outcomes like prosperity can be meaningfully interpreted as telling us anything about the Rule of Law.
This is explored in the code below by applying PCA to the WGI levels and first differences. PCA attempts to collapse collinear variables into a set of principal components (PC) that explain as much variability as possible. If the WGI’s dimensions are measuring or being driven by something common, we might expect a small number of PCs will explain a disproportional share of the variance.
Code: apply PCA to levels and first differences
#Fit a scaled PCA and pull out variance shares and PC1 loadings.
#Scaling so each index counts equally regardless of how spread out it is.
#Note the levels PCA runs on country averages (one row per country) while the
#differences PCA runs on pooled country-years - the shares are not measured on
#the same unit of analysis.
fnc_pca_tidy <- function(dta, cols, series) {
dta_input <- dta |>
select(all_of(cols)) |>
drop_na()
mod <- prcomp(dta_input, scale. = TRUE)
#share of variance picked up by each component
sum_var <- mod |>
tidy(matrix = "eigenvalues") |>
transmute(series, pc = paste0("PC", PC), var_pct = percent,
nmb_obs = nrow(dta_input))
#how strongly each index marks the first component
sum_load <- mod |>
tidy(matrix = "rotation") |>
filter(PC == 1) |>
transmute(series, index = sub("^d_", "", column), loading = abs(value))
list(var = sum_var, load = sum_load)
}
ref_pca_series <- c("WGI Levels", "WGI First Differences")
tmp_pca <- list(fnc_pca_tidy(dta_wgi_lvl_cc, ref_wgi_cols, ref_pca_series[1]),
fnc_pca_tidy(dta_wgi_dif_cc, ref_wgi_diff_cols, ref_pca_series[2]))
sum_wgi_pca_var <- tmp_pca |>
map("var") |>
list_rbind() |>
mutate(series = factor(series, ref_pca_series))
sum_wgi_pca_load <- tmp_pca |>
map("load") |>
list_rbind() |>
mutate(series = factor(series, ref_pca_series))
#drop temporary objects
rm(tmp_pca)
The first plot presents the share of variance explained by each principal component. In the case of the the WGI’s levels, the PCA indicates that almost 90 percent of the measured variance in governance can be explained by a single principal component. Indicating that the majority of information presented by the six WGI dimensions could be efficiently described by a single measure. A result that supports the idea that either the WGI is measuring a similar thing and/or that a common factor is driving all six dimensions.
Both PCAs point in a similar direction, although the first differences are much less dramatic: with the first component picks up around 40 percent of the variance, against nearly 90 percent for the levels. Bear in mind these two numbers aren’t measured on the same unit of analysis, making them not directly comparable (i.e. approximately 200 country averages vs 5,000+ first differences). It’s therefore best not to read too much into comparisons, particularly given the authors of the WGI explicitly warn against analyzing year-to-year score movements.
Figure: Variance explained by principal component
The majority of variance can be explained by a single PC for WGI levels, while a larger number of PCs are required to provide a sufficient explanation of variability for first differences

Code: how concentrated is the common dimension?
{r}
plt_wgi_pca_scree <- sum_wgi_pca_var |>
ggplot(aes(pc, var_pct)) +
geom_col(fill = "#1E298D", width = 0.7) +
geom_text(aes(label = scales::percent(var_pct, accuracy = 1)),
vjust = -0.4, size = 3, colour = "grey40") +
scale_y_continuous(labels = scales::percent, limits = c(0, 1)) +
facet_wrap(~ series) +
labs(x = NULL, y = "Variance explained (%)") +
theme_minimal(base_size = 10) +
theme(panel.grid.major.x = element_blank(),
panel.grid.minor = element_blank(),
strip.text = element_text(face = "bold", hjust = 0))
plt_wgi_pca_scree
The plot below presents PCA loadings, which measure how strongly each dimension contributes to PC1. Higher loadings mean a dimension is more closely tied to the component, and so shares more with the others. That the dimensions carry comparable loadings at both the levels and the first differences is again consistent with them measuring, or being influenced by, something common.
Political Stability sits lower than the rest, but I’d be reluctant to take my interpretation of that too far given how much variance remains unexplained. It’s also roughly what you might expect of a dimension intended to capture shocks rather than gradual shifts, or one behaving non-linearly (questions better answered in a separate post).
Figure: PCA Loadings
PCA loadings for both the levels and first differences are generally evenly spread

Code: PCA loadings
plt_wgi_pca_load <- sum_wgi_pca_load |>
mutate(index = factor(index, ref_wgi_cols, ref_wgi_labels) |> fct_reorder(loading)) |>
ggplot(aes(loading, index)) +
geom_col(fill = "#1E298D", width = 0.65) +
geom_text(aes(label = sprintf("%.2f", loading)),
hjust = -0.25, size = 3, colour = "grey40") +
scale_x_continuous(limits = c(0, 0.7), expand = expansion(c(0, 0.05))) +
facet_wrap(~ series) +
labs(x = "Absolute loading on PC1", y = NULL) +
theme_minimal(base_size = 10) +
theme(panel.grid.major.y = element_blank(),
panel.grid.minor = element_blank(),
strip.text = element_text(face = "bold", hjust = 0))
plt_wgi_pca_load
Summing up
Across three views of the same data, the six WGI dimensions behave like parts of one system rather than six separable things. Country averages correlate strongly across every pair; most associations remain even when accounting for the influence of other dimensions; and a single principal component accounts for a significant share of variation in governance scores between countries. Dimensions seem to move together too, albeit weakly, which is while not a slam dunk, does point to WGI scores having some interdependence.
But, it’s worth being clear about what this analysis doesn’t establish: Collinearity isn’t endogeneity and while a system of interdependent institutions might produce this pattern, so would WGI dimensions measuring the same thing.
But, the point of this post isn’t to support the legitimacy of the WGI or any other measure. And the distinction matters little to the point of this post. As whether dimensions move together because they co-determine each other, share a common driver, or share source data, the implications for anyone trying to understand or analyze institutions is the same: they have to be examined as a set as a correlation between one institution and an outcome may strong regardless of whether that pillar is doing the actual work or not.
How AI was used for this post: Claude was used heavily to refine the code and lightly leaned on to improve the accuracy and readability of the text. The former was mainly in an attempt to address Claude’s almost endless array of suggestions for adding more code, while the latter was mainly a result of having stared too long at my own writing.
No doubt errors remain as it’s quite the topic, which means I’ve made quite a number of edits to both the text and code. Feel free to contact me here.
Additional Note:
The motivation for writing this was the absence of descriptive analysis on the links between law and order and economic growth aimed at a general audience. Keeping the code in the post has probably cost it some readability, but my reasoning for sharing the analysis so openly was to make it easier for others to build on this post to fill the many gaps out there (including me in my future posts).
The inspiration for this series came from work I completed in 2025 for the Bingham Centre for the Rule of Law and the Law Society of England and Wales. A paper based on the work summarizing research on the topic is available here.
- For a great summary of the research, see Dr Lopez-Gomez, L. (2026, The Rule of Law and the Institutional Roots of Economic Performance. The British Institute of International and Comparative Law, (link). Also see: Haggard, S., MacIntyre, A. and Tiede, L., 2008. The rule of law and economic development. Annu. Rev. Polit. Sci., 11(1), pp.205-234; and Besley, T., Bogart, D., Chapman, J. and Nuno, P., 2025. Justices of the peace: Legal foundations of the industrial revolution (No. 20214). Centre for Economic Policy Research.
︎ - One of the better attempts I’ve seen to empirically test a causal connection between the two: Besley, T., Bogart, D., Chapman, J. and Nuno, P., 2025. Justices of the peace: Legal foundations of the industrial revolution (No. 20214). Centre for Economic Policy Research.
︎
The post Institutions Are a Package Deal: What a Correlation Can’t Tell You About the Rule of Law appeared first on Giles.
R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.