Population Structure

Population Genomics · PB 495/595 Plant Evolutionary Biology

George P. Tiley

27 August 2026

Continuity - variable diversity among species

A facultative apomict versus and obligate outcrosser can have different levels of diversity.

Nucleotide diversity can vary by an order of magnitude between similar species

Continuity - variable diversity among species

There can be biological and environmental factors underlying how much variation is maintained in a population. Inbreeding leads to structure.

Capsella is a great system to highlight how mating systems can affect inbreeding. We mentioned that C. grandiflora is self-incompatible and an obligate outcrosser. The closely related C. rubella is self-compatible and primarily selfing. They shared a common ancestor approximately 200,000 years ago, and one might assume that a shift to selfing in C. rubella has led to a dramatic reduction in genetic diversity compared to C. grandiflora. The effect is demonstrated by estimates of \(\pi\) and the site frequency spectrum.

Selfing reduces genetic diversity with respect to outcrossing species.

Diversity and selection from Capsella species from Slotte et al. (2013)

Causes of inbreeding and structure

Multiple factors contribute to genetic strcuture across plants.

Models explaining $F_{ST}$ across plants.

Models explaining \(F_{ST}\) across plants from Gamba and Muchhala (2020)

Causes of inbreeding and structure

What are some general features that reduce or increase inbreeding?

Marginal effects of features on $F_{ST}$ across plants.

Marginal effects of features on \(F_{ST}\) across plants from Gamba and Muchhala (2020)

Can the drivers of inbreeding at the population level help explain global biodiversity?

Range size and niche breadth across plants.

Range size and niche breadth across plants from Moulatlet et al. (2025)

Can the drivers of inbreeding at the population level help explain global biodiversity?

To be a cradle of diversity, a region needs to create many opportunities for inbreeding, structure, and subsequent differentiation.

Higher elevation should impose smaller ranges across plants.

Higher elevation should impose smaller ranges. Moulatlet et al. (2025)

Learning objectives

By the end you should be able to:

  • Explain why inbreeding and \(F\)-statistics are relative to a reference population.
  • Define Wright’s \(F_{IS}\), \(F_{IT}\), \(F_{ST}\).
  • Use the Wahlund effect to explain heterozygote deficits in pooled samples.
  • Explain the diversity ceiling;
  • Explain how PCA works at a conceptual level
  • Read a PCA and an STRUCTURE/ADMIXTURE plot.
  • Describe isolation-by-distance

Inbreeding is relative

Sardinia

Inbreeding suggests that two individuals are more related than random chance alone, but from which population?

Mountains create structure among villages in Sardinia.

Mountains create structure among villages in Sardinia. Fig. 1 from Chiang et al. (2018)

Sardinia

Relative isolation among some villages from multiple analyses.

Relative isolation among some villages from multiple analyses. Fig. 2 from Chiang et al. (2018)

Sardinia

Suddenly, the whole island looks highly inbred compared to the region.

Suddenly, the whole island looks highly inbred compared to the region. Fig. 4 from Chiang et al. (2018)

Expressing deficits of heterozygosity as fixation (\(F\)) statistics

Sewall Wright (1889-1988) defined the fixation indices for measuring the degree of inbreeding at nested levels: the individual (\(I\)), a subpopulation (\(S\)), and the total (\(T\)) population. Comparisons could be made to infer patterns of isolation among subpopulations, for example. The original derivations do not take four books but are lengthy and not fun.

Photograph of Sewall Wright at the University of Chicago.

Sewell Wright, undated, at the University of Chicago. Photo Credit: Ernest Martin

Wright’s F-statistics: a hierarchy

\(F_{XY}\) = the correlation between random gametes drawn at level \(X\), relative to level \(Y\). The three classic comparisons:

  • \(F_{IS}\) — individual vs. subpopulation
  • \(F_{IT}\) — individual vs. total
  • \(F_{ST}\) — subpopulation vs. total

\(F_{IS}\), \(F_{IT}\) — observed vs. expected heterozygosity

Let \(H_I = f_{12}\) be the observed heterozygosity (covered last lecture), and within subpopulation (or deme) \(S\) (allele frequency \(p_S\)) the HW expectation is \(H_S = 2p_S q_S\).

\[ F_{IS} = 1 - \frac{H_I}{H_S} = 1 - \frac{f_{12}}{2 p_S q_S}. \]

Comparing the individual instead to the total population (\(p_T\)):

\[ F_{IT} = 1 - \frac{H_I}{H_T} = 1 - \frac{f_{12}}{2 p_T q_T}. \]

Important

  • \(F_{IS} > 0\): heterozygote deficit within a deme (selfing, assortative mating).
  • \(F_{IT} > 0\): heterozygote deficit relative to the total population.

\(F_{ST}\) and the decomposition

Compare expected heterozygosity in the subpopulation to that in the total:

\[ F_{ST} = 1 - \frac{H_S}{H_T} = 1 - \frac{2 p_S q_S}{2 p_T q_T}. \]

The three statistics multiply through the nested heterozygosities:

\[ (1 - F_{IT}) = \frac{H_I}{H_S}\cdot\frac{H_S}{H_T} = (1 - F_{IS})(1 - F_{ST}). \]

Important

\(F_{ST} > 0\): heterozygote deficit due to differences among demes.

Worked example — a selfing herb

A selfing herb sampled within one meadow (\(S\)) and range-wide (\(T\)):

\[ p_S = 0.10,\quad f_{12} = 0.09,\quad p_T = 0.20. \]

\[ \begin{aligned} F_{IS} &= 1 - \frac{0.09}{2(0.1)(0.9)} = 1 - \frac{0.09}{0.18} = 0.50 \\[4pt] F_{ST} &= 1 - \frac{2(0.1)(0.9)}{2(0.2)(0.8)} = 1 - \frac{0.18}{0.32} = 0.44 \\[4pt] F_{IT} &= 1 - (1-0.50)(1-0.44) = 1 - 0.28 = 0.72 . \end{aligned} \]

  • High \(F_{IS}\) (selfing) and high \(F_{ST}\) (differentiated meadows) compound.

The Wahlund effect

Average over \(K\) subpopulations sampled equally:

\[ \bar F_{ST} = 1 - \frac{\bar H_S}{H_T},\qquad \bar H_S = \frac{1}{K}\sum_{i=1}^{K} 2 p_i q_i . \]

Because the average of \(2p_iq_i\) is always \(\leq 2\bar p\,\bar q\), pooling demes deficits heterozygotes even if every deme is in perfect HW:

Simulating the Wahlund effect

See script: wahlund_effect.R

Demonstration of Wahlund effect with two HW populations pooled together make it look like the observed heterozygosity is less than expected.

Demonstration of Wahlund effect - two HW populations pooled together make it look like the observed heterozygosity is less than expected

Some useful extensions

Beyond two alleles: \(G_{ST}\)

\(F_{ST}\) as written assumes two alleles. Nei generalized it using gene diversity (Nei 1973):

\[ H = 1 - \sum_i p_i^2, \qquad G_{ST} = \frac{H_T - H_S}{H_T}. \]

  • Identical in form to \(F_{ST}\) — just with the multi-allelic \(H\).
  • For a biallelic locus, \(H = 2pq\) and \(G_{ST} = F_{ST}\).
  • \(G_{ST}\) was more prevalent for older marker types such as microsatellites, cpDNA haplotypes, and AFLPs. Still relevant but multiallelic sites are often ignored with modern genomic data.

The ceiling problem

\(G_{ST}\) has an awkward property: it is bounded above by the within-population diversity.

\[ \max G_{ST} \;\approx\; 1 - H_S . \]

Consider a highly variable locus such as a microsatellite. If we have \(H_S \approx 0.9\), \(G_{ST}\) cannot exceed \(\approx 0.1\). This is true even if the populations share no alleles. A low \(F_{ST}\) from one type of data has no relevance on another. This has led to some concern about reliability of \(F_{ST}\) results across independent studies over time, but we should never compare \(F_{ST}/G_{ST}\) across data types.

Alternative measures of population differentiation

Measure Idea Reference Use
\(G_{ST}\) fixation index (deficit of heterozygosity) (Nei 1973) multiallelic \(F_{ST}\)
\(\rho\) average relatedness of individuals within populations compared to the whole (Ronfort et al. 1998) comparable across ploidy levels in a mixed system
\(F''_{ST}\) \(F_{ST}\) standardized by its maximum (Hedrick 2005) high-diversity markers
\(D_{est}\) effective numbers of alleles (Jost 2008) not sensitive to pop size or ploidy if heterosygosity is at equilibrium

Although these different measure exist, \(F_{ST}\) is almost always reported and maybe alternatives alongside. Correcting biases in polyploids remains an active area of research. (Meirmans, Liu, and Tienderen 2018).

Different ways to estimate \(F_{ST}\)

The formulas so far define a parameter. Real data need an estimator that accounts for finite samples. The below estimators account for sample size differences across demes.

  • Weir & Cockerham’s \(\theta\) (Weir and Cockerham 1984) — corrects for sample size and the number of demes; this is what most software reports as “\(F_{ST}\).”
  • Hudson’s estimator (Hudson, Slatkin, and Maddison 1992) — simpler and often preferred for pairwise comparisons.
  • Hudson’s estimator for sliding windows (Bhatia et al. 2013): sum numerators and denominators across loci, then divide. Averaging per-locus \(F_{ST}\) biases results, badly so with rare variants.

Estimates can come out negative. While annoying it means the between-population component is indistinguishable from zero. Report it or truncate to 0, but say which.

Visualizing structure

The data: a genotype matrix

Start with \(N\) individuals genotyped at \(S\) biallelic SNPs. Individual \(i\) at locus \(\ell\):

\[ g_{i\ell} \in \{0, 1, 2\} \quad(\text{copies of allele } A_1), \]

giving an \(N \times S\) matrix — and in genomics \(N \ll S\) (hundreds of individuals, thousands of SNPs).

  • Every SNP is an axis: the data live in \(S\)-dimensional space, far too many to look at.
  • A principal component analysis (PCA) will find a few axes that capture most of the variation, and place each individual on them.

Standardizing the genotypes

Before decomposing, each SNP is standardized to have mean 0 and variance 1. This process is sometimes called z-score normalization or centering and scaling

\[ \frac{g_{i\ell} - 2p_\ell}{\sqrt{2p_\ell(1 - p_\ell)}} . \]

  • Center the \(i^\text{th}\) individual’s genotype for site \(j\) by the mean genotype \(2p_{j}\).
  • Scale by \(\sqrt{2p_{j}(1-p_{j})}\), the binomial SD expected if alleles were drawn at frequency \(p_j\).

Note

Without scaling, common SNPs (large \(2pq\)) would dominate purely because they vary more. Standardizing gives each SNP the same weight.

Singular value decomposition

In practice softwares run an SVD on the standardized genotypematrix \(\mathbf{X}\):

\[ \mathbf{X} = \mathbf{U}\,\mathbf{D}\,\mathbf{V}^{\!\top}. \]

  • \(\mathbf{U}\mathbf{D}\) → PC scores (where each individual sits).
  • \(\mathbf{V}\) → SNP loadings (how much each locus contributes to each axis).
  • \(\mathbf{D}\) → The square roots of Eigenvalues that can be used to calculate the variance explained by each PC.

\[ \text{Var}[\text{PC}_k] = \frac{d_k^2}{\sum_j d_j^2} \]

Some programs report the square roots of Eigenvalues (\(d_k\)) directly, while others report the Eigenvalues. Check when calculating the proportions of variance explained.

Reading a PCA in practice

  • Model-free and fast — no assumed \(K\), unlike ADMIXTURE (next).
  • Admixed individuals fall between their parental clusters — a hybrid sits on the line connecting the two sources, roughly in proportion to ancestry.
  • Caveats: PCs reflect sampling design; uneven sample sizes and families/relatives can create or distort axes.

Example of population structure in a PCA.

Structure exists among inbred lines of Helianthus annuus. Fig. 5 from Darvishzadeh et al. (2026)

Model-based clustering (STRUCTURE / ADMIXTURE)

  • Genetic variation can be represented as a mixture of \(K\) HW groups.
  • Assign each individual to one or more of the \(K\) clusters. This is frequently visualized as a stacked-bar plot.
  • Determining \(K\) is a both objective and subjective
Determining the best $K$ in a model-based clustering analysis.

Population structure represented as a mixture of \(K\) clusters. Fig. 1 from Ross-Ibarra et al. (2008)

Structure in space

Isolation by distance

Not all structure comes in discrete demes. When dispersal is limited, differentiation grows continuously with geographic distance. Wright referred to this as isolation-by-distance (Wright 1943).

  • Neighbors mate more often than distant individuals.
  • IBD is usually the default expectation rather than the exception.

Rousset’s regression (Rousset 1997):

\[ \frac{F_{ST}}{1 - F_{ST}} \;\sim\; \log(\text{distance}) . \]

  • We are interested in the slope, not the absolute values of \(F_{ST}\).
  • Significance can be assessed with a Mantel test. Distances are permuted to generate a null distribution.

Effective migration surfaces (EEMS / FEEMS)

PCA tells you that samples are spatially structured. Effective migration surfaces use the variation to help visualize potential barriers to gene flow in space (Petkova, Novembre, and Stephens 2016).

  • Fit a stepping-stone model on a spatial grid overlaid on the sampling region.
  • Map deviations from expectations under isolation by distance:
    • corridors — migration higher than expected,
    • barriers — migration lower than expected.
  • FEEMS (Marcus et al. 2021) fits the same idea far faster for genome-scale data.
The EEMS for Arabidopsis thaliana.

The EEMS for Arabidopsis thaliana. Fig. 6 from Petkova, Novembre, and Stephens (2016)

→ formal migration models and \(F_{ST}=1/(1+4N_em)\) and environmental gradients are forthcoming.

Key takeaways

Summary — key equations

\[ F_{IS} = 1 - \frac{H_I}{H_S} \qquad F_{ST} = 1 - \frac{H_S}{H_T} \qquad (1-F_{IT}) = (1-F_{IS})(1-F_{ST}) \]

\[ G_{ST} = \frac{H_T - H_S}{H_T} \qquad \max G_{ST} \approx 1 - H_S \qquad \frac{F_{ST}}{1-F_{ST}} \sim \text{distance} \]

Review Questions

  • Does inbreeding reduce variation in a population?
  • How does the Wahlund effect work?
  • Define Wright’s fixation indicies.
  • Can you interpret a PCA and ADMIXTURE plot?
  • If you compare multiple populations along a gradient, do you have expectations of genetic differentatiation between them?

References

Bhatia, Gaurav, Nick Patterson, Sriram Sankararaman, and Alkes L. Price. 2013. “Estimating and Interpreting F_ST: The Impact of Rare Variants.” Genome Research 23 (9): 1514–21. https://doi.org/10.1101/gr.154831.113.
Chiang, Charleston WK, Joseph H Marcus, Carlo Sidore, Arjun Biddanda, Hussein Al-Asadi, Magdalena Zoledziewska, Maristella Pitzalis, et al. 2018. “Genomic History of the Sardinian Population.” Nature Genetics 50 (10): 1426–34.
Darvishzadeh, Reza, Hadi Alipour, Aras Türkoğlu, Kamil Haliloğlu, and Sima Fatanatvash. 2026. “Genome-Wide Assessment of Genetic Diversity and Selective Signatures in Sunflower (Helianthus Annuus l.) Using a 10 k SNP Array.” Scientific Reports 16 (1): 9439.
Gamba, Diana, and Nathan Muchhala. 2020. “Global Patterns of Population Genetic Differentiation in Seed Plants.” Molecular Ecology 29 (18): 3413–28.
Hedrick, Philip W. 2005. “A Standardized Genetic Differentiation Measure.” Evolution 59 (8): 1633–38. https://doi.org/10.1111/j.0014-3820.2005.tb01814.x.
Hudson, Richard R., Montgomery Slatkin, and Wayne P. Maddison. 1992. “Estimation of Levels of Gene Flow from DNA Sequence Data.” Genetics 132 (2): 583–89.
Jost, Lou. 2008. “G_ST and Its Relatives Do Not Measure Differentiation.” Molecular Ecology 17 (18): 4015–26. https://doi.org/10.1111/j.1365-294X.2008.03887.x.
Leffler, Ellen M, Kevin Bullaughey, Daniel R Matute, Wynn K Meyer, Laure Segurel, Aarti Venkat, Peter Andolfatto, and Molly Przeworski. 2012. “Revisiting an Old Riddle: What Determines Genetic Diversity Levels Within Species?”
Marcus, Joseph, Wooseok Ha, Rina Foygel Barber, and John Novembre. 2021. “Fast and Flexible Estimation of Effective Migration Surfaces.” eLife 10: e61927. https://doi.org/10.7554/eLife.61927.
Meirmans, Patrick G, Shenglin Liu, and Peter H van Tienderen. 2018. “The Analysis of Polyploid Genetic Data.” Journal of Heredity 109 (3): 283–96.
Moulatlet, Gabriel M, Cory Merow, Brian Maitner, Brad Boyle, Xiao Feng, Amy E Frazier, Cesar Hinojo-Hinojo, et al. 2025. “General Laws of Biodiversity: Climatic Niches Predict Plant Range Size and Ecological Dominance Globally.” Proceedings of the National Academy of Sciences 122 (46): e2517585122.
Nei, Masatoshi. 1973. “Analysis of Gene Diversity in Subdivided Populations.” Proceedings of the National Academy of Sciences 70 (12): 3321–23. https://doi.org/10.1073/pnas.70.12.3321.
Petkova, Desislava, John Novembre, and Matthew Stephens. 2016. “Visualizing Spatial Population Structure with Estimated Effective Migration Surfaces.” Nature Genetics 48 (1): 94–100. https://doi.org/10.1038/ng.3464.
Ronfort, Joëlle, Eric Jenczewski, Thomas Bataillon, and François Rousset. 1998. “Analysis of Population Structure in Autotetraploid Species.” Genetics 150 (2): 921–30.
Ross-Ibarra, Jeffrey, Stephen I Wright, John Paul Foxe, Akira Kawabe, Leah DeRose-Wilson, Gesseca Gos, Deborah Charlesworth, and Brandon S Gaut. 2008. “Patterns of Polymorphism and Demographic History in Natural Populations of Arabidopsis Lyrata.” PloS One 3 (6): e2411.
Rousset, François. 1997. “Genetic Differentiation and Estimation of Gene Flow from F-Statistics Under Isolation by Distance.” Genetics 145 (4): 1219–28. https://doi.org/10.1093/genetics/145.4.1219.
Slotte, Tanja, Khaled M Hazzouri, J Arvid Ågren, Daniel Koenig, Florian Maumus, Ya-Long Guo, Kim Steige, et al. 2013. “The Capsella Rubella Genome and the Genomic Consequences of Rapid Mating System Evolution.” Nature Genetics 45 (7): 831–35.
Weir, B. S., and C. Clark Cockerham. 1984. “Estimating F-Statistics for the Analysis of Population Structure.” Evolution 38 (6): 1358–70. https://doi.org/10.2307/2408641.
Wright, Sewall. 1943. “Isolation by Distance.” Genetics 28 (2): 114–38. https://doi.org/10.1093/genetics/28.2.114.