Local Adaptation and Environmental Associations

Population Genomics · PB 495/595 Plant Evolutionary Biology

George P. Tiley

10 September 2026

Continuing from selection

  • Selection leaves local marks on diversity against a genome-wide background.
  • Selection acts on phenotypes, and sometimes, advantageous phenotypes are dependent on the environmental context. The same allele is good in one place, bad in another.
  • We want to combine information about selection with environmental associations to understand local adaptation.

Learning objectives

By the end you should be able to:

  • Define local adaptation.
  • Understand how \(F_{ST}\) outliers and GEA detect adaptive loci, and their pitfalls.
  • Explain RDA as constrained ordination — regression then PCA of the fitted values — and how partial RDA corrects for population structure.
  • Explain why controlling for structure is critical in GEA, and how to validate candidates experimentally.

Selection and local adaptation

Local adaptation

Local adaptation occurs when a population evolves to be more fit in its native environment than other populations. It is the result of spatially heterogeneous selection.

Three examples of local adaptation (a-c) versus no local adaptation (d) from Kawecki & Ebert 2004.

Three examples of local adaptation (a-c) versus no local adapation (d) from Kawecki and Ebert (2004)

Clines and a plant example

A cline is a gradient of allele frequency across an environmental transition

Adaptation to heavy metal contamination in soils is a classic example, where populations near mine boundaries evolve metal tolerance over very short spatial scales.

Zinc tolerance in the grass Anthoxanthum odoratum (Poaceae) flips over a few meters across a mine edge, despite pollen and seed flow — selection is that strong (Antonovics, Bradshaw, and Turner 1971).

Clines and a plant example

Silene uniflora (Caryophyllaceae) similarly shows local adaptation to heavy metal contamination in soils. Perennial with a storage root that is subject to high exposure in contaminated environments.

Silene uniflora. Photo by Christopher Carlin.

Silene uniflora. Photo Credit: Christopher Carlin.

Clines and a plant example

Silene uniflora. Photo by Christopher Carlin.

Zinc and copper tolerenace from Silene in Papadopulos et al. (2021)

More cases of local adaptation in plants

Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).

Map of Mimulus guttatus from the northwestern US.

Distribution of inversion types Lowry and Willis (2010)

More cases of local adaptation in plants

Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).

Plant flowering time greatly affected by phenotype.

Visible phenotypic effect of inversion

More cases of local adaptation in plants

Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).

Plant flowering time greatly affected by phenotype.

Visible phenotypic effect of inversion

Local adaptation can be driven be a few loci

Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).

Landscape genomic analysis of candidate genes for climate adaptation in a California endemic oak, Quercus lobata.

Landscape genomic analysis of candidate genes for climate adaptation in a California endemic oak, Quercus lobata

Local adaptation can be driven be a few loci

Individual genes can be identified through selection scans or an redundancy analysis to connect loci with environemnt too (Sork et al. 2016)

Individual candidate loci indicated in table from Sork et al. 2016.

Individual candidate loci indicated in table

Finding the loci

Method I — \(F_{ST}\) outlier scans

  • Local adaptation inflates \(F_{ST}\) at the responsible loci above the genome-wide (demographic) background.
  • Flag loci in the upper tail of the genome-wide \(F_{ST}\) distribution.

Risk is that the \(F_{ST}\) outliers may not reflect local adaptation but rather neutral processes like population structure or background selection. A null distribution under the coalescent is needed to help avoid false positives.

Method I — \(F_{ST}\) outlier scans

Convergent evolution in loss of photoperiod regulation in areas with reduced photoperiods in wheat. (Zhao et al. 2023)

Individual candidate loci indicated in table from Zhao et al. 2023.

Manhattan plots map to loss of function in a candidate locus

Redundancy analysis (RDA)

Method II — genotype–environment association (GEA)

Regress allele frequency directly on an environmental variable \(E\), correcting for population structure. The correction ensure fixed variation due to distance is not confused with environmental association \[ \mathrm{logit}\,p_\ell = \beta_0 + \beta_1 E + (\text{structure / random effect}) . \]

\(p_\ell\) is the allele frequency at locus \(\ell\) in a sampled population or individual, and \(E\) is an environmental measurement at that same site. The logit is \(\ln\!\big(p/(1-p)\big)\), which transforms the allele frequency to an unbound continuous distribution that can be used for regression.

\(\beta_0\) is the intercept, and \(\beta_1\) is the slope. The slope measures how much the locus’s frequency shifts per unit change in environment.

If \(\beta_1 \neq 0\), then the locus is associated with environment more than expected due to chance alone.

Univariate limitations

The regression before is univariate: one locus, one environemntal variable. But local adaptation is often polygenic: many allele frequencies changing a little together - weak selection.

Testing one locus at a time causes two big statistical problems: - Multiple testing — millions of SNPs × several climate variables. This will lead to a high false positive rate or a punishing multiple testing correction across thousands of loci. - It misses covariation. We often want to know which loci covary across clines or gradients.

Multivariate solution: RDA

RDA analyzes everything at once. Two matrices:

  • \(\mathbf{Y}\): \(n \times L\) — populations (or individuals) × loci (allele frequencies/genotypes).
  • \(\mathbf{X}\): \(n \times p\) — the same sites × environmental predictors.

RDA is another ordination analysis, similar to our PCA. But, this is a constrained ordination. That means it is a PCA of the portion of genetic variation explained by the environmental predictors.

Multivariate solution: RDA

Multiple regression is used to associate genetic variation with environment. But here, environment is a multivariate PC axis. From Capblancq and Forester (2021).

Schematic overview of multiple regression analysis with RDA from Capblancq et al. 2021.

Schematic overview of multiple regression analysis with RDA

Multivariate solution: RDA

Example of RDA applied to pines along with their species distribution models from Capblancq et al. 2021.

Example of RDA applied to pines along with their species distribution models from Capblancq and Forester (2021)

Multivariate solution: RDA

Potential for predicting future mismatch between today's locally adapted genotypes and future environmental conditions from Capblancq et al. 2021.

Potential for predicting future mismatch between today’s locally adapted genotypes and future environmental conditions from Capblancq and Forester (2021)

\(R^2\) and significance

How much environment? The total genetic variation in \(\mathbf{Y}\) splits into the component explained by environment and a component not explained by environment. The environment fraction gives an \(R^2\) like an ordinary linear regression. Report the adjusted \(R^2\) instead of the raw one, because raw \(R^2\) always goes up when you add more predictors, even useless ones.

Is it real? A high \(R^2\) can still show up by chance, so a permutation test helps differentiate a strong result from a weak one. Here, we randomly shuffle which environmental values go with which site and refit the model many times. The \(R^2\) is recorded every permutation. If we do this enough, we get a null distribution to put a p-value on \(R^2\).

Pulling out candidate loci

Candidates are SNPs with extreme loadings on the significant RDA axes — the tails of the loading distribution (a common rule is \(\pm 3\) SD) (Forester et al. 2018).

The axis a SNP loads on tells you which environmental gradient it is associated with.

This approach is part of a family of methods that are very powerful in genetics. Here are a few main flavors of the approach.

Method Flavor Structure control Best at
RDA multivariate ordination covariates via pRDA polygenic, covarying signal
LFMM univariate + latent factors latent factors single strong-effect loci
BayPass Bayesian allele-frequency covariance null explicit demographic null

Key takeaways

Summary - Review questions

  • What would you expect if you performed reciprocal transplants of locally adapted populations along their primary environmental gradient?
  • How do \(F_{ST}\) outliers tests differ from genome scans with Tajima’s \(D\) scans?
  • How do you prevent an RDA from giving you false positives due to population structure?
  • How would you design a population genomics project to detect local adaptation along a precipitation gradient? What would your strategy be for sampling populations and how many individuals per population?

References

Antonovics, Janis, A. D. Bradshaw, and R. G. Turner. 1971. “Heavy Metal Tolerance in Plants.” Advances in Ecological Research 7: 1–85. https://doi.org/10.1016/S0065-2504(08)60202-0.
Capblancq, Thibaut, and Brenna R. Forester. 2021. “Redundancy Analysis: A Swiss Army Knife for Landscape Genomics.” Methods in Ecology and Evolution 12 (12): 2298–2309. https://doi.org/10.1111/2041-210X.13722.
Forester, Brenna R., Jesse R. Lasky, Helene H. Wagner, and Dean L. Urban. 2018. “Comparing Methods for Detecting Multilocus Adaptation with Multivariate Genotype–Environment Associations.” Molecular Ecology 27 (9): 2215–33. https://doi.org/10.1111/mec.14584.
Kawecki, Tadeusz J, and Dieter Ebert. 2004. “Conceptual Issues in Local Adaptation.” Ecology Letters 7 (12): 1225–41.
Lowry, David B, and John H Willis. 2010. “A Widespread Chromosomal Inversion Polymorphism Contributes to a Major Life-History Transition, Local Adaptation, and Reproductive Isolation.” PLoS Biology 8 (9): e1000500.
Papadopulos, Alexander ST, Andrew J Helmstetter, Owen G Osborne, Aaron A Comeault, Daniel P Wood, Edward A Straw, Laurence Mason, et al. 2021. “Rapid Parallel Adaptation to Anthropogenic Heavy Metal Pollution.” Molecular Biology and Evolution 38 (9): 3724–36.
Sork, Victoria L, Kevin Squire, Paul F Gugger, Stephanie E Steele, Eric D Levy, and Andrew J Eckert. 2016. “Landscape Genomic Analysis of Candidate Genes for Climate Adaptation in a California Endemic Oak, Quercus Lobata.” American Journal of Botany 103 (1): 33–46.
Zhao, Xuebo, Yafei Guo, Lipeng Kang, Changbin Yin, Aoyue Bi, Daxing Xu, Zhiliang Zhang, et al. 2023. “Population Genomics Unravels the Holocene History of Bread Wheat and Its Relatives.” Nature Plants 9 (3): 403–19.