Population Genomics · PB 495/595 Plant Evolutionary Biology
10 September 2026
By the end you should be able to:
Local adaptation occurs when a population evolves to be more fit in its native environment than other populations. It is the result of spatially heterogeneous selection.
Three examples of local adaptation (a-c) versus no local adapation (d) from Kawecki and Ebert (2004)
A cline is a gradient of allele frequency across an environmental transition
Adaptation to heavy metal contamination in soils is a classic example, where populations near mine boundaries evolve metal tolerance over very short spatial scales.
Zinc tolerance in the grass Anthoxanthum odoratum (Poaceae) flips over a few meters across a mine edge, despite pollen and seed flow — selection is that strong (Antonovics, Bradshaw, and Turner 1971).
Silene uniflora (Caryophyllaceae) similarly shows local adaptation to heavy metal contamination in soils. Perennial with a storage root that is subject to high exposure in contaminated environments.
Silene uniflora. Photo Credit: Christopher Carlin.
Zinc and copper tolerenace from Silene in Papadopulos et al. (2021)
Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).
Distribution of inversion types Lowry and Willis (2010)
Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).
Visible phenotypic effect of inversion
Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).
Visible phenotypic effect of inversion
Flowering time can be very important if you need to reproduce before it gets dry. Different flowering times between populations can also create prezygotic reproductive isolation. A well-known case of this is a chromosomal inversion in Mimulus guttatus that causes early-flowering annuals versus late-flowering perennials (Lowry and Willis 2010).
Landscape genomic analysis of candidate genes for climate adaptation in a California endemic oak, Quercus lobata
Individual genes can be identified through selection scans or an redundancy analysis to connect loci with environemnt too (Sork et al. 2016)
Individual candidate loci indicated in table
Risk is that the \(F_{ST}\) outliers may not reflect local adaptation but rather neutral processes like population structure or background selection. A null distribution under the coalescent is needed to help avoid false positives.
Convergent evolution in loss of photoperiod regulation in areas with reduced photoperiods in wheat. (Zhao et al. 2023)
Manhattan plots map to loss of function in a candidate locus
Regress allele frequency directly on an environmental variable \(E\), correcting for population structure. The correction ensure fixed variation due to distance is not confused with environmental association \[ \mathrm{logit}\,p_\ell = \beta_0 + \beta_1 E + (\text{structure / random effect}) . \]
\(p_\ell\) is the allele frequency at locus \(\ell\) in a sampled population or individual, and \(E\) is an environmental measurement at that same site. The logit is \(\ln\!\big(p/(1-p)\big)\), which transforms the allele frequency to an unbound continuous distribution that can be used for regression.
\(\beta_0\) is the intercept, and \(\beta_1\) is the slope. The slope measures how much the locus’s frequency shifts per unit change in environment.
If \(\beta_1 \neq 0\), then the locus is associated with environment more than expected due to chance alone.
The regression before is univariate: one locus, one environemntal variable. But local adaptation is often polygenic: many allele frequencies changing a little together - weak selection.
Testing one locus at a time causes two big statistical problems: - Multiple testing — millions of SNPs × several climate variables. This will lead to a high false positive rate or a punishing multiple testing correction across thousands of loci. - It misses covariation. We often want to know which loci covary across clines or gradients.
RDA analyzes everything at once. Two matrices:
RDA is another ordination analysis, similar to our PCA. But, this is a constrained ordination. That means it is a PCA of the portion of genetic variation explained by the environmental predictors.
Multiple regression is used to associate genetic variation with environment. But here, environment is a multivariate PC axis. From Capblancq and Forester (2021).
Schematic overview of multiple regression analysis with RDA
Example of RDA applied to pines along with their species distribution models from Capblancq and Forester (2021)
Potential for predicting future mismatch between today’s locally adapted genotypes and future environmental conditions from Capblancq and Forester (2021)
How much environment? The total genetic variation in \(\mathbf{Y}\) splits into the component explained by environment and a component not explained by environment. The environment fraction gives an \(R^2\) like an ordinary linear regression. Report the adjusted \(R^2\) instead of the raw one, because raw \(R^2\) always goes up when you add more predictors, even useless ones.
Is it real? A high \(R^2\) can still show up by chance, so a permutation test helps differentiate a strong result from a weak one. Here, we randomly shuffle which environmental values go with which site and refit the model many times. The \(R^2\) is recorded every permutation. If we do this enough, we get a null distribution to put a p-value on \(R^2\).
Candidates are SNPs with extreme loadings on the significant RDA axes — the tails of the loading distribution (a common rule is \(\pm 3\) SD) (Forester et al. 2018).
The axis a SNP loads on tells you which environmental gradient it is associated with.
This approach is part of a family of methods that are very powerful in genetics. Here are a few main flavors of the approach.
| Method | Flavor | Structure control | Best at |
|---|---|---|---|
| RDA | multivariate ordination | covariates via pRDA | polygenic, covarying signal |
| LFMM | univariate + latent factors | latent factors | single strong-effect loci |
| BayPass | Bayesian | allele-frequency covariance null | explicit demographic null |
PB 495/595 · Plant Evolutionary Biology · Population Genomics