Storytelling and story testing in domestication
Gerbault P, Allaby RG, Boivin N, Rudzinski A, Grimaldi IM, Pires JC, Climer Vigueira C, Dobney K, Gremillion KJ, Barton L, Arroyo-Kalin M, Purugganan MD, Rubio de Casas R, Bollongino R, Burger J, Fuller DQ, Bradley DG, Balding DJ, Richerson PJ, Gilbert MT, Larson G, Thomas MG.
Abstract
different histories (equifinality); (ii) any particular history can potentially give rise to a wide range of different patterns in data [evolutionary variance (12)]; and (iii) certain demographic or adaptation histories can give rise to counter intuitive patterns in data (emergence) (e.g., refs.13–16). Evolutionary histories are rarely directly “revealed” by looking only at patterns in data, such as the distribution of particular markers (e.g., morphological traits, material culture, genetic variants). This is because such data may be only weakly constrained by those histories; many different histories may explain the same data equally well. Thus, instead of simply providing narratives based on interpretations, implicit assumptions, and preconceived ideas, domestication histories need to be tested to identify those scenarios that best explain observed data; to do this, domestication histories must be modeled explicitly. We outline a range of modeling techniques that can be used in domestication research and provide examples that illustrate their utility. Although most of the discussion and examples given in this paper are based on population genetic data, most of the principles and approaches can also be applied to other datasets used to explore domestication processes. Types of Modeling Approaches A model is an explicit and simplified representation of the underlying causative mechanisms in a system and is used to make predictions about the observed outcomes (data) of that system. We consider two classes of models: discriminative models, which fit directly observed data to predicted relationships (e.g., linear regression), and generative models, which are intended to capture the main real-world mechanisms that generate data, and are typically used to produce artificial datasets. Discriminative models make assumptions, sometimes but not always explicitly, on the ways aspects of the data are correlated without specifying the actual mechanisms that generate those correlations (e.g., refs. 17–21). Generative models aim to explicitly replicate key hypothesized (i.e., assumed) processes that generate the data. Because all evolutionary processes include stochastic elements, a range of different outcomes—or patterns in empirical data—can be generated from any particular scenario or model. For this reason, when using generative models, it is often necessary to produce many datasets by simulation. In population genetics, a powerful means of simulating data is the “retrospective” approach of coalescent simulation (22), where the joining (or coalescence) of lineages is simulated backward in time under specific assumptions about such variables as population size, structure, migration, and admixture. This approach is highly efficient because it only simulates the lineage history of the sample, not of the whole population, so simulation can be very fast. However, coalescent approaches are limited in the demographic and selection scenarios that can be modeled, and
This page indexes the study's public bibliographic record. The full text belongs to the journal; follow the DOI above to read it at the source.