Interpreting f4-Statistics with AdmixPy

f4-statistics can be used to test asymmetries in allele sharing between populations. They measure the covariance between allele-frequency differences across two pairs of populations. Theory and formula An f4-statistic is the average, across SNPs, of the product of the allele-frequency differences between two pairs of populations: f4(A,B;C,D)=Ei[(pA,i−pB,i)(pC,i−pD,i)] f_4(A,B;C,D)=\mathbb{E}_i\left[(p_{A,i}-p_{B,i})(p_{C,i}-p_{D,i})\right] f4​(A,B;C,D)=Ei​[(pA,i​−pB,i​)(pC,i​−pD,i​)]Multiplying out results in: ...

September 16, 2026

F4Mix: Sample-Wise Ancestry Fitting with f4 Statistics

Last week I published F4Mix, a tool for fitting modern and ancient DNA samples against a pool of source populations, usually ancient ones. F4Mix estimates, for each target, the non-negative mixture of reference populations whose covariance-aware f4 profile best matches it. This makes it useful for testing every sample against the same sources. With a proper setup, the tool gives meaningful results, and can reveal both substructure and clear outliers within a site. ...

August 15, 2026

Testing for Admixture with f3-Statistics in AdmixPy

f3-statistics are used to test if populations are admixed or to measure shared genetic drift between two populations relative to an outgroup. This post explains the theory behind admixture f3-statistics and shows how to run admixture f3 tests with AdmixPy. If you want to skip the theoretical part, you can jump to Running admixture f3-statistics in AdmixPy. What is an f3-statistic? For three populations, the statistic is written as: f3(A;B,C)=Ei[(pA,i−pB,i)(pA,i−pC,i)] f_3(A;B,C)=\mathbb{E}_i\left[(p_{A,i}-p_{B,i})(p_{A,i}-p_{C,i})\right] f3​(A;B,C)=Ei​[(pA,i​−pB,i​)(pA,i​−pC,i​)]Here, AAA is in the target position. Populations BBB and CCC are the reference populations. The values pA,ip_{A,i}pA,i​, pB,ip_{B,i}pB,i​, and pC,ip_{C,i}pC,i​ are the allele frequencies in populations AAA, BBB, and CCC, respectively, at SNP iii. The expectation is an average across SNPs. ...

August 7, 2026

Pairwise f2 Statistics and FST in AdmixPy

This post covers how to run pairwise f2-statistics and FST in AdmixPy. They are simple to interpret, and are also useful computationally. Once f2 blocks have been computed and cached, many downstream analyses can reuse them without repeatedly reading and converting the original genotype data. What does f2 measure? The f2-statistic quantifies allele-frequency differentiation between two populations, AAA and BBB, and is defined as: f2(A,B)=E[(pA−pB)2] f_2(A, B) = E[(p_A - p_B)^2] f2​(A,B)=E[(pA​−pB​)2]where pAp_ApA​ and pBp_BpB​ are the allele frequencies of populations AAA and BBB at a SNP, and the squared allele-frequency differences are averaged across SNPs. ...

June 30, 2026

Introducing AdmixPy: f-statistics, qpAdm, and qpWave in Python

I recently published AdmixPy on GitHub, a fast implementation of f-statistics, qpAdm, and qpWave in Python that runs on Linux, macOS, and Windows. It works directly on the new AADR TGENO distribution format and is notably faster than ADMIXTOOLS 2 and simpler to set up. Supported input formats: EIGENSTRAT (.geno/.snp/.ind), packed AncestryMap (.geno/.snp/.ind), TGENO (.tgeno/.snp/.ind), and SNP-major PLINK binary (.bed/.bim/.fam). AdmixPy is implemented in Python and depends only on NumPy, SciPy, and pandas. Installation is handled through pip, and it should behave the same on every platform. ...

May 21, 2026

Fast, Transparent f4-Based Admixture Screening in R

In this post, I will build a transparent admixture-screening workflow from scratch in R using f4-statistics and constrained regression. The main advantage is automation: instead of hand-writing every candidate model, the script tests many 2-way, 3-way, and 4-way source combinations in one pass and ranks them by fit. ADMIXTOOLS 2 already includes batch tools such as qpadm_multi() and qpadm_rotate(), so the point is not that qpAdm cannot be automated. The point is that this custom workflow is compact, transparent, easy to modify, and useful for exploratory model search before you validate the strongest candidates more formally. ...

February 3, 2026

How to Run and Interpret f4-Statistics in R: AADR Examples

This post covers how to run f4-statistics using the admixtools package for R. Compared with the original ADMIXTOOLS workflow, the R implementation is more convenient for testing multiple population combinations because it can be used interactively, without repeatedly editing parameter files. For more in-depth examples, see Interpreting f4-Statistics with AdmixPy. What are f4-statistics? F4-statistics measure asymmetry in allele sharing among four populations. For four populations AAA, BBB, CCC, and DDD, the statistic is written as: ...

November 28, 2025