The Mediterranean Punic World Was Genetically Diverse, With Substantial Levantine Ancestry

Ringbauer et al. argued in their 2025 paper, “Punic people were genetically diverse with almost no Levantine ancestors”, that Punic people had almost no Levantine ancestors. Their conclusion seems to be based mainly on qpAdm rotation results. The source choices were distal. Instead of using Iron Age groups from around the Mediterranean, including the newly sequenced Akhziv Phoenician group, the study used Bronze Age groups. It also used Neolithic Ganj Dareh as a source. ...

September 24, 2026

Interpreting f4-Statistics with AdmixPy

f4-statistics can be used to test asymmetries in allele sharing between populations. They measure the covariance between allele-frequency differences across two pairs of populations. Theory and formula An f4-statistic is the average, across SNPs, of the product of the allele-frequency differences between two pairs of populations: f4(A,B;C,D)=Ei[(pA,i−pB,i)(pC,i−pD,i)] f_4(A,B;C,D)=\mathbb{E}_i\left[(p_{A,i}-p_{B,i})(p_{C,i}-p_{D,i})\right] f4​(A,B;C,D)=Ei​[(pA,i​−pB,i​)(pC,i​−pD,i​)]Multiplying out results in: ...

September 16, 2026

A Potential Early Dynastic Kish G2a Individual (I25719) with Upper Mesopotamian Neolithic Ancestry

Among the unreleased samples in the Akbari et al. dataset is an individual (IID: I25719) rumoured to be from Kish, dated in the supplementary table to the Early Dynastic III period, around 2450 BCE. His Y-DNA haplogroup is G-FT19860, a downstream branch of G2a, one of the main paternal lineages associated with Anatolian Neolithic farmers, and his mtDNA haplogroup is R0a1a. The sample has decent coverage: 598,059 “compatibility” SNPs were called, 645,472 were missing. ...

September 12, 2026

A Potential Bronze Age Anatolian-Derived V1636 Sample (I41584)

Among the Akbari et al. dataset, there is a previously unreported individual (Sample IID: I41584/I41584_preQC) potentially from Anatolia. His terminal Y-DNA subclade is Y148982/Y106006 (TMRCA ~3100 BCE according to FTDNA), a lineage that also includes most modern West Asian and Near Eastern V1636 samples. His maternal haplogroup is HV29d1, also according to FTDNA. ...

September 4, 2026

Ust-Ishim: A 45,000-Year-Old Genome at the East–West Eurasian Split

Ust-Ishim is a sample identified from only a femur bone pulled out of eroding sand on the banks of the Irtysh River in western Siberia in 2008. Later analysis established that the bone belonged to a man estimated to have lived around 45,000 years ago. His genome is one of the earliest Upper Paleolithic WGS genomes. Neither Clearly East- nor West Eurasian What makes this sample particularly interesting is that he does not fall clearly into either the East or West Eurasian category. The initial Fu et al. study placed him before, or approximately at, the separation of subsequent eastern and western Eurasian populations, which include all modern Eurasian populations. Later graph modelling reached essentially the same conclusion; the best fit placed Ust-Ishim slightly towards the West Eurasian branch, but the uncertainty overlapped the East-West split. ...

August 24, 2026

Convert 23andMe, AncestryDNA, MyHeritage & FTDNA Raw DNA to PLINK (BED/BIM/FAM)

To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (.bed, .bim, .fam), you will have to first convert the raw file to 23andMe format. You can then convert it with PLINK 1.9 using --23file. Converting Raw DNA to 23andMe Format with AWK Windows users can use WSL to access awk; see How to Download the AADR Dataset (Linux & WSL). If your DNA file is already in 23andMe format, skip this section. ...

August 19, 2026

Genetic Traits of Loschbour: Appearance, Height, Blood Type, and More

I was inferring genetic traits of ancient individuals, among them the “Cheddar Man”, whose pigmentation phenotype I thought was well established from his genotype. However, most of the 58 trait markers I was checking for could not be called reliably (with MAPQ ≥30 and base quality ≥30). Most markers had no reads at all, several others were supported by only a single read. This included markers like HERC2/OCA2 rs12913832 for eye colour, likewise SLC24A5 and SLC45A2 used to infer skin pigmentation. ...

August 18, 2026

F4Mix: Sample-Wise Ancestry Fitting with f4 Statistics

Last week I published F4Mix, a tool for fitting modern and ancient DNA samples against a pool of source populations, usually ancient ones. F4Mix estimates, for each target, the non-negative mixture of reference populations whose covariance-aware f4 profile best matches it. This makes it useful for testing every sample against the same sources. With a proper setup, the tool gives meaningful results, and can reveal both substructure and clear outliers within a site. ...

August 15, 2026

Testing for Admixture with f3-Statistics in AdmixPy

f3-statistics are used to test if populations are admixed or to measure shared genetic drift between two populations relative to an outgroup. This post explains the theory behind admixture f3-statistics and shows how to run admixture f3 tests with AdmixPy. If you want to skip the theoretical part, you can jump to Running admixture f3-statistics in AdmixPy. What is an f3-statistic? For three populations, the statistic is written as: f3(A;B,C)=Ei[(pA,i−pB,i)(pA,i−pC,i)] f_3(A;B,C)=\mathbb{E}_i\left[(p_{A,i}-p_{B,i})(p_{A,i}-p_{C,i})\right] f3​(A;B,C)=Ei​[(pA,i​−pB,i​)(pA,i​−pC,i​)]Here, AAA is in the target position. Populations BBB and CCC are the reference populations. The values pA,ip_{A,i}pA,i​, pB,ip_{B,i}pB,i​, and pC,ip_{C,i}pC,i​ are the allele frequencies in populations AAA, BBB, and CCC, respectively, at SNP iii. The expectation is an average across SNPs. ...

August 7, 2026

Are Higher qpAdm P-Values Better?

Yes. For two qpAdm models with the same target, the same right groups, and the same settings, the model with the higher p-value is the better statistical fit. qpAdm calculates a covariance-weighted discrepancy between the observed and fitted f4-statistics. The p-value reflects how well the model explains the used f4-statistics. A higher p-value means the discrepancy between the observed and fitted values is less unusual under the model. This does not mean that the model with the highest p-value for a target is automatically the best one, because qpAdm results depend on the selected right groups. Uninformative right-groups can lack the power to detect a bad model, while overly restrictive ones can make a plausible model appear to fit badly. Therefore, p-values are more comparable when models for the same target are compared using the same groups and settings. Models with different numbers of sources are also comparable since p-values account for different degrees of freedom. Z-scores can be used to assess whether an additional source is justified. ...

July 30, 2026

Why the Best PCA Fit May Still Be the Wrong Admixture Model

A Vahaduo generated PCA model for Sardinians gives: 82.8% Barcin Neolithic 11.6% Loschbour 5.6% Yamnaya Distance: 3.4303% Ganj Dareh was included in the sources but gets a weight of zero. This seems to imply that Sardinians don’t have any eastern-Farmer related ancestry. When Sardinians are modelled with qpAdm using Barcin Neolithic, Loschbour, Yamnaya, and Ganj Dareh, the model fits well: 68.6% Barcin Neolithic 11.9% Loschbour 10.2% Yamnaya 9.4% Ganj Dareh Neolithic p = 0.769 When Ganj Dareh is dropped, the model fails (p=1.18×10−12p = 1.18 \times 10^{-12}p=1.18×10−12). ...

July 22, 2026

Pairwise f2 Statistics and FST in AdmixPy

This post covers how to run pairwise f2-statistics and FST in AdmixPy. They are simple to interpret, and are also useful computationally. Once f2 blocks have been computed and cached, many downstream analyses can reuse them without repeatedly reading and converting the original genotype data. What does f2 measure? The f2-statistic quantifies allele-frequency differentiation between two populations, AAA and BBB, and is defined as: f2(A,B)=E[(pA−pB)2] f_2(A, B) = E[(p_A - p_B)^2] f2​(A,B)=E[(pA​−pB​)2]where pAp_ApA​ and pBp_BpB​ are the allele frequencies of populations AAA and BBB at a SNP, and the squared allele-frequency differences are averaged across SNPs. ...

June 30, 2026

Introducing AdmixPy: f-statistics, qpAdm, and qpWave in Python

I recently published AdmixPy on GitHub, a fast implementation of f-statistics, qpAdm, and qpWave in Python that runs on Linux, macOS, and Windows. It works directly on the new AADR TGENO distribution format and is notably faster than ADMIXTOOLS 2 and simpler to set up. Supported input formats: EIGENSTRAT (.geno/.snp/.ind), packed AncestryMap (.geno/.snp/.ind), TGENO (.tgeno/.snp/.ind), and SNP-major PLINK binary (.bed/.bim/.fam). AdmixPy is implemented in Python and depends only on NumPy, SciPy, and pandas. Installation is handled through pip, and it should behave the same on every platform. ...

May 21, 2026

Convert Raw DNA Files to EIGENSTRAT for ADMIXTOOLS and Merge with AADR

Commercial raw DNA exports are not provided in the file formats normally used by ADMIXTOOLS, ADMIXTOOLS 2, AADR-based workflows, or PLINK. Files from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, and Living DNA are usually plain-text vendor exports, while downstream workflows often require PLINK PACKEDPED or EIGENSTRAT/PACKEDANCESTRYMAP files. EIGENSTRAT is often used loosely to refer to the .geno/.snp/.ind triplet. Strictly speaking, EIGENSTRAT is the plain-text version of that triplet; PACKEDANCESTRYMAP is the packed binary form of the same three files. ADMIXTOOLS and ADMIXTOOLS 2 work with either, but PACKEDANCESTRYMAP takes far less disk space and loads much faster, which is why it’s the practical default used here. ...

May 15, 2026

The Genetic Origins of the Proto-Anatolians

The origins of the Proto-Anatolians are often treated as one of the more obscure problems, but the genetic data may be not that ambigous. Anatolian is regarded as the earliest-splitting branch of “Indo-European”, and its divergence is deep enough that some linguists distinguish a pre–Proto-Indo-European stage, sometimes called “Indo-Anatolian”, from the Proto-Indo-European reconstructed from the non-Anatolian branches. Under either framing, the relevant question is the same: whether the earlier Eneolithic steppe-related ancestry behind Yamnaya, particularly the Caucasus–Lower Volga (CLV) component, also moved south of the Caucasus into Anatolia. For this purpose, I use Progress-2 specifically as proxy for the north Caucasus-facing part of this Eneolithic steppe-related ancestry, since it sits directly at the northern end of the Caucasus and therefore serves as a good proxy for groups that may have passed through the region. ...

May 10, 2026

Downloading and Converting AADR v66

Recently, in April 2026, new AADR versions were released on Harvard Dataverse. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable. ...

April 17, 2026

dt: A Modern awk Alternative for Common Data Workflows

I recently published dt, a modern data transformation tool designed to make the awk workflows commonly used on this blog more intuitive, expressive, and fast. Dt is written in Rust because it compiles to a single binary that runs anywhere, and it uses Polars for the actual data processing, giving you columnar operations that handle large files efficiently. The syntax uses explicit functions (filter(), select(), mutate()) chained together with pipes, making common transformations easier to read and modify. There’s also an interactive REPL that shows you the result after each operation, letting you build complex pipelines step-by-step, catch mistakes early, and undo errors with .undo. ...

February 11, 2026

Fast, Transparent f4-Based Admixture Screening in R

In this post, I will build a transparent admixture-screening workflow from scratch in R using f4-statistics and constrained regression. The main advantage is automation: instead of hand-writing every candidate model, the script tests many 2-way, 3-way, and 4-way source combinations in one pass and ranks them by fit. ADMIXTOOLS 2 already includes batch tools such as qpadm_multi() and qpadm_rotate(), so the point is not that qpAdm cannot be automated. The point is that this custom workflow is compact, transparent, easy to modify, and useful for exploratory model search before you validate the strongest candidates more formally. ...

February 3, 2026

How to Merge EIGENSTRAT Datasets Using mergeit

mergeit is part of the EIGENSOFT package and can be used to merge exactly two EIGENSTRAT/PACKEDANCESTRYMAP datasets. Setting up EIGENSOFT mergeit is part of the EIGENSOFT package. You can install it via conda: conda install -c bioconda eigensoft Alternatively, if you prefer to compile from source, see: From EIGENSTRAT to PACKEDPED. Setting Up A Parameter File Like other EIGENSOFT tools, mergeit requires a parameter file: geno1: aadr.geno snp1: aadr.snp ind1: aadr.ind geno2: eigenstrat_output.geno snp2: eigenstrat_output.snp ind2: eigenstrat_output.ind genooutfilename: merged.geno snpoutfilename: merged.snp indoutfilename: merged.ind Save this as mergeit.par and run: ...

January 12, 2026

Pseudohaploid Genotyping for Ancient DNA: BAM to EIGENSTRAT

In this post, I’ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the previous post, we already know it’s aligned to hs37d5. However, if you’re starting with a BAM file, you’ll need to verify the reference genome first. I’ll start by showing how to check BAM headers to identify the reference genome. Identifying the Reference Genome from BAM Headers Before processing any BAM file, you should verify which reference genome it was aligned against. This is critical because AADR compatibility requires hs37d5 specifically. BAMs aligned to other GRCh37-based references like hg19 are also compatible (since they share the same coordinate system, differing only in chromosome naming conventions), but hg38/GRCh38 BAMs would require realignment from FASTQs. ...

January 4, 2026

Processing Ancient DNA: From FASTQ to Aligned BAM

This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workflow is based on the run I used for ERR14088885, from the Başur Höyük study PRJEB83032. Ancient DNA needs a different alignment strategy from ordinary modern whole-genome data. The molecules are short, the ends may carry post-mortem damage, and paired reads often overlap because the DNA insert is shorter than the sequencing cycles. For this sample I therefore clean poly-G tails, trim adapters, merge overlapping paired-end reads, align the merged molecules with bwa aln, remove low-confidence alignments, and deduplicate using both observed ends of each molecule. ...

January 2, 2026

Running qpAdm with ADMIXTOOLS2 in R: Testing and Interpreting Ancestry Models

This post covers using qpAdm in R to test ancestry models and estimate admixture proportions. qpAdm builds on f4-statistics and provides a framework for evaluating whether proposed source populations can explain a target population’s genetic makeup. For R and admixtools setup instructions on Debian/Ubuntu, see my previous post: Running f4-Statistics with Admixtools in R. Windows users can find R installation instructions on the R website. What is qpAdm? qpAdm is a method for testing ancestry models and estimating admixture proportions. It determines whether a target population can be modeled as a mixture of specified source populations (“left populations”), and if the model fits, calculates the contribution from each source. The method builds on f4-statistics (covered in my previous post) to evaluate these ancestry models. ...

December 16, 2025

How to Run and Interpret f4-Statistics in R: AADR Examples

This post covers how to run f4-statistics using the admixtools package for R. Compared with the original ADMIXTOOLS workflow, the R implementation is more convenient for testing multiple population combinations because it can be used interactively, without repeatedly editing parameter files. For more in-depth examples, see Interpreting f4-Statistics with AdmixPy. What are f4-statistics? F4-statistics measure asymmetry in allele sharing among four populations. For four populations AAA, BBB, CCC, and DDD, the statistic is written as: ...

November 28, 2025

MyHeritage's $30 WGS: Technical Analysis

Recently, MyHeritage quietly announced that all new DNA kits will be processed using low-pass whole-genome sequencing (WGS) instead of traditional genotyping arrays, for just $30 per kit. The coverage depth will be roughly 2x (as announced in their blog post), compared to the 30x used in clinical sequencing. It’s nowhere near diagnostic quality, but it is whole-genome data nevertheless. The surprising part is the price. Getting an entire genome, even at shallow depth, for less than what MyHeritage charges to “unlock” an uploaded SNP file (~$35) seems almost too cheap. Below are a few of the issues I see: technical, economic, and data-security related. ...

November 11, 2025

How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)

In an earlier post, PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output), I showed the manual way to build a subset from the .ind/.fam. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using awk that generates a PLINK --keep file automatically from a list of populations. Prepare a list of populations to keep: Create a text file (e.g. pops) in the same directory as your reference .ind and .fam. Put one population label per line: ...

November 10, 2025

Running ADMIXTURE in Supervised Mode

This post is a short follow-up to the previous one on Estimating Ancestry Components Using ADMIXTURE. Here, we’ll explore supervised ADMIXTURE, a mode that allows you to explicitly define ancestral populations and infer the ancestry proportions of unassigned individuals based on those references. What Is a Supervised Run? In supervised mode, ADMIXTURE skips the component discovery step and instead uses user-defined groupings to represent ancestral components. The benefit: if you already have solid candidates for reference populations, you can use them to quickly infer ancestry proportions for target or admixed individuals. ...

August 5, 2025

How to Run ADMIXTURE (Unsupervised): Full Tutorial & Python Plotting Script

In this post, I’ll demonstrate how to estimate ancestry proportions using one of the most widely used tools in population genetics: ADMIXTURE. ADMIXTURE is a model-based clustering algorithm that estimates individual ancestry proportions and ancestral allele frequencies from multilocus SNP genotypes. Preparing the Dataset Download the appropriate ADMIXTURE binary and either place it in your dataset directory or make it globally accessible. For this run, I included a subset of West Asian populations along with a few adjacent populations (around 150 samples in total). Linkage Disequilibrium (LD) pruning was applied beforehand. If you’re unsure how to prune your dataset, refer to the previous post. ...

August 2, 2025

SmartPCA Tutorial: How to Run PCA on Genetic Data

This post is a continuation of the previous one, where I demonstrated how to perform PCA with PLINK. While PLINK’s PCA is great for quick, exploratory analysis, smartpca (part of the EIGENSOFT toolset) is particularly common in population-genetic and ancient-DNA studies. Smartpca can be compiled from the EIGENSOFT source or installed through conda. I covered the installation process in this earlier post: From EIGENSTRAT to PACKEDPED. As before, I’ll use a small subset. The focus here is on the technical process. One key difference in this post is that I’ll perform Linkage Disequilibrium (LD) pruning, which reduces redundancy between correlated SNPs before PCA. ...

July 30, 2025

PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)

In this post, I’ll demonstrate how to perform a PCA on a PLINK dataset. Before we begin, we need to prepare a subset of samples we’re interested in analyzing. To do this, we’ll extract sample information from the .fam file. But first, we need to identify the samples of interest. For example, those from a specific population such as Sardinians. The easiest way is to open the corresponding .ind file and look at the population column, which is the third column in each row. Open the file in a text editor, and search for the population name, in this case, Sardinian. ...

July 29, 2025

Converting EIGENSTRAT/PACKEDANCESTRYMAP to PACKEDPED

The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style .geno/.snp/.ind dataset. This naming can be confusing: the .snp and .ind files are the usual EIGENSTRAT metadata files, but the .geno file may either be plain-text EIGENSTRAT or binary PACKEDANCESTRYMAP. PACKEDPED format allows for easier downstream processing using the PLINK toolset. With PLINK, it becomes straightforward to extract sample subsets, filter SNPs, and perform a wide range of analyses. ...

July 29, 2025