A Potential Bronze Age Anatolian-Derived V1636 Sample (I41584)

Among the Akbari et al. dataset, there is a previously unreported individual (Sample IID: I41584/I41584_preQC) potentially from Anatolia. His terminal Y-DNA subclade is Y148982/Y106006 (TMRCA ~3100 BCE according to FTDNA), a lineage that also includes most modern West Asian and Near Eastern V1636 samples. His maternal haplogroup is HV29d1, also according to FTDNA. ...

September 4, 2026

Ust-Ishim: A 45,000-Year-Old Genome at the East–West Eurasian Split

Ust-Ishim is a sample identified from only a femur bone pulled out of eroding sand on the banks of the Irtysh River in western Siberia in 2008. Later analysis established that the bone belonged to a man estimated to have lived around 45,000 years ago. His genome is one of the earliest Upper Paleolithic WGS genomes. Neither Clearly East- nor West Eurasian What makes this sample particularly interesting is that he does not fall clearly into either the East or West Eurasian category. The initial Fu et al. study placed him before, or approximately at, the separation of subsequent eastern and western Eurasian populations, which include all modern Eurasian populations. Later graph modelling reached essentially the same conclusion; the best fit placed Ust-Ishim slightly towards the West Eurasian branch, but the uncertainty overlapped the East-West split. ...

August 24, 2026

Genetic Traits of Loschbour: Appearance, Height, Blood Type, and More

I was inferring genetic traits of ancient individuals, among them the “Cheddar Man”, whose pigmentation phenotype I thought was well established from his genotype. However, most of the 58 trait markers I was checking for could not be called reliably (with MAPQ ≥30 and base quality ≥30). Most markers had no reads at all, several others were supported by only a single read. This included markers like HERC2/OCA2 rs12913832 for eye colour, likewise SLC24A5 and SLC45A2 used to infer skin pigmentation. ...

August 18, 2026

The Genetic Origins of the Proto-Anatolians

The origins of the Proto-Anatolians are often treated as one of the more obscure problems, but the genetic data may be not that ambigous. Anatolian is regarded as the earliest-splitting branch of “Indo-European”, and its divergence is deep enough that some linguists distinguish a pre–Proto-Indo-European stage, sometimes called “Indo-Anatolian”, from the Proto-Indo-European reconstructed from the non-Anatolian branches. Under either framing, the relevant question is the same: whether the earlier Eneolithic steppe-related ancestry behind Yamnaya, particularly the Caucasus–Lower Volga (CLV) component, also moved south of the Caucasus into Anatolia. For this purpose, I use Progress-2 specifically as proxy for the north Caucasus-facing part of this Eneolithic steppe-related ancestry, since it sits directly at the northern end of the Caucasus and therefore serves as a good proxy for groups that may have passed through the region. ...

May 10, 2026

Downloading and Converting AADR v66

Recently, in April 2026, new AADR versions were released on Harvard Dataverse. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable. ...

April 17, 2026

Pseudohaploid Genotyping for Ancient DNA: BAM to EIGENSTRAT

In this post, I’ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the previous post, we already know it’s aligned to hs37d5. However, if you’re starting with a BAM file, you’ll need to verify the reference genome first. I’ll start by showing how to check BAM headers to identify the reference genome. Identifying the Reference Genome from BAM Headers Before processing any BAM file, you should verify which reference genome it was aligned against. This is critical because AADR compatibility requires hs37d5 specifically. BAMs aligned to other GRCh37-based references like hg19 are also compatible (since they share the same coordinate system, differing only in chromosome naming conventions), but hg38/GRCh38 BAMs would require realignment from FASTQs. ...

January 4, 2026

Processing Ancient DNA: From FASTQ to Aligned BAM

This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workflow is based on the run I used for ERR14088885, from the Başur Höyük study PRJEB83032. Ancient DNA needs a different alignment strategy from ordinary modern whole-genome data. The molecules are short, the ends may carry post-mortem damage, and paired reads often overlap because the DNA insert is shorter than the sequencing cycles. For this sample I therefore clean poly-G tails, trim adapters, merge overlapping paired-end reads, align the merged molecules with bwa aln, remove low-confidence alignments, and deduplicate using both observed ends of each molecule. ...

January 2, 2026

How to Download the AADR Dataset (Linux & WSL)

Note: This post uses an older AADR release and parts of it may now be outdated. For the latest AADR v66 download, including TGENO conversion and ADMIXTOOLS2 compatibility notes, see Downloading and Converting AADR v66. A Linux environment is unavoidable when it comes to bioinformatical data processing and preparation. You can use your favorite distribution. For Windows users, the Windows Subsystem for Linux (WSL) provides a good alternative to dual booting or setting up a full virtual machine. ...

July 29, 2025