Convert 23andMe, AncestryDNA, MyHeritage & FTDNA Raw DNA to PLINK (BED/BIM/FAM)

To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (.bed, .bim, .fam), you will have to first convert the raw file to 23andMe format. You can then convert it with PLINK 1.9 using --23file. Converting Raw DNA to 23andMe Format with AWK Windows users can use WSL to access awk; see How to Download the AADR Dataset (Linux & WSL). If your DNA file is already in 23andMe format, skip this section. ...

August 19, 2026

Convert Raw DNA Files to EIGENSTRAT for ADMIXTOOLS and Merge with AADR

Commercial raw DNA exports are not provided in the file formats normally used by ADMIXTOOLS, ADMIXTOOLS 2, AADR-based workflows, or PLINK. Files from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, and Living DNA are usually plain-text vendor exports, while downstream workflows often require PLINK PACKEDPED or EIGENSTRAT/PACKEDANCESTRYMAP files. EIGENSTRAT is often used loosely to refer to the .geno/.snp/.ind triplet. Strictly speaking, EIGENSTRAT is the plain-text version of that triplet; PACKEDANCESTRYMAP is the packed binary form of the same three files. ADMIXTOOLS and ADMIXTOOLS 2 work with either, but PACKEDANCESTRYMAP takes far less disk space and loads much faster, which is why it’s the practical default used here. ...

May 15, 2026

Downloading and Converting AADR v66

Recently, in April 2026, new AADR versions were released on Harvard Dataverse. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable. ...

April 17, 2026

dt: A Modern awk Alternative for Common Data Workflows

I recently published dt, a modern data transformation tool designed to make the awk workflows commonly used on this blog more intuitive, expressive, and fast. Dt is written in Rust because it compiles to a single binary that runs anywhere, and it uses Polars for the actual data processing, giving you columnar operations that handle large files efficiently. The syntax uses explicit functions (filter(), select(), mutate()) chained together with pipes, making common transformations easier to read and modify. There’s also an interactive REPL that shows you the result after each operation, letting you build complex pipelines step-by-step, catch mistakes early, and undo errors with .undo. ...

February 11, 2026

How to Merge EIGENSTRAT Datasets Using mergeit

mergeit is part of the EIGENSOFT package and can be used to merge exactly two EIGENSTRAT/PACKEDANCESTRYMAP datasets. Setting up EIGENSOFT mergeit is part of the EIGENSOFT package. You can install it via conda: conda install -c bioconda eigensoft Alternatively, if you prefer to compile from source, see: From EIGENSTRAT to PACKEDPED. Setting Up A Parameter File Like other EIGENSOFT tools, mergeit requires a parameter file: geno1: aadr.geno snp1: aadr.snp ind1: aadr.ind geno2: eigenstrat_output.geno snp2: eigenstrat_output.snp ind2: eigenstrat_output.ind genooutfilename: merged.geno snpoutfilename: merged.snp indoutfilename: merged.ind Save this as mergeit.par and run: ...

January 12, 2026

Pseudohaploid Genotyping for Ancient DNA: BAM to EIGENSTRAT

In this post, I’ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the previous post, we already know it’s aligned to hs37d5. However, if you’re starting with a BAM file, you’ll need to verify the reference genome first. I’ll start by showing how to check BAM headers to identify the reference genome. Identifying the Reference Genome from BAM Headers Before processing any BAM file, you should verify which reference genome it was aligned against. This is critical because AADR compatibility requires hs37d5 specifically. BAMs aligned to other GRCh37-based references like hg19 are also compatible (since they share the same coordinate system, differing only in chromosome naming conventions), but hg38/GRCh38 BAMs would require realignment from FASTQs. ...

January 4, 2026

Processing Ancient DNA: From FASTQ to Aligned BAM

This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workflow is based on the run I used for ERR14088885, from the Başur Höyük study PRJEB83032. Ancient DNA needs a different alignment strategy from ordinary modern whole-genome data. The molecules are short, the ends may carry post-mortem damage, and paired reads often overlap because the DNA insert is shorter than the sequencing cycles. For this sample I therefore clean poly-G tails, trim adapters, merge overlapping paired-end reads, align the merged molecules with bwa aln, remove low-confidence alignments, and deduplicate using both observed ends of each molecule. ...

January 2, 2026

How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)

In an earlier post, PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output), I showed the manual way to build a subset from the .ind/.fam. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using awk that generates a PLINK --keep file automatically from a list of populations. Prepare a list of populations to keep: Create a text file (e.g. pops) in the same directory as your reference .ind and .fam. Put one population label per line: ...

November 10, 2025

Converting EIGENSTRAT/PACKEDANCESTRYMAP to PACKEDPED

The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style .geno/.snp/.ind dataset. This naming can be confusing: the .snp and .ind files are the usual EIGENSTRAT metadata files, but the .geno file may either be plain-text EIGENSTRAT or binary PACKEDANCESTRYMAP. PACKEDPED format allows for easier downstream processing using the PLINK toolset. With PLINK, it becomes straightforward to extract sample subsets, filter SNPs, and perform a wide range of analyses. ...

July 29, 2025

How to Download the AADR Dataset (Linux & WSL)

Note: This post uses an older AADR release and parts of it may now be outdated. For the latest AADR v66 download, including TGENO conversion and ADMIXTOOLS2 compatibility notes, see Downloading and Converting AADR v66. A Linux environment is unavoidable when it comes to bioinformatical data processing and preparation. You can use your favorite distribution. For Windows users, the Windows Subsystem for Linux (WSL) provides a good alternative to dual booting or setting up a full virtual machine. ...

July 29, 2025