<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Data Preparation on PopGenetics Blog</title><link>https://popgenetics.dev/topics/data-preparation/</link><description>Recent content in Data Preparation on PopGenetics Blog</description><generator>Hugo -- 0.148.2</generator><language>en-us</language><lastBuildDate>Wed, 19 Aug 2026 22:17:00 +0200</lastBuildDate><atom:link href="https://popgenetics.dev/topics/data-preparation/index.xml" rel="self" type="application/rss+xml"/><item><title>Convert 23andMe, AncestryDNA, MyHeritage &amp; FTDNA Raw DNA to PLINK (BED/BIM/FAM)</title><link>https://popgenetics.dev/posts/raw-dna-to-plink/</link><pubDate>Wed, 19 Aug 2026 22:17:00 +0200</pubDate><guid>https://popgenetics.dev/posts/raw-dna-to-plink/</guid><description>&lt;p>To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (&lt;code>.bed&lt;/code>, &lt;code>.bim&lt;/code>, &lt;code>.fam&lt;/code>), you will have to first convert the raw file to 23andMe format. You can then convert it with PLINK 1.9 using &lt;code>--23file&lt;/code>.&lt;/p>
&lt;hr>
&lt;h2 id="converting-raw-dna-to-23andme-format-with-awk">Converting Raw DNA to 23andMe Format with AWK&lt;/h2>
&lt;p>Windows users can use WSL to access &lt;code>awk&lt;/code>; see &lt;a href="https://popgenetics.dev/posts/download-ancient-modern-dna-aadr/">How to Download the AADR Dataset (Linux &amp;amp; WSL)&lt;/a>.&lt;/p>
&lt;p>If your DNA file is already in 23andMe format, skip this section.&lt;/p></description></item><item><title>Convert Raw DNA Files to EIGENSTRAT for ADMIXTOOLS and Merge with AADR</title><link>https://popgenetics.dev/posts/raw-dna-to-admixtools/</link><pubDate>Fri, 15 May 2026 15:55:55 +0200</pubDate><guid>https://popgenetics.dev/posts/raw-dna-to-admixtools/</guid><description>&lt;p>Commercial raw DNA exports are not provided in the file formats normally used by ADMIXTOOLS, ADMIXTOOLS 2, AADR-based workflows, or PLINK. Files from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, and Living DNA are usually plain-text vendor exports, while downstream workflows often require PLINK PACKEDPED or EIGENSTRAT/PACKEDANCESTRYMAP files.&lt;/p>
&lt;p>EIGENSTRAT is often used loosely to refer to the &lt;code>.geno&lt;/code>/&lt;code>.snp&lt;/code>/&lt;code>.ind&lt;/code> triplet. Strictly speaking, EIGENSTRAT is the plain-text version of that triplet; PACKEDANCESTRYMAP is the packed binary form of the same three files. ADMIXTOOLS and ADMIXTOOLS 2 work with either, but PACKEDANCESTRYMAP takes far less disk space and loads much faster, which is why it&amp;rsquo;s the practical default used here.&lt;/p></description></item><item><title>Downloading and Converting AADR v66</title><link>https://popgenetics.dev/posts/aadr-v66-download-and-conversion/</link><pubDate>Fri, 17 Apr 2026 13:55:43 +0200</pubDate><guid>https://popgenetics.dev/posts/aadr-v66-download-and-conversion/</guid><description>&lt;p>Recently, in April 2026, new AADR versions were released on &lt;a href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW">Harvard Dataverse&lt;/a>. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable.&lt;/p></description></item><item><title>dt: A Modern awk Alternative for Common Data Workflows</title><link>https://popgenetics.dev/posts/data-transform/</link><pubDate>Wed, 11 Feb 2026 14:37:59 +0100</pubDate><guid>https://popgenetics.dev/posts/data-transform/</guid><description>&lt;p>I recently published &lt;a href="https://github.com/system0x7/dt">dt&lt;/a>, a modern data transformation tool designed to make the awk workflows commonly used on this blog more intuitive, expressive, and fast. Dt is written in Rust because it compiles to a single binary that runs anywhere, and it uses Polars for the actual data processing, giving you columnar operations that handle large files efficiently. The syntax uses explicit functions (&lt;code>filter()&lt;/code>, &lt;code>select()&lt;/code>, &lt;code>mutate()&lt;/code>) chained together with pipes, making common transformations easier to read and modify. There&amp;rsquo;s also an interactive REPL that shows you the result after each operation, letting you build complex pipelines step-by-step, catch mistakes early, and undo errors with &lt;code>.undo&lt;/code>.&lt;/p></description></item><item><title>How to Merge EIGENSTRAT Datasets Using mergeit</title><link>https://popgenetics.dev/posts/mergeit-tutorial/</link><pubDate>Mon, 12 Jan 2026 06:31:35 +0100</pubDate><guid>https://popgenetics.dev/posts/mergeit-tutorial/</guid><description>&lt;p>mergeit is part of the EIGENSOFT package and can be used to merge exactly two EIGENSTRAT/PACKEDANCESTRYMAP datasets.&lt;/p>
&lt;hr>
&lt;h2 id="setting-up-eigensoft">Setting up EIGENSOFT&lt;/h2>
&lt;p>mergeit is part of the EIGENSOFT package. You can install it via conda:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>conda install -c bioconda eigensoft
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Alternatively, if you prefer to compile from source, see: &lt;a href="https://popgenetics.dev/posts/convert-eigenstrat-to-packedped/">From EIGENSTRAT to PACKEDPED&lt;/a>.&lt;/p>
&lt;hr>
&lt;h2 id="setting-up-a-parameter-file">Setting Up A Parameter File&lt;/h2>
&lt;p>Like other EIGENSOFT tools, mergeit requires a parameter file:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>geno1: aadr.geno
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>snp1: aadr.snp
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>ind1: aadr.ind
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>geno2: eigenstrat_output.geno
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>snp2: eigenstrat_output.snp
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>ind2: eigenstrat_output.ind
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>genooutfilename: merged.geno
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>snpoutfilename: merged.snp
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>indoutfilename: merged.ind
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Save this as &lt;code>mergeit.par&lt;/code> and run:&lt;/p></description></item><item><title>Pseudohaploid Genotyping for Ancient DNA: BAM to EIGENSTRAT</title><link>https://popgenetics.dev/posts/pileup-to-eigenstrat/</link><pubDate>Sun, 04 Jan 2026 00:44:12 +0100</pubDate><guid>https://popgenetics.dev/posts/pileup-to-eigenstrat/</guid><description>&lt;p>In this post, I&amp;rsquo;ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the previous post, we already know it&amp;rsquo;s aligned to hs37d5. However, if you&amp;rsquo;re starting with a BAM file, you&amp;rsquo;ll need to verify the reference genome first. I&amp;rsquo;ll start by showing how to check BAM headers to identify the reference genome.&lt;/p>
&lt;hr>
&lt;h2 id="identifying-the-reference-genome-from-bam-headers">Identifying the Reference Genome from BAM Headers&lt;/h2>
&lt;p>Before processing any BAM file, you should verify which reference genome it was aligned against. This is critical because AADR compatibility requires hs37d5 specifically. BAMs aligned to other GRCh37-based references like hg19 are also compatible (since they share the same coordinate system, differing only in chromosome naming conventions), but hg38/GRCh38 BAMs would require realignment from FASTQs.&lt;/p></description></item><item><title>Processing Ancient DNA: From FASTQ to Aligned BAM</title><link>https://popgenetics.dev/posts/ancient-dna-alignment-bwa-tutorial/</link><pubDate>Fri, 02 Jan 2026 14:48:44 +0100</pubDate><guid>https://popgenetics.dev/posts/ancient-dna-alignment-bwa-tutorial/</guid><description>&lt;p>This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workflow is based on the run I used for ERR14088885, from the Başur Höyük study &lt;a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB83032">PRJEB83032&lt;/a>.&lt;/p>
&lt;p>Ancient DNA needs a different alignment strategy from ordinary modern whole-genome data. The molecules are short, the ends may carry post-mortem damage, and paired reads often overlap because the DNA insert is shorter than the sequencing cycles. For this sample I therefore clean poly-G tails, trim adapters, merge overlapping paired-end reads, align the merged molecules with &lt;code>bwa aln&lt;/code>, remove low-confidence alignments, and deduplicate using both observed ends of each molecule.&lt;/p></description></item><item><title>How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)</title><link>https://popgenetics.dev/posts/awk-subset-populations-genetics/</link><pubDate>Mon, 10 Nov 2025 15:25:02 +0100</pubDate><guid>https://popgenetics.dev/posts/awk-subset-populations-genetics/</guid><description>&lt;p>In an earlier post, &lt;a href="https://popgenetics.dev/posts/plink-pca-tutorial/">PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)&lt;/a>, I showed the manual way to build a subset from the &lt;code>.ind/.fam&lt;/code>. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using &lt;code>awk&lt;/code> that generates a PLINK &lt;code>--keep&lt;/code> file automatically from a list of populations.&lt;/p>
&lt;hr>
&lt;ol>
&lt;li>Prepare a list of populations to keep:&lt;/li>
&lt;/ol>
&lt;p>Create a text file (e.g. &lt;code>pops&lt;/code>) in the same directory as your reference &lt;code>.ind&lt;/code> and &lt;code>.fam&lt;/code>. Put one population label per line:&lt;/p></description></item><item><title>Converting EIGENSTRAT/PACKEDANCESTRYMAP to PACKEDPED</title><link>https://popgenetics.dev/posts/convert-eigenstrat-to-packedped/</link><pubDate>Tue, 29 Jul 2025 15:30:00 +0000</pubDate><guid>https://popgenetics.dev/posts/convert-eigenstrat-to-packedped/</guid><description>&lt;p>The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style &lt;code>.geno/.snp/.ind&lt;/code> dataset. This naming can be confusing: the &lt;code>.snp&lt;/code> and &lt;code>.ind&lt;/code> files are the usual EIGENSTRAT metadata files, but the &lt;code>.geno&lt;/code> file may either be plain-text EIGENSTRAT or binary PACKEDANCESTRYMAP.&lt;/p>
&lt;p>PACKEDPED format allows for easier downstream processing using the &lt;strong>PLINK&lt;/strong> toolset. With PLINK, it becomes straightforward to extract sample subsets, filter SNPs, and perform a wide range of analyses.&lt;/p></description></item><item><title>How to Download the AADR Dataset (Linux &amp; WSL)</title><link>https://popgenetics.dev/posts/download-ancient-modern-dna-aadr/</link><pubDate>Tue, 29 Jul 2025 15:00:00 +0000</pubDate><guid>https://popgenetics.dev/posts/download-ancient-modern-dna-aadr/</guid><description>Find out how to download the v62.0 AADR dataset using wget. Commands for .geno, .snp, and .ind files specifically for ancient DNA research.</description></item></channel></rss>