<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Ancestry Modeling on PopGenetics Blog</title><link>https://popgenetics.dev/topics/ancestry-modeling/</link><description>Recent content in Ancestry Modeling on PopGenetics Blog</description><generator>Hugo -- 0.148.2</generator><language>en-us</language><lastBuildDate>Sat, 15 Aug 2026 11:44:17 +0200</lastBuildDate><atom:link href="https://popgenetics.dev/topics/ancestry-modeling/index.xml" rel="self" type="application/rss+xml"/><item><title>F4Mix: Sample-Wise Ancestry Fitting with f4 Statistics</title><link>https://popgenetics.dev/posts/f4mix/</link><pubDate>Sat, 15 Aug 2026 11:44:17 +0200</pubDate><guid>https://popgenetics.dev/posts/f4mix/</guid><description>&lt;p>Last week I published &lt;a href="https://github.com/system0x7/f4mix">F4Mix&lt;/a>, a tool for fitting modern and ancient DNA samples against a pool of source populations, usually ancient ones. F4Mix estimates, for each target, the non-negative mixture of reference populations whose covariance-aware f4 profile best matches it. This makes it useful for testing every sample against the same sources.&lt;/p>
&lt;p>With a proper setup, the tool gives meaningful results, and can reveal both substructure and clear outliers within a site.&lt;/p></description></item><item><title>Are Higher qpAdm P-Values Better?</title><link>https://popgenetics.dev/posts/higher-p-values-better-qpadm/</link><pubDate>Thu, 30 Jul 2026 17:40:54 +0200</pubDate><guid>https://popgenetics.dev/posts/higher-p-values-better-qpadm/</guid><description>&lt;p>Yes. For two qpAdm models with the same target, the same right groups, and the same settings, the model with the higher p-value is the better statistical fit.&lt;/p>
&lt;p>qpAdm calculates a covariance-weighted discrepancy between the observed and fitted f4-statistics. The p-value reflects how well the model explains the used f4-statistics. A higher p-value means the discrepancy between the observed and fitted
values is less unusual under the model.&lt;/p>
&lt;p>This does not mean that the model with the highest p-value for a target is automatically the best one, because qpAdm results depend on the selected right groups. Uninformative right-groups can lack the power to detect a bad model, while overly restrictive ones can make a plausible model appear to fit badly. Therefore, p-values are more comparable when models for the same target are compared using the same groups and settings. Models with different numbers of sources are also comparable since p-values account for different degrees of freedom. Z-scores can be used to assess whether an additional source is justified.&lt;/p></description></item><item><title>Why the Best PCA Fit May Still Be the Wrong Admixture Model</title><link>https://popgenetics.dev/posts/best-pca-fit-wrong-admixture-model/</link><pubDate>Wed, 22 Jul 2026 11:31:54 +0200</pubDate><guid>https://popgenetics.dev/posts/best-pca-fit-wrong-admixture-model/</guid><description>&lt;p>A Vahaduo generated PCA model for Sardinians gives:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>82.8% Barcin Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>11.6% Loschbour
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>5.6% Yamnaya
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>Distance: 3.4303%
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Ganj Dareh was included in the sources but gets a weight of zero. This seems to imply that Sardinians don&amp;rsquo;t have any eastern-Farmer related ancestry.&lt;/p>
&lt;p>When Sardinians are modelled with qpAdm using Barcin Neolithic, Loschbour, Yamnaya, and Ganj Dareh, the model fits well:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>68.6% Barcin Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>11.9% Loschbour
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>10.2% Yamnaya
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>9.4% Ganj Dareh Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>p = 0.769
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>When Ganj Dareh is dropped, the model fails (&lt;span class="katex">&lt;span class="katex-mathml">&lt;math xmlns="http://www.w3.org/1998/Math/MathML">&lt;semantics>&lt;mrow>&lt;mi>p&lt;/mi>&lt;mo>=&lt;/mo>&lt;mn>1.18&lt;/mn>&lt;mo>×&lt;/mo>&lt;msup>&lt;mn>10&lt;/mn>&lt;mrow>&lt;mo>−&lt;/mo>&lt;mn>12&lt;/mn>&lt;/mrow>&lt;/msup>&lt;/mrow>&lt;annotation encoding="application/x-tex">p = 1.18 \times 10^{-12}&lt;/annotation>&lt;/semantics>&lt;/math>&lt;/span>&lt;span class="katex-html" aria-hidden="true">&lt;span class="base">&lt;span class="strut" style="height:0.625em;vertical-align:-0.1944em;">&lt;/span>&lt;span class="mord mathnormal">p&lt;/span>&lt;span class="mspace" style="margin-right:0.2778em;">&lt;/span>&lt;span class="mrel">=&lt;/span>&lt;span class="mspace" style="margin-right:0.2778em;">&lt;/span>&lt;/span>&lt;span class="base">&lt;span class="strut" style="height:0.7278em;vertical-align:-0.0833em;">&lt;/span>&lt;span class="mord">1.18&lt;/span>&lt;span class="mspace" style="margin-right:0.2222em;">&lt;/span>&lt;span class="mbin">×&lt;/span>&lt;span class="mspace" style="margin-right:0.2222em;">&lt;/span>&lt;/span>&lt;span class="base">&lt;span class="strut" style="height:0.8141em;">&lt;/span>&lt;span class="mord">1&lt;/span>&lt;span class="mord">&lt;span class="mord">0&lt;/span>&lt;span class="msupsub">&lt;span class="vlist-t">&lt;span class="vlist-r">&lt;span class="vlist" style="height:0.8141em;">&lt;span style="top:-3.063em;margin-right:0.05em;">&lt;span class="pstrut" style="height:2.7em;">&lt;/span>&lt;span class="sizing reset-size6 size3 mtight">&lt;span class="mord mtight">&lt;span class="mord mtight">−&lt;/span>&lt;span class="mord mtight">12&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>).&lt;/p></description></item><item><title>Introducing AdmixPy: f-statistics, qpAdm, and qpWave in Python</title><link>https://popgenetics.dev/posts/admixpy/</link><pubDate>Thu, 21 May 2026 19:05:53 +0200</pubDate><guid>https://popgenetics.dev/posts/admixpy/</guid><description>&lt;p>I recently published &lt;a href="https://github.com/system0x7/admixpy">AdmixPy&lt;/a> on GitHub, a fast implementation of f-statistics, qpAdm, and qpWave in Python that runs on Linux, macOS, and Windows. It works directly on the new AADR TGENO distribution format and is notably faster than ADMIXTOOLS 2 and simpler to set up. Supported input formats: EIGENSTRAT (&lt;code>.geno/.snp/.ind&lt;/code>), packed AncestryMap (&lt;code>.geno/.snp/.ind&lt;/code>), TGENO (&lt;code>.tgeno/.snp/.ind&lt;/code>), and SNP-major PLINK binary (&lt;code>.bed/.bim/.fam&lt;/code>).&lt;/p>
&lt;p>AdmixPy is implemented in Python and depends only on NumPy, SciPy, and pandas. Installation is handled through pip, and it should behave the same on every platform.&lt;/p></description></item><item><title>Fast, Transparent f4-Based Admixture Screening in R</title><link>https://popgenetics.dev/posts/deterministic-f4-solver/</link><pubDate>Tue, 03 Feb 2026 18:45:43 +0100</pubDate><guid>https://popgenetics.dev/posts/deterministic-f4-solver/</guid><description>&lt;p>In this post, I will build a transparent admixture-screening workflow from scratch in R using f4-statistics and constrained regression. The main advantage is automation: instead of hand-writing every candidate model, the script tests many 2-way, 3-way, and 4-way source combinations in one pass and ranks them by fit. ADMIXTOOLS 2 already includes batch tools such as &lt;code>qpadm_multi()&lt;/code> and &lt;code>qpadm_rotate()&lt;/code>, so the point is not that qpAdm cannot be automated. The point is that this custom workflow is compact, transparent, easy to modify, and useful for exploratory model search before you validate the strongest candidates more formally.&lt;/p></description></item><item><title>Running qpAdm with ADMIXTOOLS2 in R: Testing and Interpreting Ancestry Models</title><link>https://popgenetics.dev/posts/qpadm-tutorial-admixture-modeling/</link><pubDate>Tue, 16 Dec 2025 14:37:15 +0100</pubDate><guid>https://popgenetics.dev/posts/qpadm-tutorial-admixture-modeling/</guid><description>&lt;p>This post covers using qpAdm in R to test ancestry models and estimate admixture proportions. qpAdm builds on f4-statistics and provides a framework for evaluating whether proposed source populations can explain a target population&amp;rsquo;s genetic makeup.&lt;/p>
&lt;p>For R and admixtools setup instructions on Debian/Ubuntu, see my previous post: &lt;a href="https://popgenetics.dev/posts/interpreting-f4-tests-admixtools/">Running f4-Statistics with Admixtools in R&lt;/a>. Windows users can find R installation instructions on the R website.&lt;/p>
&lt;hr>
&lt;h2 id="what-is-qpadm">What is qpAdm?&lt;/h2>
&lt;p>qpAdm is a method for testing ancestry models and estimating admixture proportions. It determines whether a target population can be modeled as a mixture of specified source populations (&amp;ldquo;left populations&amp;rdquo;), and if the model fits, calculates the contribution from each source. The method builds on f4-statistics (covered in my previous post) to evaluate these ancestry models.&lt;/p></description></item><item><title>Running ADMIXTURE in Supervised Mode</title><link>https://popgenetics.dev/posts/admixture-supervised-tutorial/</link><pubDate>Tue, 05 Aug 2025 13:21:16 +0200</pubDate><guid>https://popgenetics.dev/posts/admixture-supervised-tutorial/</guid><description>&lt;p>This post is a short follow-up to the previous one on &lt;a href="https://popgenetics.dev/posts/admixture-unsupervised/">Estimating Ancestry Components Using ADMIXTURE&lt;/a>. Here, we’ll explore supervised ADMIXTURE, a mode that allows you to explicitly define ancestral populations and infer the ancestry proportions of unassigned individuals based on those references.&lt;/p>
&lt;hr>
&lt;h3 id="what-is-a-supervised-run">What Is a Supervised Run?&lt;/h3>
&lt;p>In supervised mode, ADMIXTURE skips the component discovery step and instead uses &lt;strong>user-defined groupings&lt;/strong> to represent ancestral components. The benefit: if you already have solid candidates for reference populations, you can use them to quickly infer ancestry proportions for target or admixed individuals.&lt;/p></description></item><item><title>How to Run ADMIXTURE (Unsupervised): Full Tutorial &amp; Python Plotting Script</title><link>https://popgenetics.dev/posts/admixture-unsupervised/</link><pubDate>Sat, 02 Aug 2025 22:47:12 +0200</pubDate><guid>https://popgenetics.dev/posts/admixture-unsupervised/</guid><description>&lt;p>In this post, I’ll demonstrate how to estimate ancestry proportions using one of the most widely used tools in population genetics: &lt;a href="https://dalexander.github.io/admixture/download.html">ADMIXTURE&lt;/a>. ADMIXTURE is a model-based clustering algorithm that estimates individual ancestry proportions and ancestral allele frequencies from multilocus SNP genotypes.&lt;/p>
&lt;hr>
&lt;h3 id="preparing-the-dataset">Preparing the Dataset&lt;/h3>
&lt;p>Download the appropriate ADMIXTURE binary and either place it in your dataset directory or make it globally accessible. For this run, I included a subset of West Asian populations along with a few adjacent populations (around 150 samples in total). Linkage Disequilibrium (LD) pruning was applied beforehand. If you&amp;rsquo;re unsure how to prune your dataset, refer to the previous post.&lt;/p></description></item></channel></rss>