<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>PCA on PopGenetics Blog</title><link>https://popgenetics.dev/topics/pca/</link><description>Recent content in PCA on PopGenetics Blog</description><generator>Hugo -- 0.148.2</generator><language>en-us</language><lastBuildDate>Wed, 22 Jul 2026 11:31:54 +0200</lastBuildDate><atom:link href="https://popgenetics.dev/topics/pca/index.xml" rel="self" type="application/rss+xml"/><item><title>Why the Best PCA Fit May Still Be the Wrong Admixture Model</title><link>https://popgenetics.dev/posts/best-pca-fit-wrong-admixture-model/</link><pubDate>Wed, 22 Jul 2026 11:31:54 +0200</pubDate><guid>https://popgenetics.dev/posts/best-pca-fit-wrong-admixture-model/</guid><description>&lt;p>A Vahaduo generated PCA model for Sardinians gives:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>82.8% Barcin Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>11.6% Loschbour
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>5.6% Yamnaya
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>Distance: 3.4303%
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Ganj Dareh was included in the sources but gets a weight of zero. This seems to imply that Sardinians don&amp;rsquo;t have any eastern-Farmer related ancestry.&lt;/p>
&lt;p>When Sardinians are modelled with qpAdm using Barcin Neolithic, Loschbour, Yamnaya, and Ganj Dareh, the model fits well:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>68.6% Barcin Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>11.9% Loschbour
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>10.2% Yamnaya
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>9.4% Ganj Dareh Neolithic
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>p = 0.769
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>When Ganj Dareh is dropped, the model fails (&lt;span class="katex">&lt;span class="katex-mathml">&lt;math xmlns="http://www.w3.org/1998/Math/MathML">&lt;semantics>&lt;mrow>&lt;mi>p&lt;/mi>&lt;mo>=&lt;/mo>&lt;mn>1.18&lt;/mn>&lt;mo>×&lt;/mo>&lt;msup>&lt;mn>10&lt;/mn>&lt;mrow>&lt;mo>−&lt;/mo>&lt;mn>12&lt;/mn>&lt;/mrow>&lt;/msup>&lt;/mrow>&lt;annotation encoding="application/x-tex">p = 1.18 \times 10^{-12}&lt;/annotation>&lt;/semantics>&lt;/math>&lt;/span>&lt;span class="katex-html" aria-hidden="true">&lt;span class="base">&lt;span class="strut" style="height:0.625em;vertical-align:-0.1944em;">&lt;/span>&lt;span class="mord mathnormal">p&lt;/span>&lt;span class="mspace" style="margin-right:0.2778em;">&lt;/span>&lt;span class="mrel">=&lt;/span>&lt;span class="mspace" style="margin-right:0.2778em;">&lt;/span>&lt;/span>&lt;span class="base">&lt;span class="strut" style="height:0.7278em;vertical-align:-0.0833em;">&lt;/span>&lt;span class="mord">1.18&lt;/span>&lt;span class="mspace" style="margin-right:0.2222em;">&lt;/span>&lt;span class="mbin">×&lt;/span>&lt;span class="mspace" style="margin-right:0.2222em;">&lt;/span>&lt;/span>&lt;span class="base">&lt;span class="strut" style="height:0.8141em;">&lt;/span>&lt;span class="mord">1&lt;/span>&lt;span class="mord">&lt;span class="mord">0&lt;/span>&lt;span class="msupsub">&lt;span class="vlist-t">&lt;span class="vlist-r">&lt;span class="vlist" style="height:0.8141em;">&lt;span style="top:-3.063em;margin-right:0.05em;">&lt;span class="pstrut" style="height:2.7em;">&lt;/span>&lt;span class="sizing reset-size6 size3 mtight">&lt;span class="mord mtight">&lt;span class="mord mtight">−&lt;/span>&lt;span class="mord mtight">12&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>&lt;/span>).&lt;/p></description></item><item><title>SmartPCA Tutorial: How to Run PCA on Genetic Data</title><link>https://popgenetics.dev/posts/smartpca-tutorial/</link><pubDate>Wed, 30 Jul 2025 20:47:41 +0200</pubDate><guid>https://popgenetics.dev/posts/smartpca-tutorial/</guid><description>&lt;p>This post is a continuation of the previous one, where I demonstrated how to perform PCA with PLINK. While PLINK’s PCA is great for quick, exploratory analysis, smartpca (part of the EIGENSOFT toolset) is particularly common in population-genetic and ancient-DNA studies.&lt;/p>
&lt;p>Smartpca can be compiled from the EIGENSOFT source or installed through conda. I covered the installation process in this earlier post: &lt;a href="https://popgenetics.dev/posts/convert-eigenstrat-to-packedped/">From EIGENSTRAT to PACKEDPED&lt;/a>.&lt;/p>
&lt;p>As before, I’ll use a small subset. The focus here is on the technical process. One key difference in this post is that I’ll perform Linkage Disequilibrium (LD) pruning, which reduces redundancy between correlated SNPs before PCA.&lt;/p></description></item><item><title>PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)</title><link>https://popgenetics.dev/posts/plink-pca-tutorial/</link><pubDate>Tue, 29 Jul 2025 16:00:00 +0000</pubDate><guid>https://popgenetics.dev/posts/plink-pca-tutorial/</guid><description>&lt;p>In this post, I’ll demonstrate how to perform a PCA on a PLINK dataset.
Before we begin, we need to prepare a subset of samples we&amp;rsquo;re interested in analyzing.&lt;/p>
&lt;p>To do this, we’ll extract sample information from the &lt;code>.fam&lt;/code> file.
But first, we need to identify the samples of interest. For example, those from a specific population such as Sardinians.&lt;/p>
&lt;p>The easiest way is to open the corresponding &lt;code>.ind&lt;/code> file and look at the population column, which is the third column in each row. Open the file in a text editor, and search for the population name, in this case, Sardinian.&lt;/p></description></item></channel></rss>