Wednesday, September 28, 2022
Exploring the sources of the 'Southern ancestry' in the Steppe
Friday, May 6, 2022
New 7300 year old samples from the Middle Don region in Europe & the oldest sample yet from SC Asia
New Preprint:
Population Genomics of Stone Age Eurasia. bioRxiv. Published online 2022. doi:10.1101/2022.05.04.490594
They found some new cool samples from 5300 bce. These are samples from Middle Don region with CHG ancestry. These should be the ancestors of the Khvalynsk samples who show no IranN/IndiaN as per my modeling.
From the preprint:
Interestingly, two herein reported ~7,300-year-old imputed genomes from the Middle Don River region in the Pontic-Caspian steppe (Golubaya Krinitsa, NEO113 & NEO212) derive ~20-30% of their ancestry from a source cluster of hunter-gatherers from the Caucasus (Caucasus_13000BP_10000BP) (Fig. 3). Additional lower coverage (non imputed) genomes from the same site project in the same PCA space (Fig. 1D), shifted away from the European hunter-gatherer cline towards Iran and the Caucasus. Our results thus document genetic contact between populations from the Caucasus and the Steppe region as early as 7,300 years ago, providing documentation of continuous admixture prior to the advent of later nomadic Steppe cultures, in contrast to recent hypotheses, and also further to the west than previously reported.
We demonstrate that this “steppe” ancestry (Steppe_5000BP_4300BP) can be modelled as a mixture of ~65% ancestry related to herein reported hunter-gatherer genomes from the Middle Don River region (MiddleDon_7500BP) and ~35% ancestry related to hunter-gatherers from Caucasus (Caucasus_13000BP_10000BP) (Extended Data Fig. 4). Thus, Middle Don hunter-gatherers, who already carry ancestry related to Caucasus hunter-gatherers (Fig. 2), serve as a hitherto unknown proximal source for the majority ancestry contribution into Yamnaya genomes.
Now this 35% CHG/IranN which came in later, this should contain the IndiaN/IranN that I have written about at length in previous posts.
Let me get my hands on the geno files, I should have a lot of fun working with these samples.
UPDATE:
The paper has 3 new samples from the neolithic Western Iranian site of Tepe Guran dated to around 7000BCE. Should be similar to the Ganj Dareh samples. The male sample is J2a.
Most importantly, they have a 4600 BCE sample from the south Turkmenistan site of Monjukli-Depe.
Y-HG is L1a. The paper says this
"The Neolithic individual from Turkmenistan (~6,500 BP) clusters close to Neolithic Iranians."
So I think we will finally have a sample with the IndiaN ancestry rather than IranN. IranN is the western version, IndiaN being the eastern one. This ancestry gets diluted due to later influx from West Asia as seen in the later AnatoliaN heavy Geoksyur, Anau and Namazga samples.
Thursday, January 20, 2022
Saturday, November 27, 2021
A Reply to the Anthrogenica clique's criticism of my Steppe Eneolithic Post
I recently wrote a post in which my analysis showed that The same ancestors which provided iranian like ancestry to Irula tribals also provided ancestry to Steppe Eneolithic as well as South Central Asia (Sarazm aDna etc). One can read the post here.
For feedback as well as spreading the post, I posted the link to a popular DNA & population genomics forum - Anthrogenica. Quite unexpectedly, immediately after, I was suspended from Anthrogenica with no reason or message whatsoever. I always knew it is a Kurganist bastion, but never quite expected this level of censorship to opposing ideas. I am glad though, starting this blog is now worth it. Also, credit to Davidski at Eurogenes blog for allowing me on his blog comments section even with my opposing views.
I came to know that the Moderator who banned me is a handle named Coldmountains, a handle who I have in the past ridiculed quite a lot (in Eurogenes comment section) for not finding a single R-L657 indian y haplogroup in the steppe since 2015. Poor guy comes empty handed after each successive paper when new samples from the steppe are published. His search still goes on. Meanwhile the only L657+ sample we have so far in aDna is from Roopkund lake India 800CE.
The link to my Anthrogenica thread is here. Please register and show Anthrogenica some love in this thread and elsewhere. The moderator clique there is in an echochamber and needs some awakening.
Anyway, let move on to the criticism of my post. There is just one, and sadly i couldn't reply because I was banned. Hence this post.
Kale on 25-Nov-2021 wrote
Kotias is a pseudo-haploid sample > That means rather than having two different sets of chromosomes like a real person, it is treated as having two exactly identical sets > That means the drift going to itself it going to be crazy high > If you have an edge coming out of an artificially crazy high drift, the percentage contribution has to be artificially crazy small to avoid overfitting.
This graph is completely uninformative until structured properly.
Kale is absolutely wrong here. The pseudo-haploid* samples do not cause artificial high drift edges, rather, the artificially high drift is due to just 1 sample in the label because of which heterozygosity cannot be computed for the label. This problem is solved by using 2 samples in the label even if samples are pseudo-haploid. This is not a problem for .DG samples as these are diploid genomes and allow for heterozygous calls.
This is exactly what I have done in the graph below (later). I lumped Satsurblia & Kotias into 1 label known as CHG. I will show that my conclusion does not change.
Proof of my claim is from the programmers of Admixtools in their qpGraph readme pasted below. Should have been basic reading right?
Genotypes are expected to be pseudo-haploid -- 2 samples at least per population or drift lengths on leaves are not meaningful.
As far as edges coming out of artificially high drifts are concerned, sister clades of Kotias also did not help Kales case. See, i spent weeks on the model trying every possibility. Them not being able to read the graph is not my problem.
Below I will paste my new qpGraph for Steppe Eneolithic which follows these principles and should be acceptable to the Kurganists as well.
- Worst residual ZScore below 3.
- CHG label now has 2 samples and therefore allows for heterozygous calls.
- Each admixture node has a drift edge following it rather than an immediate admixture edge. 1 or 2 admixture edges in my graph follow an immediate admixture node because the drift edge length was 0 (hence i omitted them)
- Non 0 drift edges implying that the edge is a true one. (This is not a strict need. qpGraph disallows immediate admixture edge if the admixed node is labeled as a number. But qpGraph allows the admixture node to be a source to another node if the label given to it is alphanumeric. This is useful if multiple admixtures together are to be modeled.)
Please click on the graph for high res mobile view. On Desktop download image and zoom in a zoomable picture viewer.
DISCUSSION
After correcting all criticisms, the need for IndiaN component in Steppe eneolithic does not go away. I again prove that the same ancestors who ultimately provided ancestry to Steppe Eneolithic in 5th mil BCE also provided ancestry to Irula tribe (and by extension most of indian groups). The minute criticisms which Kurganists come up with are immaterial now, because of course they will come up with them. So far, they have been busy denying even Iranian inflow into steppe (its a mater of purity of course!), so to accept South/ SC Asian origin is a different matter altogether.
