Saturday, September 9, 2023

Genomic analysis of Seleucid and Parthian period population in North Iran

Iranic Genomes Project posted this abstract on Twitter. I need to find out the origin, If you find the link please post it in the comments.

11 ancient individuals from the Seleucid-Parthian era (~300 BCE - 200 CE) from North Iran (Mazandaran, Gilan, Semnan provinces)

Site location
: Seleucid-Parthian new sites 🟣: Tepe Hissar Chalcolithic-Bronze age
Base Map Credit: Wikipedia (CC BY-SA 4.0)

 

Saturday, July 29, 2023

The Hybrid Model for Indo-European languages

New paper - Paul Heggarty et al. ,Language trees with sampled ancestors support a hybrid model for the origin of Indo-European languages.Science381,eabg0818(2023).DOI:10.1126/science.abg0818

*Note, in the paper the authors use dates in BP (before present) where present = 2000 CE. While I appreciate the strictly secular nomenclature, I prefer BCE dates as they are easier to comprehend for my brain (for now). So I convert these into BCE by subtracting 2000. eg. 6000 BP = 4000 BCE. Do keep this in mind while reading. Thank you.

Brief Overview of the Method


The authors aver that this method used Bayesian phylogenetic inference which is not similar to either Lexicostatistics or Glottochronology, both of which they consider deeply flawed.

This paper's Bayesian phylogenetic inference analysis is based on a new improved database (IE - CoR 1.0). The IE‑CoR 1.0 database contains data on relationships of cognacy (shared word origin) between 161 Indo-European languages in a reference set of 170 basic meanings. The new languages include the Nuristani branch, extinct Iranic languages from central Asia and a representative of sub-branches of Celtic which was missing from previous databases (Gaulish). The coverage prioritizes non-modern languages, providing a deeper phylogenetic signal and better chronological estimation. This database was contributed by 80 experts of different language sub-families to maximize data accuracy.

The authors state they improved the cognate encoding (keeping 1 lexeme for each cognate set rather than many synonyms used in previous databases which created lots of cognate sets per lexeme. This, for example, artificially elongated the branch length of modern Greek and the age of old Greek). The IE-CoR data set has highly consistent counts of cognate sets across all languages, very close to the target of 1 cognate set per meaning, per language. They also removed the constraints previously placed on ancient languages to be directly ancestral to modern languages which need not be the case. This previously forced 0 branch length (and therefore no divergence), simply forced the changes onto the next branch and elongated branch lengths artificially.

The database also solves the loanword problem in computational cladistics. "IE-CoR introduced the concept of loanword event, through which it has become possible to encode correctly both non-cognacy to the source lexeme, and subsequent cognacy between vertical descendants of that lexeme, once borrowed and fully integrated into the borrower language."

The IE-Cor database can be found here https://iecor.clld.org/


Important Discussion and Conclusions from the Paper


Heggarty et al reaffirm the position of the earliest Indo-European speakers in the south of the Caucasus around ~6100 BCE. They support a hybrid model in which the steppe was a secondary staging ground for European languages. Notably, the beginning of the split from Indo-Iranian into Indo-Aryan and Iranic is dated to ~3500 BCE, a finding wholly incompatible with the Andronovo hypothesis.

DensiTree showing final IE Tree with probability of topologies
DensiTree final output of the paper shows the probability distribution of various topologies. Orphan branches are sampled ancient languages in the database (some examples in red and yellow box markers)


Saturday, February 18, 2023

TTK001 from Tutkaul, Tajikistan

Yu, He, 2022, "Paleogenomics of Upper Paleolithic to Neolithic European hunter-gatherers", https://doi.org/10.17617/3.Y1KJMF, Edmond, V2

The dataset for the above paper is available, but the paper/preprint is not. One of the samples is TTK001

Tutkaul map
Tutkaul, Sarazm, Khvalynsk locations


qpAdm on this sample revealed that they possessed mainly ANE ancestry with additional Iran Neolithic component.

Target: TTK001
Russia_Kolyma_M.SG: 2.7 ± 2.7%
Iran_GanjDareh_N: 17.8 ± 3.1%
Russia_AfontovaGora3: 79.5 ± 3.2%
p-value: 0.065

Thursday, December 22, 2022

Genetic History of the Tajiks

12 metre Buddha in Nirvana: Ajina Tepe, Tajikistan 6th - 7th Century CE)


I was going through the paper by Guarino-Vignon et al (2022) again and was struck by some very good insights. The paper is titled "Genetic continuity of Indo-Iranian speakers since the Iron Age in southern Central Asia".

The paper studies the modern Tajik and Yaghnobi people of Tajikistan. While the Tajik speak a Persian dialect (Iranian > West Iranian > SW Iranian > Persian > Tajik), the Yaghnobis speak an Eastern Iranian language, a descendant of ancient Sogdian (Iranian > East Iranian > Sogdian > Yaghnobi).

Sunday, December 18, 2022

Steppe Ancestry definitely arrived in India post 1000 BCE - The Final Blow


Karna
कर्ण (karNa), by @zdrava



In my first post on this topic - 'The True source of steppe ancestry in modern Indians' - I laid out the genetic claim that the best source for steppe in Indians seems to be a sample from the Yaz II Iron Age culture site at Takhirbai 3, Turkmenistan (short name of the sample - TKM_IA), dated to approximately 850 BCE. Based on other archaeological, anthropological, epigraphic and literary evidence which also supported this claim, I proposed that only a post-1000 BCE movement of these 'proto-Śāka' people could explain the steppe ancestry present in modern Indians. The main reason to reject the view that steppe folks came to the core Vedic region (of Punjab, Haryana and East UP) around ~1500 BCE, and gave Indians the Vedic culture and Indo-European languages is the fact that absolutely no archaeological evidence exists to support such a big claim. This difference is important - a post-1000 BCE movement of Iranian-speaking proto-Śāka into India cannot bring the Vedic culture and language. Neither does it match the evidence from Rig Veda, which describes a pre-Iron Age life. Do read that article to understand all the evidence in support of my claim. 

Friday, December 16, 2022

The true source of the Steppe ancestry in modern Indians (continued)




In my previous post, I concluded that the steppe ancestry in Indians is most likely to have arrived after 1000 BCE via the East Iranian-speaking Śāka. the ~850 CE TKM_IA sample turned out to be the best steppe source among the options provided.

Since then, a particular commentator on Twitter, who is adamant about the ~1500 BCE invisible steppe invasion into India has been putting alternative viewpoints in support of the old theory favoured by Kurganists. 

In his first attempt, he argued for an Inner Asian Mountain Corridor (IAMC) specific ancestry (Aigyrzhal_BA from Kyrgystan 2000 BCE as the proxy).


Saturday, December 3, 2022

The true source of the steppe ancestry in modern Indians

 

Mauryan figurines

This post aims to clarify the source of the bronze age steppe ancestry in modern Indians. For this, I have chosen the following targets.

Target list and details
Modern targets and their details


Monday, November 28, 2022

Mitanni at Hasanlu - Did they have Sintashta ancestry or not?

I want to make a quick post on a new article by Nezih Seven. You can find it here.

A Genetic Analysis of Historical Population Movements Around The Zagros Mountains


He analyzed the NW Iranian samples from the 'Southern Arc' paper and reached some interesting conclusions.

Nezih model hasanlu_lba 85% bmac + 15% sintashta
Model for Hasanlu_LBA_A by Nezih


The p-value of the above model is ~0.20 (therefore passing as p>0.05). His qpAdm output file can be seen here

Saturday, November 5, 2022

Ancestry trends of 150+ groups from the Indian subcontinent



In this article, I will lay out general ancestry trends of 150+ groups from the Indian subcontinent. And try to make it make sense to a lay audience.

Before I present the data table, there are some important caveats:

1. The purpose of the table is to concurrently compare the ancestry trends for 100+ groups of the Indian subcontinent. If the chosen sources are incorrect, they will be incorrect for all the groups but will still allow us to compare ancestry % across groups and make some conclusions. The fit (distance %) is not too relevant, and the distances of some groups will indeed be bad (>3%). However, that does not take away from our purpose of finding broad trends within the data.

Tuesday, November 1, 2022

Did Y haplogroup R1a-Y3 go into hiding?

 

As per the Harvard database, there are 454 male samples dated between 3000BCE and 0CE in all the *stan countries north of India and Russia combined. Of these, 166 are R1a but none is R1a-Y2 or L657+ which is supposed to have been formed around 2600BCE from Steppe R1a-Z94 > Y3.

In Narasimhan et al 2019 supplement too, there are 250 male samples analyzed between the dates 3000BCE - 0 CE, and none of them is R1a-Y3+. These include all the Sintashta, Andronovo, Saka and descended culture samples, all of these have significant steppe_mlba admixture.

This is the modern country-wise frequency of R-Y3+ lineages as per YFull (23 countries, 13250 samples). As you can see, only the Indian subcontinent sees R1a-Y3+ samples, with an average of ~15-20% of all males sampled in Pakistan, India, Nepal, Bangladesh and Sri Lanka. The only thing common between these and the steppe R1a is a common ancestor who lived around 2600BCE. But if his descendant Y3 was born in the steppe, we should have seen much more Y3 in the ancient samples and in the modern distribution in the countries north of India. So where are they?

R1a-Y3 modern map


Friday, October 14, 2022

R1a Explained

M780 r1a distribution


Y Haplogroups - A Primer


Y Chromosome haplogroup is a defined set of mutations in the (non-recombinant) portion of the DNA from male specific Y-chromosome. Y-Chr is passed down from father to son along with all the mutations that have been accumulated till that point. All the male descendants of a man will share the mutations in the Y-Chr of that man, plus all the additional mutations that have been collected in the generations between that man and the descendant on his specific male ancestral lineage chain. This feature helps us in identifying the most recent common paternal ancestor of a group of men and is a great tool for population geneticists. Do note that the same set of additional mutations cannot be created in two different men at the same time, the probability of that happening randomly is 0.

mtDNA haplogroups are mutations found in the mitochondrial DNA. Unlike autosomal chromosomes and Y chromosomes, mtDNA is found outside the cell nucleus. mtDNA is passed down from mother to children only, and therefore is informative only of the maternal lineage of a person. 

In this article, we will discuss Y haplogroups which can only inform us of the paternal ancestry of males.


Wednesday, September 28, 2022

Exploring the sources of the 'Southern ancestry' in the Steppe



The proximal source of the southern (IranN/CHG) ancestry in the Eneolithic Steppe as well as Yamnaya populations has been a mystery. Chintalapati et al 2022 infer that the admixture between Eastern European Hunter-Gatherers and Iran Neolithic ancestry occurred between 4400-4000BCE, using the DATES algorithm developed by Moorjani Lab.

In this post, powered with newly published samples as well as better tools than before, I will attempt to decipher the true sources of these populations.

Disclaimer: The analysis presented below is based on the currently available published samples from the Pontic Caspian Steppe. The samples from the Middle Don region (Allentoft et al 2022 preprint) are unavailable, and this analysis will be updated once those samples become available.

Monday, September 19, 2022

The Lady from Central Asia who was found dead in a Bronze Age Turkish well

 

Site of the Alalakh 'Well Lady'. Courtesy: Skourtanioti et al 2020


The remains of the Alalakh 'Well Lady' were discovered in the archaeological site of Tell Atchana / Alalakh / Alalah, Turkey dating to ~1550BCE. Her aDna was published in 2020 (1) and she was found to be of central Asian origin, far from where she was found. She was found at the bottom of a deep well that was still in use at the time. The individual showed evidence of healed trauma on the skull's frontal bone and two healed fractured ribs, and her manner of death has been suggested as a homicide due to or before being thrown into the well (2). Her teeth also showed clear signs of enamel defects (2), a possible sign of malnutrition during childhood. Her age at death was 40-45 years.

Friday, September 9, 2022

Map of ancestry distribution in West to SC Asia during the Neolithic




All the relevant populations from the early neolithic to chalcolithic are represented here. Most up to date map of population movements in the neolithic age, including the latest neolithic samples from Lazaridis et al 2022. All results have been arrived at using rigorous qpAdm rotating models.

If alternate model exists, I have noted them and will be visible once you hover on a label.

To pan the map, click the play button on the map and then and select the + button or Cross which enables the Pan cursor. Map can then be panned in all 4 directions. Pinch in/out to zoom. There is also a date filter.

Latitude/longitude of some labels might not be accurate. Co-ordinates of some labels have been shifted a bit so that there is less overlap between labels close to each other in location (eg. Seh_Gabi_LN, Seh_Gabi_C, Ganj_Dareh, Hajji_Firuz cluster; Arm_Masis_N, Arm_Aknashen, Aze_Lowlands cluster; Geoksyur & Gonur cluster; Boncuklu & Pinarbasi)


Wednesday, August 31, 2022

Armenia vs NW Iran - Two different short stories

 The Armenian Data from the new papers


1. Some EHG ancestry from north of caucasus is seen at Areni Cave, 4200BCE.

2. EHG dilutes by the time of Kura Araxes 3600-2000BCE, Arm_EBA (green below).

3. Arm_MBA/LBA/IA cluster clearly shifts towards steppe, with lots of R1b & I2 Y hg. According to Lazaridis et al, this explains the Armenian language branch. For now I will accept this conclusion as it makes sense, and explains the different language of NW Iran.

4. Post IA, intense mixing with more southwestern populations (but not NW Iran) so much so that all steppe autosomal ancestry gets diluted.


Armenia PCA
PCA - Ancient & Modern Armenia

Friday, August 26, 2022

The Southern Arc paper and it's data upends the Steppe Theory in multiple ways


Hasanlu Indra chariot



Lazaridis, Iosif, et al. “The Genetic History of the Southern Arc: A Bridge between West Asia and Europe.” Science, vol. 377, no. 6609, 2022, https://doi.org/10.1126/science.abm4247.


This is a recently published paper ('the paper', 'this paper') with many new samples from what the authors call 'The Southern Arc'. The authors define it as 'a region centred on the large Anatolian peninsula (Turkey), including in the west (in Europe) the Balkans and the Aegean, and in the south and east, Cyprus, Mesopotamia, the Levant, Armenia, Azerbaijan, and Iran.'

The main conclusion of this paper which upends one part of the Steppe theory is that West Asia is now considered to be the homeland of the first Indo-European speakers, and the steppe to be only a secondary homeland. From the paper:

Thursday, August 18, 2022

A Relook at the Rakhigarhi ancestor I6113


In my previous post, I analyzed the so-called 'Indus Periphery' samples from the data published in Narasimhan et al 2019. 

My conclusion was that the IVCp samples can be best modeled distally as 

Ganj_Dareh + Ancient North Eurasian (Tarim_Basin or West Siberian HG) + Onge + Levant_PPN or Anatolian Farmer.

IVC periphery models

In this post, I shall do an in depth analysis of the genome of the lone Rakhigarhi sample (Id: I6113, female) dated to around 2000 BCE, published in Shinde et al 2019. 

Burial picture of Rakhigarhi woman I6113
Picture of I6113 burial, Rakhigarhi. From Shinde et al 2019


Sunday, August 14, 2022

Update: IVC & Swat Valley Genetics, Bonus - Kashmiri Pandit Ancestry

SECTION A: THE INDUS PERIPHERY SAMPLES


In the Narasimhan et al 2019 paper, 13 outlier samples were published - 10 from the Iranian site of Shahr-i-Sokhta (abbrv. SiS, various dates 3100-2000 BCE) and 3 from the Bactria Margiana Archaeological Complex (abbrv. BMAC) site of Gonur (2400-2000 BCE). These samples showed an elevated ancestry component found in the Onge Andamanese and modern Indians, which was not found in the main samples from SiS and BMAC. Based on this, the hypothesis was made that these people were migrants from one or more sites of the Indus Valley Civilization (abbrv. IVC) and were labeled as 'Indus Periphery Samples'.

Various Site Locations


Sunday, June 19, 2022

Marija Gimbutas's Kurgan Hypothesis rejected by Harvard

 

The Genetic History of the Southern Arc: A Bridge between West Asia & Europe

Below are some excerpts from the notes of a talk which will be delivered by Dr David Reich at an Israeli university in July 2022.

Our comprehensive sampling shows that Anatolia received hardly any genetic input from Europe or the Eurasian steppe from the Chalcolithic to the Iron Age; this contrasts with Southeastern Europe and Armenia that were impacted by major gene flow from Yamnaya steppe pastoralists.

The impermeability of Anatolia to exogenous migration contrasts with our finding that the Yamnaya had two distinct gene flows, both from West Asia, suggesting that the Indo-Anatolian language family originated in the eastern wing of the Southern Arc and that the steppe served only as a secondary staging area of Indo-European language dispersal. 

Here, Dr Reich confirms what I have already said before - That the Anatolian branch of the IE language family has its origin in Armenia and has nothing to do with Steppes. I believe this is the first time that Harvard will officially debar Steppefrom being a PIE homeland contender. 

Read my post - The Armenian Origin of the Anatolian branch of IE language family

Friday, May 6, 2022

New 7300 year old samples from the Middle Don region in Europe & the oldest sample yet from SC Asia

 

New Preprint: 

Population Genomics of Stone Age Eurasia. bioRxiv. Published online 2022. doi:10.1101/2022.05.04.490594


They found some new cool samples from 5300 bce. These are samples from Middle Don region with CHG ancestry. These should be the ancestors of the Khvalynsk samples who show no IranN/IndiaN as per my modeling. 

From the preprint:

 

Interestingly, two herein reported ~7,300-year-old imputed genomes from the Middle Don River region in the Pontic-Caspian steppe (Golubaya Krinitsa, NEO113 & NEO212) derive ~20-30% of their ancestry from a source cluster of hunter-gatherers from the Caucasus (Caucasus_13000BP_10000BP) (Fig. 3). Additional lower coverage (non imputed) genomes from the same site project in the same PCA space (Fig. 1D), shifted away from the European hunter-gatherer cline towards Iran and the Caucasus. Our results thus document genetic contact between populations from the Caucasus and the Steppe region as early as 7,300 years ago, providing documentation of continuous admixture prior to the advent of later nomadic Steppe cultures, in contrast to recent hypotheses, and also further to the west than previously reported.

We demonstrate that this “steppe” ancestry (Steppe_5000BP_4300BP) can be modelled as a mixture of ~65% ancestry related to herein reported hunter-gatherer genomes from the Middle Don River region (MiddleDon_7500BP) and ~35% ancestry related to hunter-gatherers from Caucasus (Caucasus_13000BP_10000BP) (Extended Data Fig. 4). Thus, Middle Don hunter-gatherers, who already carry ancestry related to Caucasus hunter-gatherers (Fig. 2), serve as a hitherto unknown proximal source for the majority ancestry contribution into Yamnaya genomes.


Now this 35% CHG/IranN which came in later, this should contain the IndiaN/IranN that I have written about at length in previous posts. 

Let me get my hands on the geno files, I should have a lot of fun working with these samples.

UPDATE:

The paper has 3 new samples from the neolithic Western Iranian site of Tepe Guran dated to around 7000BCE. Should be similar to the Ganj Dareh samples. The male sample is J2a.

Most importantly, they have a 4600 BCE sample from the south Turkmenistan site of  Monjukli-Depe.

Y-HG is L1a. The paper says this

"The Neolithic individual from Turkmenistan (~6,500 BP) clusters close to Neolithic Iranians."

So I think we will finally have a sample with the IndiaN ancestry rather than IranN. IranN is the western version, IndiaN being the eastern one. This ancestry gets diluted due to later influx from West Asia as seen in the later AnatoliaN heavy Geoksyur, Anau and Namazga samples.