Proto-Indo-European: The Mother of All Languages

Notes — verified against academic sources, with margin annotations

The left column reproduces Navin's original notes. The right column carries verification notes: claims sourced to standard handbooks and clickable online references, flags for debated points, and corrections for outright errors. A short ordered reading list appears at the end.

SUPPORTED well-attested in standard references  DEBATED scholarly disagreement or oversimplified  NUANCE right idea, but worth refining

Abstract / framing

Proto Indo European: The Mother of All Languages.
How do we know so much about a language of which zero records survive? The Latin word for “father” is pater: how do we know both are descended from PIE *ph₂tḗr, and more importantly, why was the original p-based and not f-based? How do we know Sanskrit is descended from PIE and not the other way round?
SUPPORTED The reconstruction *ph₂tḗr “father” is the standard PIE form (laryngeal h₂ conditions the lengthened grade).
Fortson, Indo-European Language and Culture (2nd ed., 2010), §3.36; Mallory & Adams (2006), §13.2.
Online: Wiktionary: *ph₂tḗr; U Texas LRC Indo-European Documentation Center.
SUPPORTED The “why p and not f?” question is the standard motivating puzzle for Grimm's Law — the same one Jacob Grimm formalized in 1822.

How reconstruction works

How to reconstruct

  • There are rules.
  • They go forward.
  • Run them backward: what language could have evolved into both of these given the known rules?

Where do forward rules come from?

  • Which languages are related?
    • Cognate sets — Swadesh list.
  • Find correspondence sets.
  • Posit proto-sounds.
  • Two steps: easy cases first, then hard ones. Majority rules. Occam's Razor.
  • Example cher / caro / caro / caru (“dear” in French, Italian, Spanish, Portuguese)
    • *karo = proto-word by majority rules.
    • 2/4 is a majority if the remaining are different.
    • * indicates reconstructed.
  • When there isn't a majority, look across the dataset for which direction is more consistent.
SUPPORTED This is the textbook comparative method: cognate sets → sound correspondences → reconstructed protoforms, judged by typological plausibility and economy.
Hock, Principles of Historical Linguistics (2nd ed., 1991), chs. 14–16; Campbell, Historical Linguistics: An Introduction (3rd ed., 2013), ch. 5.
Online: Comparative method; U Texas LRC: PIE Phonology.
NUANCE The Swadesh list is one tool (Morris Swadesh, 1950s), but modern comparative linguists generally don't rely on it as the sole basis — it's better suited to lexicostatistics and rough subgrouping than to formal reconstruction.
NUANCE “Majority rules” is a useful first heuristic but it can mislead — e.g. Grimm's Law has all of Germanic going one way and the rest going the other, yet PIE *p (not Germanic f) is reconstructed because the change p→f is typologically common and the reverse rare. Directionality of change matters more than headcount.
Hock (1991), ch. 16 (“Internal reconstruction” and the role of typological plausibility).

Directionality of sound changes

  • Assimilation
  • Degemination
  • Lengthening of vowel
  • Lenition: weakening of a consonant from one that takes more effort to less. Stop → affricate or fricative.
  • Sandhi: conditioned changes at word boundaries. English: Frank is → Frank's.
SUPPORTED All five are textbook categories of sound change. Lenition is well-attested as a strongly directional process (the reverse, fortition, is rarer and usually conditioned).
Campbell (2013), ch. 2; Hock (1991), ch. 5.
Online: Lenition; Sandhi.
NUANCE “Sandhi” is a Sanskrit term (saṁdhi “joining”) used in modern phonology — Pāṇini described it systematically > 2,300 years ago, and the term is used cross-linguistically today.

Where's the proof?

  • Hittite
  • Records
  • Archaeology:
    • Linguistics forced archaeology:
      • Settlement archaeology.
      • Linguistic reconstruction came first; archaeologists then tried to find cultures matching the language tree.
SUPPORTED Hittite (deciphered by Bedřich Hrozný in 1915–1917) was a major proof-test — it pushed the earliest written IE language back to ~1650 BCE and confirmed Saussure's laryngeal predictions a generation later.
Fortson (2010), ch. 9; Watkins, “Hittite” in The Cambridge Encyclopedia of the World's Ancient Languages (Woodard ed., 2004).
Online: Bedřich Hrozný.
SUPPORTED “Linguistic palaeontology” (using reconstructed vocabulary to constrain homeland and date) is exactly the method Anthony, Mallory, and others have used — the wheel/wagon/horse cluster of words is the most famous case.
Anthony, The Horse, the Wheel, and Language (2007), chs. 2–3; Mallory & Adams (2006), ch. 10.

Early signs: Sassetti and Jones

Florence merchant in Goa, 1580s

Filippo Sassetti noticed that Sanskrit sarpa resembled Italian serpe, deva resembled dio, and the numerals six, seven, eight looked nearly identical.

Sir William Jones (1786)

  • What did Jones see?
    • Sanskrit pitṛ / Greek patēr / Latin pater / Old English fæder
    • Sanskrit trayas / Greek treis / Latin trēs / English three
  • Jones: “... common source, which, perhaps, no longer exists.”
SUPPORTED Sassetti (1540–1588), a Florentine merchant in Cochin/Goa, wrote to Bernardo Davanzati in 1585 noting Sanskrit-Italian similarities: deva/dio, sarpa/serpe, sapta/sette, aṣṭa/otto, nava/nove. The letters weren't published in his lifetime.
Karttunen, “Sassetti, Filippo,” Persons of Indian Studies;
Online: Wikipedia: Filippo Sassetti; whowaswho-indology.info.
SUPPORTED Jones's full sentence (Third Anniversary Discourse to the Asiatic Society, 2 February 1786): “... no philologer could examine them all three, without believing them to have sprung from some common source, which, perhaps, no longer exists.”
NUANCE Several predecessors had noticed similar correspondences (the Jesuit Thomas Stephens in Goa, Sassetti himself, Coeurdoux's 1767 manuscript). Jones's contribution was the explicit family-tree formulation, including Persian, Gothic and Celtic.
Auroux et al., History of the Language Sciences, vol. 2 (2001).
Online: Indo-European studies history.

From coincidences to systems: Grimm's Law

  • Where Latin has p, native English vocabulary consistently has f: pater/father, piscis/fish, pēs/foot, plēnus/full, prō/for.
  • Where Latin has k (written c), English has h: centum/hundred, cor/heart, canis/hound, caput/head.
  • Grimm's Law.
  • Nine consonant shifts.

Neogrammarian principle: sound laws have no exceptions

  • Verner's Law (move to overflow).

Borrowing

  • schedule doesn't follow Grimm's Law (entered via Latin/French).
  • Hindi kitāb (Arabic loan) vs pustak (via Sanskrit).
  • Tamil: puttakam (via Sanskrit) vs native ēṭu.
  • Mama-papa: universal.
SUPPORTED Cognate sets and Grimm's three-step chain shift (PIE voiceless stops → Gmc voiceless fricatives; voiced → voiceless; voiced aspirates → plain voiced or fricatives). First sketched by Rasmus Rask (1818); systematized by Jacob Grimm in Deutsche Grammatik vol. 2 (1822).
Fortson (2010), §15.6; Ringe, From PIE to Proto-Germanic (2nd ed., 2017).
Online: Grimm's Law; Britannica; Ringe handout, U Penn (PDF).
NUANCE “Nine consonant shifts” depends on how you count. The standard presentation gives three series × three places of articulation = 9 changes: p→f, t→θ, k→x; b→p, d→t, g→k; bʰ→b, dʰ→d, gʰ→g. (Plus the labiovelars complicate this slightly.)
SUPPORTED Neogrammarian thesis (Ausnahmslosigkeit): Karl Verner's 1875 paper “Eine Ausnahme der ersten Lautverschiebung” (publ. 1876 in KZ 23) cleaned up Grimm's residue by tying voicing to PIE accent.
Fortson (2010), §15.10.
Online: Verner's Law; Britannica: Verner's Law.
SUPPORTED Loans that violate Grimm's Law are the classic diagnostic of post-shift borrowing — schedule, pedal, cardiac, etc.
DEBATED “Mama-papa: universal” is overstated. Murdock's data (1959) showed ~52% of sampled societies use [ma/na] for “mother” and ~55% use [pa/ta] for “father” — common, but not universal. Jakobson (1962) gave the standard articulatory explanation; reduplicated bilabials/dentals emerge naturally from infant babbling.
Jakobson, “Why ‘mama’ and ‘papa’?” in Selected Writings I (1962); Murdock, “Cross-Language Parallels in Parental Kin Terms” (1959).
Online: Sentence first: mamas & papas.

Laryngeal triumph: science = prediction

  • Saussure (1879): PIE must have had “lost segments,” not found in any daughter language.
    • Saussure predicted *peh₂s-:
      • Latin: pāscō (laryngeal gone, but it left the vowel long).
    • Hittite (1927): found h exactly where predicted.
      • Hittite paḫš- (laryngeal visible).
SUPPORTED Saussure, Mémoire sur le système primitif des voyelles dans les langues indo-européennes, Leipzig, 1879 (often cited as 1878 because the printing was finished in December 1878 with title-page year 1879). Pure internal reconstruction — he posited “coefficients sonantiques” with no direct evidence in any known daughter language.
Fortson (2010), §3.4; Beekes, Comparative Indo-European Linguistics (rev. 2011), ch. 9.
Online: Laryngeal theory; Original 1879 Mémoire at Gallica/BnF.
SUPPORTED Correct date. Hittite was deciphered by Bedřich Hrozný in 1915 (proved Indo-European in 1916–1917), but the recognition that Hittite's ḫ matched Saussure's predicted laryngeals came with Jerzy Kuryłowicz, “ə indoeuropéen et ḫ hittite” (1927), in Symbolae Grammaticae in honorem Joannis Rozwadowski II. Saussure had died in 1913, 14 years before vindication.
Lehmann, Theoretical Bases of Indo-European Linguistics (1993), ch. 5.
Online: Wikipedia: Kuryłowicz and Hittite; U Texas LRC: Laryngeal theory.
SUPPORTED The *peh₂s- “protect, shepherd” example is canonical: Latin pāscō, Hittite paḫš-, with the lengthened ā in Latin and the surviving ḫ in Hittite.
LIV² (Rix et al., 2001), s.v. *peh₂(s)-.
Online: Wiktionary: *peh₂-.

Family tree

  • 1860s.
  • Shared innovation: only valid grouping criterion. Shared retentions prove nothing.
  • centum/satem split = ?
  • Anatolian = first branch to break off.
  • Tocharian.
SUPPORTED August Schleicher published the first IE Stammbaum in 1853, with a fuller version in his Compendium (1861–1862) — so “1860s” is essentially right but the seed is from 1853.
SUPPORTED The shared-innovations criterion (Brugmann's gemeinsame Neuerungen) is the standard cladistic principle in historical linguistics.
Fortson (2010), §1.10; Hock (1991), ch. 18.
DEBATED “centum/satem split”: this is not a primary phylogenetic split. Since the early 20th century centum/satem has been treated as an areal isogloss, not a clean binary subgrouping. Tocharian (geographically eastern) is centum; the satem changes look like waves through a dialect continuum.
Fortson (2010), §3.16; Mallory & Adams (2006), ch. 4.
Online: Centum and satem languages.
SUPPORTED Anatolian-first (the “Indo-Hittite” hypothesis, Sturtevant 1926/1933) is now the majority view: Anatolian lacks features (e.g., feminine gender, the full present/aorist system) that the other branches share, suggesting it split off before those innovations.
Sturtevant, A Comparative Grammar of the Hittite Language (1933); Kloekhorst, “The Anatolian stop system and the Indo-Hittite hypothesis” (2016).
Online: Indo-Hittite; Kloekhorst PDF.
NUANCE Tocharian (NW China, 6th–9th c. CE manuscripts) is famously centum despite its easternmost location — one of the strongest pieces of evidence that centum/satem isn't a clean east/west split.
Mallory & Mair, The Tarim Mummies (2000); Adams, A Dictionary of Tocharian B (rev. 2013).

Culture and archaeology

  • PIE words: wheel, axle, yoke, wool, horse, honey, bee.
    • wheel:
      • time = after wheel was invented (4th millennium BCE).
      • place = where horses were domesticated (Yamnaya).
    • 3500 to 3000 BCE.
SUPPORTED Reconstructible vocabulary: *kʷékʷlos “wheel”, *h₂eḱs- “axle”, *yugóm “yoke”, *h₁éḱwos “horse”, *médʰu “honey/mead”, *bʰi- “bee”, *h₂wĺ̥h₁neh₂ “wool”.
Mallory & Adams, Oxford Introduction to PIE (2006), chs. 14–19.
Online: Wiktionary: *kʷékʷlos; Language Log: horse & wheel in IE.
SUPPORTED The wheel argument: wheeled vehicles appear archaeologically ca. 3500 BCE; the wheel-word is reconstructible to all major branches; therefore PIE breakup is post-3500 BCE. This is David Anthony's signature argument, building on earlier work by Bill Darden, Don Ringe and others.
Anthony (2007), ch. 4 (“Language and Time 2”); Ringe et al., “IE and computational cladistics” (2002).
Online: Anthony, ch. 4 (online excerpt).
NUANCE The 3500–3000 BCE range fits the Steppe (Yamnaya) homeland model. The competing Anatolian/farming-spread model (Renfrew 1987; Bouckaert et al. 2012) puts PIE much earlier (~7000 BCE) but is rejected by most historical linguists precisely because of the wheel/wagon vocabulary. The 2015 ancient-DNA papers (Haak et al., Allentoft et al.) substantially favor the Steppe model.
Haak et al., “Massive migration from the steppe”, Nature 522 (2015).
Online: Haak et al. (Nature 2015); Kurgan hypothesis.

Ablaut and the imperishable-fame formula

The famous e/o/zero ablaut you see in English sing/sang/sung is a direct inheritance from PIE, and you see the same pattern in Greek leíp-ō / lé-loip-a / é-lip-on (“I leave / I have left / I left”). We can even reconstruct poetic formulas: imperishable fame survives as Greek kléos áphthiton and Vedic śravas akṣitam — almost certainly the same phrase, sung by Indo-European bards before the daughter languages parted.

SUPPORTED PIE ablaut (e/o/zero/lengthened grades) is one of the most robust reconstructions; English sing/sang/sung (with a layer of Verner alternation) and Greek leíp-/loip-/lip- are textbook illustrations.
Fortson (2010), ch. 5; Beekes (2011), ch. 11.
Online: Indo-European ablaut.
SUPPORTED kléos áphthiton = śravas akṣitam “imperishable fame”: the canonical IE poetic formula. The match was first noted by Adalbert Kuhn (1853); it became foundational for the field of Indo-European poetics, especially in Calvert Watkins's How to Kill a Dragon (1995).
Watkins, How to Kill a Dragon: Aspects of Indo-European Poetics (Oxford, 1995), part II.4; Schmitt, Dichtung und Dichtersprache in indogermanischer Zeit (1967).
Online: Encyclopaedia Iranica: PIE prosody.
NUANCE “Almost certainly the same phrase” is the dominant view but not unanimous — Margalit Finkelberg and others have argued the Greek phrase may be a later inner-Greek formation. The Indo-European reading remains the consensus among IE poetics specialists (Watkins, Nagy, West).
Finkelberg, “κλέος ἄφθιτον revisited”, CP 102 (2007).

Proto-Indo-Iranian (~2500–2000 BCE)

The satem shift

  • PIE *ḱm̥tóm = hundred
    • Sanskrit śatam, Avestan satəm, Old Persian θata, Modern Persian sad, Hindi sau.
      • Palatovelar *ḱ becomes a sibilant ś.
    • Latin centum, Greek hekatón, English hundred.

Vowel merger

  • PIE had distinct *e, *o, *a. Indo-Iranian merges all three into *a.
  • Greek pherō / Latin ferō / Sanskrit bharā́mi “I carry.”

RUKI rule

  • PIE *s becomes *š (later Sanskrit ṣ) after r, u, k, i. PIE *nisdós “nest” gives Sanskrit nīḍa.

Brugmann's Law

  • PIE *o in open syllables lengthens to PII *ā. PIE *bʰórom → Sanskrit bhāram “load.”
SUPPORTED The satem shift (palatovelars → sibilants) and the śatam/centum contrast are textbook. Hindi sau from śatam via expected MIA developments.
Fortson (2010), §10.3.
Online: Centum/satem.
SUPPORTED PIE three-way e/o/a → Indo-Iranian single a is the “a-merger”; the pherō / ferō / bharāmi trio is the canonical illustration (with PIE *bʰ giving Skt bh, Gk ph, Lat f-).
Fortson (2010), §10.4; Mayrhofer, Indogermanische Grammatik I.2 (1986).
SUPPORTED RUKI: PIE *s → š/ṣ after r, u, k, i. Formulated by Holger Pedersen for satem languages; nearly exceptionless in Indo-Iranian, with restorative exceptions in compound boundaries and a few morphologically conditioned cases. *nisdós > Skt nīḍa is the textbook example.
Fortson (2010), §3.21.
Online: RUKI sound law.
SUPPORTED Brugmann's Law (1876): PIE *o → PII *ā in open non-final syllables. Classic examples: *dóru > dā́ru, *bʰórom > bhā́ram. The law is widely accepted but its precise conditioning is still debated (Kiparsky 2010 argues for a morphologized version).
Fortson (2010), §10.4.
Online: Brugmann's Law.

Indo-Aryan / Iranian split (~1800 BCE)

Small but systematic divergences

  • PII *s → Iranian h:
    • sapta “seven” / Avestan hapta / Old Persian hafta / Modern Persian haft.
    • Sanskrit soma / Avestan haoma.
    • Sanskrit Sindhu / Old Persian Hindu — whence Greek Indos, English India, Hindu.
  • Iranian removes aspiration from *bʰ, *dʰ, *gʰ → b, d, g.
    • Sanskrit bhrātar / Avestan brātar / Persian barādar “brother.”
  • PII *ś (from PIE *ḱ) + v → Iranian sp.
    • Sanskrit aśva / Avestan aspa / Old Persian asa / Modern Persian asb.
  • The deva/daēva inversion.
    • PIE *deywós “celestial, god” → Sanskrit deva “god” but Avestan daēva “demon.”
    • Sanskrit asura (lord in Rigveda; demon later) / Avestan ahura.
SUPPORTED PII *s → Iranian h: confirmed; one of the cleanest Iranian sound laws. Examples are correct (sapta/hapta, soma/haoma, Sindhu/Hindu).
Hoffmann & Forssman, Avestische Laut- und Flexionslehre (rev. 2004); Skjærvø in Cambridge Encyclopedia (2004).
Online: Britannica: Indo-Iranian phonology.
NUANCE The rule has conditions: s survives in Iranian before nonnasal stops, and after RUKI-environment triggers (i, u, r, k) where it had already become š. So Avestan keeps s in some positions (e.g. asti “is”).
Beekes, A Grammar of Gatha-Avestan (1988).
SUPPORTED Deaspiration of voiced aspirates in Iranian; bhrātar/brātar/barādar is the classic example.
Fortson (2010), §10.5.
SUPPORTED aśva/aspa: PIE *h₁éḱwos > PII *Háćwa- > Skt aśva / Av. aspa. Saka languages (e.g. Khotanese aśśa) show a different reflex, which is one piece of evidence used to subgroup within Iranian.
Mayrhofer, EWAia I, s.v. aśva.
Online: Encyclopaedia Iranica: Eastern Iranian.
DEBATED The deva/daēva & asura/ahura inversion is real but its date is disputed. In the oldest layer (the Gāthās), Avestan daēva doesn't yet have the full “demon” meaning — the demonization sharpens later, in Younger Avestan. So this looks like a post-PII religious reform (Zarathustra's), not a shared inheritance. Likewise the Vedic demonization of asura develops over time (positive in early Rigveda, negative by the Atharvaveda).
Skjærvø, The Spirit of Zoroastrianism (2011); Hale, &Ā;sura- in Early Vedic Religion (1986); Boyce, A History of Zoroastrianism, vol. 1 (1975).
Online: Daeva; Encyclopaedia Iranica: daiva-.

Vedic Sanskrit to Classical Sanskrit

  • Vedic Sanskrit: messy, freer syntax; more verbal forms; pitch/accent.
  • Classical Sanskrit: Pāṇini's Aṣṭādhyāyī (~5th c. BCE).
SUPPORTED Vedic had pitch accent (preserved in recitation traditions), a richer system of moods/tenses (subjunctive, injunctive, plural of the aorist), and freer word order. Pāṇini (likely 5th–4th c. BCE, Gandhāra) codified Classical Sanskrit in ~4,000 sūtras.
Cardona, Pāṇini: His Work and its Traditions (2nd ed., 1997); Witzel, “Tracing the Vedic Dialects” (1989).
Online: Pāṇini.

Sanskrit to Prakrits to modern IA

  • Regional “natural” languages:
    • Māhārāṣṭrī (ancestor of Marathi/Konkani)
    • Śaurasenī (Hindi belt)
    • Māgadhī (Bengali, Odia, Assamese, Bihari)
    • Ardha-Māgadhī (Jain canon)
    • Pali (Theravāda Buddhism — essentially a Western Prakrit)
  • Simplification:
    • Simplify clusters via gemination (consonant doubling via assimilation).
    • Intervocalic consonants weaken (lenition).
    • Vowels assimilate.
    • Degemination with compensatory lengthening of the preceding vowel.
  • Examples:
    • Sanskrit sapta → Pali satta → Hindi sāt “seven.”
    • Sanskrit hasta “hand” → Prakrit hattha → Hindi hāth / Marathi hāt.
    • Sanskrit karma → Prakrit kamma → Hindi kām “work.”
    • Sanskrit agni “fire” → Prakrit aggi → Hindi āg.
    • Sanskrit dugdha “milk” → Prakrit duddha → Hindi dūdh.
    • Sanskrit akṣi “eye” → Prakrit acchi → Hindi ā̃kh.
    • Latin noctem → Italian notte → Spanish noche — parallel.

What's in a name

  • PIE *h₁nómn̥
    • PII *Hnā́ma
      • Sanskrit nā́ma → Pali nāma → Hindi/Marathi nām
      • Avestan nąman → Old Persian nāma → Middle Persian nām → Modern Persian nām
    • Latin nōmen
    • Greek ónoma, English name
SUPPORTED The Prakrit list and their descendants are standardly grouped this way; for the lineages see Masica (1991) and Cardona & Jain (2003). Pali's regional affiliation is debated (some put it closer to a Western/Mid-Indo-Aryan dialect rather than Magadhi); but it's certainly a Prakrit, not a continuation of Pāṇinian Sanskrit.
Masica, The Indo-Aryan Languages (1991), ch. 2; Cardona & Jain (eds.), The Indo-Aryan Languages (Routledge, 2003).
Online: Prakrit; Pali.
SUPPORTED Cluster simplification by assimilation (sapta>satta, karma>kamma, akṣi>acchi) and intervocalic lenition with compensatory lengthening are the canonical MIA → NIA changes. Pischel's Grammatik der Prakrit-Sprachen (1900) is the standard reference; Turner's CDIAL documents each lineage.
Pischel, Grammatik der Prakrit-Sprachen (1900, Eng. tr. 1957); Turner, A Comparative Dictionary of the Indo-Aryan Languages (CDIAL, 1962–1966).
Online: CDIAL at DSAL/Chicago; Phonological history of Hindustani.
NUANCE “The weaker consonant becomes a copy of the stronger” (in the original notes' XXX) — this is regressive assimilation in MIA clusters: a stop + non-stop generally yields a geminate of the stop (e.g. -rm- > -mm-, -kṣ- > -cch- via *-tsy-/*-kṣ-). It's the more sonorous member that loses its identity, not strictly the “weaker.”
Bloch, Indo-Aryan from the Vedas to Modern Times (1965), §§129–134.
SUPPORTED Latin>Italian>Spanish noctem>notte>noche shows exactly the parallel: cluster > geminate > affricate. The IA and Romance trajectories are typologically similar.
SUPPORTED *h₁nómn̥ “name”: well-reconstructed neuter r/n-stem.
Mallory & Adams (2006), §13.4.
Online: Wiktionary: *h₁nómn̥.

Dravidian influence on Sanskrit

  • AASI, ASI, ANI, IVC.
  • Dravidian, Munda: older languages, substrate (gone now, since 1500 BCE).
  • Sanskrit:
    • Already different from PII because of Dravidian influence.
    • Retroflexes (ṭ, ḍ, ṇ, ṣ):
      • Sanskrit has way too many retroflexes compared to all its other PIE-descended cousins.
      • Some of these retroflexes are from internal Sanskrit drift, e.g.:
        • RUKI rule: Sanskrit nīḍa “nest” from PIE *nisdós.
        • RUKI + assimilation: Sanskrit iṣṭa “wanted” from PIE *h₂is-tós — retroflex ṭ by assimilation after a RUKI-produced ṣ.
        • From PII: Sanskrit aṣṭā(u) “eight” from PIE *h₃eḱtō — RUKI-like substitution of *ḱt gives Sanskrit ṣṭ, and Avestan has the parallel ašta.
      • But a whole bunch of other Sanskrit words are: 1) retroflex, 2) not found in other PIE descendants, and 3) found in Tamil:
        • kuṭa, kuṭī — “hut, dwelling” — Tamil kuṭi “dwelling.”
        • daṇḍa — “stick, staff” — Tamil taṇṭu “stalk, staff.”
      • Some have been in Sanskrit since the Rigveda:
        • naḷa / naḍa — “reed” — Tamil naḷ.
        • aṇu — “small, atomic” — Tamil aṇu “small.”
      • So: lots of retroflexes appearing where no internal rule predicts them, especially in words for culturally Indian things. Explanation: Dravidian loanwords.
SUPPORTED The three-bucket presentation (Bucket 1 = internal, Bucket 2 = Tamil cognates, Bucket 3 = debated) is the responsible way to argue for substrate influence: rule out internal causation first, then point at the residue. This is the structure used by Krishnamurti (2003) and Kuiper (1991), and it survives Hock's critique better than a flat “retroflexes = Dravidian” claim.
Kuiper, Aryans in the Rigveda (1991); Krishnamurti, The Dravidian Languages (Cambridge, 2003), §1.6.
SUPPORTED Bucket 1 examples (internal):
  • nīḍa < PIE *nisdós (cf. Latin nīdus, English nest): RUKI s → ṣ after i, then zd → ḍ (via the voiced allophone before d).
  • iṣṭa < PIE *h₂is-tós (root *h₂eys- “wish”): retroflex ṭ by assimilation to the RUKI-produced ṣ.
  • aṣṭā(u): the satem palatovelar's RUKI-like behavior before t gives ṣṭ in Sanskrit; Avestan ašta shows the parallel sibilantization. Internal Indo-Iranian, not Dravidian.
Fortson (2010), §10.3; Mayrhofer, EWAia, s.vv.
Online: RUKI sound law; Wiktionary: iṣṭa.
NUANCE The reconstruction for “eight” varies by school. Beekes/LIV-style notation gives *h₃eḱtṓw (with laryngeal); many handbooks write bare *oḱtṓw with no laryngeal; some have *h₁oḱtṓw. The *h₃eḱtō form used here is defensible but readers may meet other notations in standard textbooks.
Beekes, Comparative Indo-European Linguistics (rev. 2011), §13.1; Mallory & Adams (2006), §16.1; Fortson (2010), §6.66.
Online: Wiktionary: *oḱtṓw.
SUPPORTED Bucket 2 examples (Dravidian loans): kuṭa/kuṭī, daṇḍa, naḷa, aṇu are standard candidates in Burrow & Emeneau's Dravidian Etymological Dictionary. The direction (Dravidian → Sanskrit) is reasonably secure: these words have clean Dravidian cognates, no PIE etymology, and many appear in early Vedic.
Burrow & Emeneau, A Dravidian Etymological Dictionary (DED², Oxford 1984); Krishnamurti (2003), §1.6.
Online: DED at DSAL/Chicago.
DEBATED The strong claim — that Dravidian specifically is the source — is one of three positions in the field:
  • Dravidian substrate (Kuiper, Emeneau, Southworth, mostly Krishnamurti): retroflexes spread to IA from Dravidian.
  • Internal & NW areal (Hock 1975, 1996; Tikkanen): retroflexion is largely a NW-South-Asian areal feature with substantial internal causation; the “Burushaski zone” also has retroflexes.
  • Para-Munda mixed substrate (Witzel 1999): the earliest non-IE Rigvedic loans look non-Dravidian (“Para-Munda”); Dravidian contact comes only by middle-Rigvedic times.
All three accept that some retroflexes are loans — they disagree on the donor and the date.
Witzel, “Substrate Languages in Old Indo-Aryan”, EJVS 5.1 (1999); Hock, “Substratum Influence on (Rig-Vedic) Sanskrit?”, SLS 5.2 (1975); Tikkanen, “Burushaski as an aberrant Indo-Iranian language” (1988).
Online: Witzel 1999 (full PDF); Wikipedia: Substratum in Vedic.
NUANCE “Retroflexes not in PII / Iranian” needs one footnote: Iranian doesn't have contrastive retroflexes, but Nuristani (Kafiri) languages and some NW IA dialects have developed them independently, suggesting the NW South Asian convergence zone is a real factor (Tikkanen, Bashir).
Bashir, “The development of retroflexion in NW South Asia”, in Linguistic Convergence (2003).
NUANCE “AASI, ASI, ANI, IVC”: the post-2015 ancient-DNA framing (Ancient Ancestral South Indian / Ancestral South Indian / Ancestral North Indian / Indus Valley Civilization). The genetic story is broadly consistent with a Steppe-derived IA migration into a substrate that included both Dravidian and other (likely non-Dravidian, non-IE) languages.
Narasimhan et al., “The formation of human populations in South and Central Asia”, Science 365 (2019).
Online: Narasimhan et al. 2019 (Science).

Syntactic features attributed to Dravidian / South Asia

SOV

  • Hindi: rām-ne mohan-ko kitāb dī. Tamil: rāmaṉ mōhaṉukku puttakam koṭuttāṉ.
  • PIE was probably partially SOV; Classical Latin was SOV; Old Persian is SOV.
  • All modern European languages: partly or fully SVO. Vedic Sanskrit allowed SV/VS/other patterns. Classical Sanskrit & all later IA: rigidly SOV. Because of Dravidian influence.

Quotative

  • Tamil eṉṟu, Kannada anta/endu, Telugu ani, Malayalam ennu.
  • Sanskrit iti; Hindi ki; Marathi mhaṇūn; Bengali bole.
  • All have a quotative derived from “say” placed after the quoted material, parallel to Dravidian.

Echo reduplication

  • Hindi chāy-vāy, kitāb-vitāb, pānī-vānī.
  • Tamil tēṉīr-kīṉīr, puttakam-kittakam.
  • Kannada chahā-gihā, pustaka-gistaka.
  • Marathi chahā-bihā, pustak-bistak.
  • Bengali chā-ṭā. Telugu ṭī-gīṭī.
  • This is almost everywhere in India. Outside India, not common:
    • English has only marginal Yiddish-borrowed schm- (fancy-schmancy) — used differently.
    • Turkish and Armenian have an m- form: Turkish kitap-mitap, Armenian seġan-meġan.
    • But the density in Indian languages is best explained as contact with Dravidian.

Dative subjects for experiencers

  • Hindi mujhe bhūkh lagī hai = “to-me hunger is felt.”
  • Other IE: French j'ai faim, German ich habe Hunger.

Conjunctive participles

  • Hindi ghar jā-kar khānā khā-yā; Tamil vīṭṭukku pōy cāppiṭṭēṉ.
  • In Western IE this is rare; in Sanskrit it's one option; in Hindi/Tamil it's default.
SUPPORTED Broad consensus (Hock 2013, 2015) is that PIE was an SOV language with V-final tendencies (though with considerable flexibility). The shift towards SVO in many branches happened independently.
Hock, “Proto-Indo-European verb-finality: Reconstruction, typology, validation”, JHL 3 (2013); Lehmann, Proto-Indo-European Syntax (1974) for the original detailed argument.
Online: Hock 2013 (JHL).
NUANCE The claim “Classical Sanskrit became rigidly SOV because of Dravidian influence” is plausible but only one explanation. Hock has argued the verb-final tightening in IA is largely an internal continuation of PIE V-final tendencies amplified by typological drift; Emeneau (1956) and Krishnamurti put more weight on contact. Some areal pressure is widely accepted.
Hock (1991), ch. 18; Emeneau, “India as a Linguistic Area”, Language 32 (1956).
SUPPORTED Quotative constructions: iti in Sanskrit (already in Rigveda), ki/bole/mhaṇūn in NIA, eṉṟu/ani/endu/ennu in Dravidian. The clause-final “say”-quotative is a textbook South Asian areal feature.
Emeneau (1956); Hock (1991), ch. 18; Steever, Analysis to Synthesis: The Development of Complex Verbal Forms (1993).
Online: IAS: Linguistic history of India.
SUPPORTED Correctly framed: dense in South Asia, marginal-but-present outside. Echo / m-reduplication belongs to a broad Eurasian areal feature running from Turkic and Armenian through Mongolic and Persian into the South Asian linguistic area, where it reaches its densest and most productive form. The “density argument” (it's pan-Indic across IA, Dravidian, Munda, Tibeto-Burman) is one of Emeneau's classic India-as-a-linguistic-area diagnostics.
Emeneau, “India as a Linguistic Area”, Language 32 (1956); Abbi, Reduplication in South Asian Languages (Allied 1992); Stolz, Stroh & Urdze, Total Reduplication (Akademie 2011); Southern, Contagious Couplings: Transmission of Expressives in Yiddish Echo Phrases (Praeger 2005).
Online: Echo word; Reduplication: echo-reduplication.
NUANCE Two small additions if you want to be even more defensible: (1) Mongolic (Khalkha) and Dargwa (NE Caucasian) also have m-reduplication, so the “Eurasian belt” is broader than just Turkish + Armenian; (2) the English schm- pattern is itself Yiddish-borrowed, so it isn't an independent Germanic development.
Stolz et al. (2011), §4.3; Nevins & Vaux, “Metalinguistic, Shmetalinguistic” (CLS 2003).
SUPPORTED Dative-experiencer subjects are an Emeneau (1956) hallmark of the Indian linguistic area. Verma & Mohanan, Experiencer Subjects in South Asian Languages (1990) is the standard collection.
Verma & Mohanan (eds., 1990); Bhaskararao & Subbarao (eds.), Non-nominative Subjects (2004).
SUPPORTED Conjunctive participles (“converbs”): pan-Indic, shared across Indo-Aryan, Dravidian, Munda, Tibeto-Burman. Sanskrit -tvā/-ya, Hindi -kar, Tamil -i/-y, etc. The structural parallelism is one of the strongest Sprachbund features.
Masica, Defining a Linguistic Area: South Asia (1976); Slade, “Verb concatenation in Asian linguistics” (2020).
Online: Slade 2020 (PDF).

Dating: how do we get absolute and relative dates?

Hard dates

  • Modern: DNA (201x).
  • Earlier written records:
    • Old Persian: Behistun inscription, ~520 BCE (Darius I).
    • Hittite: cuneiform tablets, ~1650–1200 BCE.
    • Mycenaean Greek: Linear B tablets, ~1400–1200 BCE.
    • First written Sanskrit: Ashokan-era, 3rd c. BCE.
      • But Vedic Sanskrit (Rigveda) ~1500–1200 BCE.
    • Latin: earliest inscriptions ~600 BCE.

Relative dating: layered sound changes

  • Sanskrit sapta → Prakrit satta → Hindi sāt.
    • #1 cluster simplification (pt → tt).
    • #2 degemination (tt → t) and vowel lengthening.
    • #1 must precede #2.
  • Iranian s → h: must be after Indo-Aryan split (2000 BCE), before Old Persian hafta (600 BCE).

Borrowed words freeze at time of borrowing

  • Finnish kuningas “king” borrowed from Proto-Germanic *kuningaz. Germanic itself moved on (English king, German König).

Mitanni treaty

  • ~1380 BCE, northern Syria. Contains Indo-Aryan god-names (Mitra, Varuna, Indra, Nasatya) and horse-training terms (aika-, tera-, panza-, satta-, nava-). Forms more archaic than Vedic. Proves Indo-Aryan existed as a distinct branch by 1400 BCE.

Linguistics + archaeology

  • PIE has solid reconstructions for wheel *kʷékʷlos, axle *h₂eks-, yoke *yugóm, wagon/wain, horse *h₁éḱwos. Wagons appear ~3500 BCE. PIE breakup post-3500 BCE.
  • No reconstructible word for iron → breakup before Iron Age (~1200 BCE).
  • Reconstructed words for bee, honey (*médʰu) but not tropical species → temperate homeland.
  • Indo-Iranian shared vocabulary for chariot, spoke, horse-training dates the common period after the spoked-wheel chariot (Sintashta, ~2000 BCE).
SUPPORTED All the hard dates check out. Behistun: trilingual rock inscription of Darius I, ~520 BCE; Hittite cuneiform: ca. 1650 BCE earliest (Anitta text); Linear B: 1400–1200 BCE; Aśokan edicts: mid-3rd c. BCE; Latin Praeneste fibula / Duenos inscription debated but 7th–6th c. BCE.
Watkins (ed.), The Cambridge Encyclopedia of the World's Ancient Languages (2004).
Online: Behistun; Linear B.
NUANCE Rigveda dating (~1500–1200 BCE) is a linguistic-internal estimate; there are no contemporary external attestations until the Mitanni evidence. Range is widely accepted; some scholars push earlier (Witzel allows down to ~1700 BCE for early hymns).
Witzel, “The Development of the Vedic Canon” (1997); Erdosy (ed.), The Indo-Aryans of Ancient South Asia (1995).
SUPPORTED Relative-chronology argument is sound. The sapta > satta > sāt ordering can't be reversed because there's no rule that takes single intervocalic t to a long-vowel + t; you need the geminate intermediate to license compensatory lengthening.
Masica (1991), §7.2.
SUPPORTED Finnish kuningas < Proto-Germanic *kuningaz: textbook example of a loanword preserving a stage Germanic has since lost. The final -az shows the original PGmc masculine nominative which all Germanic languages have since dropped.
Kallio, “On the Earliest Germanic Loanwords in Uralic” (2012); Koivulehto, Verba mutuata (1999).
Online: Proto-Germanic.
NUANCE Mitanni: god-names appear in the Hittite–Mitanni treaty between Suppiluliuma I and Šattiwaza (~1380 BCE). The horse-training numerals (aika-, tera-, panza-, satta-, nava-) and the term vartana- “turn/lap” come from a separate Hittite manual by Kikkuli (~1400 BCE), not from the treaty itself. The notes conflate two related documents into one.
Mayrhofer, Die Indo-Arier im alten Vorderasien (1966); Burrow, The Sanskrit Language (3rd ed., 1973), §1.3; Thieme, “The 'Aryan' Gods of the Mitanni Treaties”, JAOS 80 (1960).
Online: Indo-Aryan superstrate in Mitanni; Kikkuli.
SUPPORTED The wheel/wagon argument and the absence-of-iron argument are the two pillars of linguistic palaeontology for dating PIE breakup. The temperate-flora-and-fauna argument is older (Schrader 1890s) and still used as a soft constraint.
Anthony (2007), chs. 4–5; Mallory (1989), ch. 5.
SUPPORTED Sintashta (~2100–1800 BCE, S. Urals) is the type-site for the spoke-wheel chariot; matches the Indo-Iranian shared chariot vocabulary date well.
Anthony (2007), ch. 16; Kuznetsov, “The Emergence of Bronze Age Chariots in Eastern Europe” (2006).
Online: Sintashta culture.

Why simplification?

  • Small populations = increasing complexity.
  • Increasing population (conquest/administration, trade, religion) = simplification.
DEBATED This is the “Lupyan–Dale hypothesis”: languages with more L2 speakers tend to simplify morphology. Real correlation in some studies, but contested.
  • Pro: Lupyan & Dale (2010); Trudgill (2011, Sociolinguistic Typology).
  • Skeptical: Bentz & Winter (2013); Sinnemäki & Di Garbo (2018) find mixed results.
Calling it a settled mechanism overstates the field. It's better stated as: contact/L2 acquisition can drive analytic restructuring (e.g. case-loss in English, Persian, Bengali).
Lupyan & Dale, “Language Structure Is Partly Determined by Social Structure”, PLoS ONE 5 (2010); Trudgill, Sociolinguistic Typology: Social Determinants of Linguistic Complexity (Oxford 2011).
Online: Lupyan & Dale 2010 (PLoS ONE).

Reading list (in order of difficulty)

If you want to build up real knowledge of this field, here's a sequenced path. The first three books are accessible; from item 4 onwards you're reading textbooks; from 7 onwards you're reading specialist literature.

  1. David W. Anthony, The Horse, the Wheel, and Language: How Bronze-Age Riders from the Eurasian Steppes Shaped the Modern World (Princeton, 2007).
    The best general-audience entry point. Combines linguistics, archaeology, and a dash of genetics into a single coherent argument for the Steppe homeland. Read this first.
    Princeton UP
  2. J. P. Mallory, In Search of the Indo-Europeans: Language, Archaeology and Myth (Thames & Hudson, 1989).
    Slightly older but extremely readable. Covers the homeland question, branch-by-branch, and the cultural reconstruction. Good complement to Anthony.
  3. Calvert Watkins, How to Kill a Dragon: Aspects of Indo-European Poetics (Oxford, 1995).
    If you want to be persuaded that we can reach back to PIE poetry, this is the canonical text. Builds out the kléos áphthiton case and many other formulas.
    OUP
  4. Benjamin W. Fortson IV, Indo-European Language and Culture: An Introduction (2nd ed., Wiley-Blackwell, 2010).
    The current standard university textbook. Branch-by-branch, with worked examples of sound laws. If you want a single book that teaches the field, get this one.
    Wiley
  5. J. P. Mallory & D. Q. Adams, The Oxford Introduction to Proto-Indo-European and the Proto-Indo-European World (Oxford, 2006).
    Hybrid: half textbook, half encyclopaedia. Strongest for cultural reconstruction (kinship, society, religion, material culture). Pair with Fortson.
  6. Robert S. P. Beekes, Comparative Indo-European Linguistics: An Introduction (2nd revised ed. by Michiel de Vaan, John Benjamins, 2011).
    Tighter and more technical than Fortson; Leiden-school perspective on phonology and morphology. Good once you've finished Fortson.
    Benjamins
  7. Colin P. Masica, The Indo-Aryan Languages (Cambridge, 1991).
    If you care specifically about Sanskrit → Prakrit → modern IA, this is the standard reference. Dense but comprehensive.
  8. Bhadriraju Krishnamurti, The Dravidian Languages (Cambridge, 2003).
    Counterpart for the Dravidian side. Indispensable for any serious work on the substrate question.
    CUP
  9. Michael Witzel, “Substrate Languages in Old Indo-Aryan,” Electronic Journal of Vedic Studies 5.1 (1999).
    The case for a non-Dravidian (“Para-Munda”) substrate, with extensive lexical evidence. Read alongside Hock's substrate paper for the other side.
    Full PDF
  10. Hans Henrich Hock, Principles of Historical Linguistics (2nd ed., Mouton de Gruyter, 1991).
    Not Indo-European-specific, but the best methodological grounding. After this you can read anything in the field.
  11. Calvert Watkins (ed.), The American Heritage Dictionary of Indo-European Roots (3rd ed., HMH, 2011).
    For browsing: every English word traced back to its PIE root. The best thing to keep on your desk and dip into.
  12. Helmut Rix et al., Lexikon der indogermanischen Verben (LIV²; Reichert, 2001).
    Reference work for verbal roots. Specialist; use as a lookup tool, not bedtime reading.

Online standing references:

If you want only one book: Anthony (#1). It's the most fun, and it'll send you to the right next-reads on its own.