The Indo-Austronesian Hypothesis

Models, physics, mind, society, and culture

An Indo-European and Austronesian comparison list becomes a testable research program through lexical provenance, regular sound correspondences, held-out predictions, and controls for coincidence.

I want to put a dangerous linguistic idea into a form that comparative linguists can work with.

The proposal is that the Indo-European and Austronesian language families may share a remote common ancestor. Similar words are plentiful: languages contain short words, human mouths reuse a limited set of sounds, meanings drift, and a long search produces coincidences.

The path forward is finite. Publish the seed list, establish every reconstruction inside its own family, derive candidate correspondences, and use those correspondences to predict material that was held back. Writing down the pattern gives the comparative method something exact to correct, extend, or reject.

Check the 2022 Comparison List Against the Dictionaries

These are the comparisons that first attracted my attention:

Proposed meaning Indo-European-side form as written Austronesian-side form as written Working reading
three *treys *telu Both forms have dictionary trails; test the correspondence
hand / five *ronk *lima Remove this row as written
flow *ser *qalur Attach exact reconstructions and reflexes
skin *skend *qanic Attach exact reconstructions and reflexes
walk / foot *stembh *qaqay Attach exact reconstructions and reflexes
smoke / ash *smew *qabu Attach exact reconstructions and reflexes
watch / day *serw *qalayaw Attach exact reconstructions and reflexes
salt *sal *qasira Attach exact reconstructions and reflexes
talk / mouth *bhas *baqbaq Attach exact reconstructions and reflexes
bow *bhewgh *busur Attach exact reconstructions and reflexes
give birth / woman *bher *bay Attach exact reconstructions and reflexes
two *dwoh₁ *dusa Attach exact reconstructions and reflexes
sky / upward *dyews *daya Attach exact reconstructions and reflexes
river / lake *danu *danaw Attach exact reconstructions and reflexes
tongue *dn̥ǵʰwéh₂ *dilaq Keep *dilaq as “lick” and rebuild the semantic route
wide *delh₁ *dempad Attach exact reconstructions and reflexes
I *eǵoh₂ *aku Attach exact reconstructions and reflexes
think *men *nemnem Attach exact reconstructions and reflexes
name *h₁nómn̥ Proto-Polynesian *hingoa Compare at the Proto-Polynesian level only

The asterisks reproduce the notation of the original proposal. Each row now needs its exact published source and proto-language level beside it.

Checking modern comparative lexicons immediately changed the list. The Indo-European Lexicon’s set for “hand” does not contain *ronk. The Austronesian Comparative Dictionary distinguishes PAN *lima, “five,” from PAN *qalima, “hand.” The original hand/five row slides between two glosses and leaves the comparison as written. The same dictionary gives PAN *dilaq as “to lick,” rather than “tongue.” That semantic relationship may explain where the comparison came from, but it requires a separately documented route before it can reenter the table.

I am keeping the corrected rows visible because they teach me how to rebuild the rest. Every surviving pair gets citations, proto-language levels, daughter-language reflexes, and within-family derivations. The “name” comparison is explicitly Proto-Polynesian, a far later node than Proto-Austronesian, so I keep it at that level instead of silently moving it deeper in time.

The table turns a feeling into a finite object another person can inspect, reproduce, and improve.

It also changes how I search. I begin with a fixed concept, retrieve the published reconstructions from each family’s own reference works, follow their daughter reflexes, and only then compare the forms across families. A semantic bridge such as “lick” to “tongue” becomes its own historical argument instead of a quiet substitution inside a row.

Start Inside Each Language Family

Before comparing two families, each side must be valid inside its own family.

For every row I need a record containing:

This is the first experiment. A form reconstructed only for Proto-Polynesian stays at that level. A root glossed “lick” can reach “tongue” only through a separately argued semantic history. A nonexistent reconstruction leaves nothing to compare.

The Austronesian Comparative Dictionary is especially useful because it distinguishes several reconstruction levels and separates canonical comparisons, near comparisons, loans, and chance look-alikes. That is exactly the discipline this proposal needs. The list should shrink before it grows.

The sound correspondences I suspect

The correspondences I originally suspected are:

A real historical relationship would not merely generate similar-looking words. It would generate repeated, conditioned sound correspondences. A sound should change in a regular way depending on its position, neighboring sounds, stress, and the history of the daughter languages.

I write s : q as a candidate correspondence because I selected it after seeing the seed list. It now has to earn generality across sourced forms and new material.

I therefore search systematically for the places where each proposed correspondence breaks instead of collecting another page of attractive pairs.

The resulting correspondence table should be conditioned by sound environment: initial, medial, or final position; neighboring vowels and consonants; stress; and the chronology of changes already established within each daughter family. One symbol pair repeated without those conditions is a visual pattern. A conditioned system can begin making new lexical predictions.

For every candidate s : q match, I ask:

  1. How many words with the same sound environment do not match?
  2. Does s correspond to several unrelated Austronesian sounds whenever convenient?
  3. Does q correspond to several Indo-European sounds with equal freedom?
  4. Are the meanings historically plausible rather than merely imaginable?
  5. Could known contact or borrowing explain the pair?
  6. Do inflections, pronouns, or other grammatical material preserve the same pattern?

The hypothesis becomes interesting only when the rules predict comparisons I did not use to invent them.

Count the Search, Not Just the Match

I originally estimated the chance of obtaining seven apparent s : q matches from short roots. The calculation produced an attractively small number because it treated the correspondence as though I had selected it before looking.

In the actual search, multiple sound pairings were available. Meanings could drift. Candidate roots could enter or leave the list. The words are statistically related, and reconstructed phoneme inventories are structured systems rather than uniform bags of consonants.

This is the multiple-search problem: the probability of one predeclared pattern can be small while the probability of finding some interesting pattern after trying many patterns is large.

A better test needs a procedure fixed before examining the held-out data:

  1. Freeze an explicit list of meanings and reconstructions.
  2. Derive candidate sound laws from only part of the list.
  3. Score predictions on untouched meanings.
  4. Compare the score with unrelated language-family pairs processed by the same method.
  5. Penalize semantic flexibility and one-off exceptions.
  6. Repeat the analysis under competing reconstructions.

The useful comparison is whether the proposed relationship predicts withheld structure better than unrelated family pairs and better than alternative explanations.

What a Family Relationship Should Predict

Vocabulary alone is vulnerable to coincidence and borrowing. A relationship reaching back beyond both families should predict several kinds of traces:

Controls complete the test. If the sound correspondences collapse outside the selected pairs, grammatical systems add no matching structure, or unrelated family pairs score just as well, the procedure has answered the question.

A proposal is an invitation

A conjecture becomes useful when it generates work another person can repeat. A clear dataset and a fixed procedure can reveal a better question, a bad statistical habit, or a neglected comparison regardless of where this particular family tree ends.

The list now has a concrete job. Two rows leave the seed set as written. The candidate correspondences came from that same seed set, so the next lexical and grammatical material stays untouched until the rules are frozen. That sequence gives specialists something much better than a resemblance between three and telu: a comparison they can run through the machinery of historical linguistics.

The list is here. The candidate correspondences are here. The corrections are part of the route.

Now the idea can meet experts and become a real comparative-linguistics project.

Further reading