clepsydra

Source

Hierarchical Bayesian lexical dating — when did these languages diverge, and how sure can we be?

This page runs the whole inference engine locally in your browser via WebAssembly — no data is sent anywhere. It is the same C code the command-line tool runs, so a date computed here is the date the CLI computes. The bundled datasets are real, published cognate data (CC-BY-4.0; each is credited where it is selected).

What this does

Given which languages share cognates for each concept, plus a tree and some age calibrations, it estimates the age of every node with a posterior distribution — a range, not a number. The pairwise model does the same for two languages without a tree.

What this does not do

It does not infer the tree, judge whether your cognate coding is right, or turn a weak signal into a confident date. A wide interval is the answer, not a failure. Deep-time estimates depend heavily on the rate prior, and the defaults here are shortened for responsiveness — a result you intend to publish belongs on the CLI.

Two languages, no tree: the hierarchical constant-rate pairwise model (P1). It runs in well under a second, so it updates as you change things.

to

Use your own data

A forms table in Clepsydra's TSV format: columns language_id, concept_id, cognate_id, and optionally primary_weight. It is read in this tab and goes nowhere else.

years BP (median)

posterior over divergence time

The whole distribution, not just its middle. The shaded band is the central 90%; the dashed line is the median.