-
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
Authors:
Schyan Zafar,
Geoff K. Nicholls
Abstract:
Word meanings change over time, and word senses evolve, emerge or die out in the process. For ancient languages, where the corpora are often small and sparse, modelling such changes accurately proves challenging, and quantifying uncertainty in sense-change estimates consequently becomes important. GASC (Genre-Aware Semantic Change) and DiSC (Diachronic Sense Change) are existing generative models…
▽ More
Word meanings change over time, and word senses evolve, emerge or die out in the process. For ancient languages, where the corpora are often small and sparse, modelling such changes accurately proves challenging, and quantifying uncertainty in sense-change estimates consequently becomes important. GASC (Genre-Aware Semantic Change) and DiSC (Diachronic Sense Change) are existing generative models that have been used to analyse sense change for target words from an ancient Greek text corpus, using unsupervised learning without the help of any pre-training. These models represent the senses of a given target word such as "kosmos" (meaning decoration, order or world) as distributions over context words, and sense prevalence as a distribution over senses. The models are fitted using Markov Chain Monte Carlo (MCMC) methods to measure temporal changes in these representations. This paper introduces EDiSC, an Embedded DiSC model, which combines word embeddings with DiSC to provide superior model performance. It is shown empirically that EDiSC offers improved predictive accuracy, ground-truth recovery and uncertainty quantification, as well as better sampling efficiency and scalability properties with MCMC methods. The challenges of fitting these models are also discussed.
△ Less
Submitted 25 June, 2024; v1 submitted 1 November, 2023;
originally announced November 2023.
-
TraitLab: a Matlab package for fitting and simulating binary tree-like data
Authors:
Luke J. Kelly,
Geoff K. Nicholls,
Robin J. Ryder,
David Welch
Abstract:
TraitLab is a software package for simulating, fitting and analysing tree-like binary data under a stochastic Dollo model of evolution. The model also allows for rate heterogeneity through catastrophes, evolutionary events where many traits are simultaneously lost while new ones arise, and borrowing, whereby traits transfer laterally between species as well as through ancestral relationships. The…
▽ More
TraitLab is a software package for simulating, fitting and analysing tree-like binary data under a stochastic Dollo model of evolution. The model also allows for rate heterogeneity through catastrophes, evolutionary events where many traits are simultaneously lost while new ones arise, and borrowing, whereby traits transfer laterally between species as well as through ancestral relationships. The core of the package is a Markov chain Monte Carlo (MCMC) sampling algorithm that enables the user to sample from the Bayesian joint posterior distribution for tree topologies, clade and root ages, and the trait loss, catastrophe and borrowing rates for a given data set. Data can be simulated according to the fitted Dollo model or according to a number of generalized models that allow for heterogeneity in the trait loss rate, biases in the data collection process and borrowing of traits between lineages. Coupled pairs of Markov chains can be used to diagnose MCMC mixing and convergence and to debias MCMC estimators. The raw data, MCMC run output, and model fit can be inspected using a number of useful graphical and analytical tools provided within the package or imported into other popular analysis programs. TraitLab is freely available and runs within the Matlab computing environment with its Statistics and Machine Learning toolbox, no other additional toolboxes are required.
△ Less
Submitted 17 August, 2023;
originally announced August 2023.
-
Bayesian Inference for Vertex-Series-Parallel Partial Orders
Authors:
Chuxuan,
Jiang,
Geoff K. Nicholls,
Jeong Eun Lee
Abstract:
Partial orders are a natural model for the social hierarchies that may constrain "queue-like" rank-order data. However, the computational cost of counting the linear extensions of a general partial order on a ground set with more than a few tens of elements is prohibitive. Vertex-series-parallel partial orders (VSPs) are a subclass of partial orders which admit rapid counting and represent the sor…
▽ More
Partial orders are a natural model for the social hierarchies that may constrain "queue-like" rank-order data. However, the computational cost of counting the linear extensions of a general partial order on a ground set with more than a few tens of elements is prohibitive. Vertex-series-parallel partial orders (VSPs) are a subclass of partial orders which admit rapid counting and represent the sorts of relations we expect to see in a social hierarchy. However, no Bayesian analysis of VSPs has been given to date. We construct a marginally consistent family of priors over VSPs with a parameter controlling the prior distribution over VSP depth. The prior for VSPs is given in closed form. We extend an existing observation model for queue-like rank-order data to represent noise in our data and carry out Bayesian inference on "Royal Acta" data and Formula 1 race data. Model comparison shows our model is a better fit to the data than Plackett-Luce mixtures, Mallows mixtures, and "bucket order" models and competitive with more complex models fitting general partial orders.
△ Less
Submitted 27 June, 2023;
originally announced June 2023.
-
Biclustering random matrix partitions with an application to classification of forensic body fluids
Authors:
Chieh-Hsi Wu,
Amy D. Roeder,
Geoff K. Nicholls
Abstract:
Classification of unlabeled data is usually achieved by supervised learning from labeled samples. Although there exist many sophisticated supervised machine learning methods that can predict the missing labels with a high level of accuracy, they often lack the required transparency in situations where it is important to provide interpretable results and meaningful measures of confidence. Body flui…
▽ More
Classification of unlabeled data is usually achieved by supervised learning from labeled samples. Although there exist many sophisticated supervised machine learning methods that can predict the missing labels with a high level of accuracy, they often lack the required transparency in situations where it is important to provide interpretable results and meaningful measures of confidence. Body fluid classification of forensic casework data is the case in point. We develop a new Biclustering Dirichlet Process for Class-assignment with Random Matrices (BDP-CaRMa), with a three-level hierarchy of clustering, and a model-based approach to classification that adapts to block structure in the data matrix. As the class labels of some observations are missing, the number of rows in the data matrix for each class is unknown. BDP-CaRMa handles this and extends existing biclustering methods by simultaneously biclustering multiple matrices each having a randomly variable number of rows. We demonstrate our method by applying it to the motivating problem, which is the classification of body fluids based on mRNA profiles taken from crime scenes. The analyses of casework-like data show that our method is interpretable and produces well-calibrated posterior probabilities. Our model can be more generally applied to other types of data with a similar structure to the forensic data.
△ Less
Submitted 14 October, 2023; v1 submitted 27 June, 2023;
originally announced June 2023.
-
Bayesian inference for partial orders from random linear extensions: power relations from 12th Century Royal Acta
Authors:
Geoff K. Nicholls,
Jeong Eun Lee,
Nicholas Karn,
David Johnson,
Rukuang Huang,
Alexis Muir-Watt
Abstract:
We give a new class of models for time series data in which actors are listed in order of precedence. We model the lists as a realisation of a queue in which queue-position is constrained by an underlying social hierarchy. We model the hierarchy as a partial order so that the lists are random linear extensions. We account for noise via a random queue-jum** process. We give a marginally consisten…
▽ More
We give a new class of models for time series data in which actors are listed in order of precedence. We model the lists as a realisation of a queue in which queue-position is constrained by an underlying social hierarchy. We model the hierarchy as a partial order so that the lists are random linear extensions. We account for noise via a random queue-jum** process. We give a marginally consistent prior for the stochastic process of partial orders based on a latent variable representation for the partial order. This allows us to introduce a parameter controlling partial order depth and incorporate actor-covariates informing the position of actors in the hierarchy.
We fit the model to witness lists from Royal Acta from England, Wales and Normandy in the eleventh and twelfth centuries. Witnesses are listed in order of social rank, with any bishops present listed as a group. Do changes in the order in which the bishops appear reflect changes in their personal authority? The underlying social order which constrains the positions of bishops within lists need not be a complete order and so we model the evolving social order as an evolving partial order. The status of an Anglo-Norman bishop was at the time partly determined by the length of time they had been in office. This enters our model as a time-dependent covariate. We fit the model, estimate partial orders and find evidence for changes in status over time. We interpret our results in terms of court politics. Simpler models, based on Bucket Orders and vertex-series-parallel orders, are rejected. We compare our results with a time-series extension of the Plackett-Luce model.
△ Less
Submitted 1 August, 2023; v1 submitted 11 December, 2022;
originally announced December 2022.
-
Scalable Semi-Modular Inference with Variational Meta-Posteriors
Authors:
Chris U. Carmona,
Geoff K. Nicholls
Abstract:
The Cut posterior and related Semi-Modular Inference are Generalised Bayes methods for Modular Bayesian evidence combination. Analysis is broken up over modular sub-models of the joint posterior distribution. Model-misspecification in multi-modular models can be hard to fix by model elaboration alone and the Cut posterior and SMI offer a way round this. Information entering the analysis from missp…
▽ More
The Cut posterior and related Semi-Modular Inference are Generalised Bayes methods for Modular Bayesian evidence combination. Analysis is broken up over modular sub-models of the joint posterior distribution. Model-misspecification in multi-modular models can be hard to fix by model elaboration alone and the Cut posterior and SMI offer a way round this. Information entering the analysis from misspecified modules is controlled by an influence parameter $η$ related to the learning rate. This paper contains two substantial new methods. First, we give variational methods for approximating the Cut and SMI posteriors which are adapted to the inferential goals of evidence combination. We parameterise a family of variational posteriors using a Normalising Flow for accurate approximation and end-to-end training. Secondly, we show that analysis of models with multiple cuts is feasible using a new Variational Meta-Posterior. This approximates a family of SMI posteriors indexed by $η$ using a single set of variational parameters.
△ Less
Submitted 1 April, 2022;
originally announced April 2022.
-
Valid belief updates for prequentially additive loss functions arising in Semi-Modular Inference
Authors:
Geoff K. Nicholls,
Jeong Eun Lee,
Chieh-Hsi Wu,
Chris U. Carmona
Abstract:
Model-based Bayesian evidence combination leads to models with multiple parameteric modules. In this setting the effects of model misspecification in one of the modules may in some cases be ameliorated by cutting the flow of information from the misspecified module. Semi-Modular Inference (SMI) is a framework allowing partial cuts which modulate but do not completely cut the flow of information be…
▽ More
Model-based Bayesian evidence combination leads to models with multiple parameteric modules. In this setting the effects of model misspecification in one of the modules may in some cases be ameliorated by cutting the flow of information from the misspecified module. Semi-Modular Inference (SMI) is a framework allowing partial cuts which modulate but do not completely cut the flow of information between modules. We show that SMI is part of a family of inference procedures which implement partial cuts. It has been shown that additive losses determine an optimal, valid and order-coherent belief update. The losses which arise in Cut models and SMI are not additive. However, like the prequential score function, they have a kind of prequential additivity which we define. We show that prequential additivity is sufficient to determine the optimal valid and order-coherent belief update and that this belief update coincides with the belief update in each of our SMI schemes.
△ Less
Submitted 24 January, 2022;
originally announced January 2022.
-
Tree based credible set estimation
Authors:
Jeong Eun. Lee,
Geoff K. Nicholls
Abstract:
Estimating a joint Highest Posterior Density credible set for a multivariate posterior density is challenging as dimension gets larger. Credible intervals for univariate marginals are usually presented for ease of computation and visualisation. There are often two layers of approximation, as we may need to compute a credible set for a target density which is itself only an approximation to the tru…
▽ More
Estimating a joint Highest Posterior Density credible set for a multivariate posterior density is challenging as dimension gets larger. Credible intervals for univariate marginals are usually presented for ease of computation and visualisation. There are often two layers of approximation, as we may need to compute a credible set for a target density which is itself only an approximation to the true posterior density. We obtain joint Highest Posterior Density credible sets for density estimation trees given by Li et al. (2016) approximating a density truncated to a compact subset of R^d as this is preferred to a copula construction. These trees approximate a joint posterior distribution from posterior samples using a piecewise constant function defined by sequential binary splits. We use a consistent estimator to measure of the symmetric difference between our credible set estimate and the true HPD set of the target density samples. This quality measure can be computed without the need to know the true set. We show how the true-posterior-coverage of an approximate credible set estimated for an approximate target density may be estimated in doubly intractable cases where posterior samples are not available. We illustrate our methods with simulation studies and find that our estimator is competitive with existing methods.
△ Less
Submitted 27 May, 2021; v1 submitted 26 December, 2020;
originally announced December 2020.
-
Distortion estimates for approximate Bayesian inference
Authors:
Hanwen Xing,
Geoff K. Nicholls,
Jeong Eun Lee
Abstract:
Current literature on posterior approximation for Bayesian inference offers many alternative methods. Does our chosen approximation scheme work well on the observed data? The best existing generic diagnostic tools treating this kind of question by looking at performance averaged over data space, or otherwise lack diagnostic detail. However, if the approximation is bad for most data, but good at th…
▽ More
Current literature on posterior approximation for Bayesian inference offers many alternative methods. Does our chosen approximation scheme work well on the observed data? The best existing generic diagnostic tools treating this kind of question by looking at performance averaged over data space, or otherwise lack diagnostic detail. However, if the approximation is bad for most data, but good at the observed data, then we may discard a useful approximation. We give graphical diagnostics for posterior approximation at the observed data. We estimate a "distortion map" that acts on univariate marginals of the approximate posterior to move them closer to the exact posterior, without recourse to the exact posterior.
△ Less
Submitted 19 June, 2020;
originally announced June 2020.
-
Semi-Modular Inference: enhanced learning in multi-modular models by tempering the influence of components
Authors:
Chris U. Carmona,
Geoff K. Nicholls
Abstract:
Bayesian statistical inference loses predictive optimality when generative models are misspecified.
Working within an existing coherent loss-based generalisation of Bayesian inference, we show existing Modular/Cut-model inference is coherent, and write down a new family of Semi-Modular Inference (SMI) schemes, indexed by an influence parameter, with Bayesian inference and Cut-models as special c…
▽ More
Bayesian statistical inference loses predictive optimality when generative models are misspecified.
Working within an existing coherent loss-based generalisation of Bayesian inference, we show existing Modular/Cut-model inference is coherent, and write down a new family of Semi-Modular Inference (SMI) schemes, indexed by an influence parameter, with Bayesian inference and Cut-models as special cases. We give a meta-learning criterion and estimation procedure to choose the inference scheme. This returns Bayesian inference when there is no misspecification.
The framework applies naturally to Multi-modular models. Cut-model inference allows directed information flow from well-specified modules to misspecified modules, but not vice versa. An existing alternative power posterior method gives tunable but undirected control of information flow, improving prediction in some settings. In contrast, SMI allows tunable and directed information flow between modules.
We illustrate our methods on two standard test cases from the literature and a motivating archaeological data set.
△ Less
Submitted 15 March, 2020;
originally announced March 2020.
-
Large Scale Tensor Regression using Kernels and Variational Inference
Authors:
Robert Hu,
Geoff K. Nicholls,
Dino Sejdinovic
Abstract:
We outline an inherent weakness of tensor factorization models when latent factors are expressed as a function of side information and propose a novel method to mitigate this weakness. We coin our method \textit{Kernel Fried Tensor}(KFT) and present it as a large scale forecasting tool for high dimensional data. Our results show superior performance against \textit{LightGBM} and \textit{Field Awar…
▽ More
We outline an inherent weakness of tensor factorization models when latent factors are expressed as a function of side information and propose a novel method to mitigate this weakness. We coin our method \textit{Kernel Fried Tensor}(KFT) and present it as a large scale forecasting tool for high dimensional data. Our results show superior performance against \textit{LightGBM} and \textit{Field Aware Factorization Machines}(FFM), two algorithms with proven track records widely used in industrial forecasting. We also develop a variational inference framework for KFT and associate our forecasts with calibrated uncertainty estimates on three large scale datasets. Furthermore, KFT is empirically shown to be robust against uninformative side information in terms of constants and Gaussian noise.
△ Less
Submitted 11 February, 2020;
originally announced February 2020.
-
Calibration procedures for approximate Bayesian credible sets
Authors:
Jeong Eun Lee,
Geoff K. Nicholls,
Robin J. Ryder
Abstract:
We develop and apply two calibration procedures for checking the coverage of approximate Bayesian credible sets including intervals estimated using Monte Carlo methods. The user has an ideal prior and likelihood, but generates a credible set for an approximate posterior which is not proportional to the product of ideal likelihood and prior. We estimate the realised posterior coverage achieved by t…
▽ More
We develop and apply two calibration procedures for checking the coverage of approximate Bayesian credible sets including intervals estimated using Monte Carlo methods. The user has an ideal prior and likelihood, but generates a credible set for an approximate posterior which is not proportional to the product of ideal likelihood and prior. We estimate the realised posterior coverage achieved by the approximate credible set. This is the coverage of the unknown ``true'' parameter if the data are a realisation of the user's ideal observation model conditioned on the parameter, and the parameter is a draw from the user's ideal prior. In one approach we estimate the posterior coverage at the data by making a semi-parametric logistic regression of binary coverage outcomes on simulated data against summary statistics evaluated on simulated data. In another we use Importance Sampling from the approximate posterior, windowing simulated data to fall close to the observed data. We illustrate our methods on four examples.
△ Less
Submitted 8 April, 2019; v1 submitted 15 October, 2018;
originally announced October 2018.
-
Lateral transfer in Stochastic Dollo models
Authors:
Luke J. Kelly,
Geoff K. Nicholls
Abstract:
Lateral transfer, a process whereby species exchange evolutionary traits through non-ancestral relationships, is a frequent source of model misspecification in phylogenetic inference. Lateral transfer obscures the phylogenetic signal in the data as the histories of affected traits are mosaics of the overall phylogeny. We control for the effect of lateral transfer in a Stochastic Dollo model and a…
▽ More
Lateral transfer, a process whereby species exchange evolutionary traits through non-ancestral relationships, is a frequent source of model misspecification in phylogenetic inference. Lateral transfer obscures the phylogenetic signal in the data as the histories of affected traits are mosaics of the overall phylogeny. We control for the effect of lateral transfer in a Stochastic Dollo model and a Bayesian setting. Our likelihood is highly intractable as the parameters are the solution of a sequence of large systems of differential equations representing the expected evolution of traits along a tree. We illustrate our method on a data set of lexical traits in Eastern Polynesian languages and obtain an improved fit over the corresponding model without lateral transfer.
△ Less
Submitted 16 March, 2017; v1 submitted 28 January, 2016;
originally announced January 2016.
-
Scalable Bayesian Inference for the Inverse Temperature of a Hidden Potts Model
Authors:
Matthew T. Moores,
Geoff K. Nicholls,
Anthony N. Pettitt,
Kerrie Mengersen
Abstract:
The inverse temperature parameter of the Potts model governs the strength of spatial cohesion and therefore has a major influence over the resulting model fit. A difficulty arises from the dependence of an intractable normalising constant on the value of this parameter and thus there is no closed-form solution for sampling from the posterior distribution directly. There are a variety of computatio…
▽ More
The inverse temperature parameter of the Potts model governs the strength of spatial cohesion and therefore has a major influence over the resulting model fit. A difficulty arises from the dependence of an intractable normalising constant on the value of this parameter and thus there is no closed-form solution for sampling from the posterior distribution directly. There are a variety of computational approaches for sampling from the posterior without evaluating the normalising constant, including the exchange algorithm and approximate Bayesian computation (ABC). A serious drawback of these algorithms is that they do not scale well for models with a large state space, such as images with a million or more pixels. We introduce a parametric surrogate model, which approximates the score function using an integral curve. Our surrogate model incorporates known properties of the likelihood, such as heteroskedasticity and critical temperature. We demonstrate this method using synthetic data as well as remotely-sensed imagery from the Landsat-8 satellite. We achieve up to a hundredfold improvement in the elapsed runtime, compared to the exchange algorithm or ABC. An open source implementation of our algorithm is available in the R package "bayesImageS."
△ Less
Submitted 17 August, 2018; v1 submitted 27 March, 2015;
originally announced March 2015.
-
Coupled MCMC with a randomized acceptance probability
Authors:
Geoff K. Nicholls,
Colin Fox,
Alexis Muir Watt
Abstract:
We consider Metropolis Hastings MCMC in cases where the log of the ratio of target distributions is replaced by an estimator. The estimator is based on m samples from an independent online Monte Carlo simulation. Under some conditions on the distribution of the estimator the process resembles Metropolis Hastings MCMC with a randomized transition kernel. When this is the case there is a correction…
▽ More
We consider Metropolis Hastings MCMC in cases where the log of the ratio of target distributions is replaced by an estimator. The estimator is based on m samples from an independent online Monte Carlo simulation. Under some conditions on the distribution of the estimator the process resembles Metropolis Hastings MCMC with a randomized transition kernel. When this is the case there is a correction to the estimated acceptance probability which ensures that the target distribution remains the equilibrium distribution. The simplest versions of the Penalty Method of Ceperley and Dewing (1999), the Universal Algorithm of Ball et al. (2003) and the Single Variable Exchange algorithm of Murray et al. (2006) are special cases. In many applications of interest the correction terms cannot be computed. We consider approximate versions of the algorithms. We show that on average O(m) of the samples realized by a simulation approximating a randomized chain of length n are exactly the same as those of a coupled (exact) randomized chain. Approximation biases Monte Carlo estimates with terms O(1/m) or smaller. This should be compared to the Monte Carlo error which is O(1/sqrt(n)).
△ Less
Submitted 30 May, 2012;
originally announced May 2012.
-
On building and fitting a spatio-temporal change-point model for settlement and growth at Bourewa, Fiji Islands
Authors:
Geoff K. Nicholls,
Patrick D. Nunn
Abstract:
The Bourewa beach site on the Rove Peninsula of Viti Levu is the earliest known human settlement in the Fiji Islands. How did the settlement at Bourewa develop in space and time? We have radiocarbon dates on sixty specimens, found in association with evidence for human presence, taken from pits across the site. Owing to the lack of diagnostic stratigraphy, there is no direct archaeological evidenc…
▽ More
The Bourewa beach site on the Rove Peninsula of Viti Levu is the earliest known human settlement in the Fiji Islands. How did the settlement at Bourewa develop in space and time? We have radiocarbon dates on sixty specimens, found in association with evidence for human presence, taken from pits across the site. Owing to the lack of diagnostic stratigraphy, there is no direct archaeological evidence for distinct phases of occupation through the period of interest. We give a spatio-temporal analysis of settlement at Bourewa in which the deposition rate for dated specimens plays an important role. Spatio-temporal map** of radiocarbon date intensity is confounded by uneven post-depositional thinning. We assume that the confounding processes act in such a way that the absence of dates remains informative of zero rate for the original deposition process. We model and fit the onset-field, that is, we estimate for each location across the site the time at which deposition of datable specimens began. The temporal process generating our spatial onset-field is a model of the original settlement dynamics.
△ Less
Submitted 29 June, 2010;
originally announced June 2010.
-
Missing data in a stochastic Dollo model for cognate data, and its application to the dating of Proto-Indo-European
Authors:
Robin J. Ryder,
Geoff K. Nicholls
Abstract:
Nicholls and Gray (2008) describe a phylogenetic model for trait data. They use their model to estimate branching times on Indo-European language trees from lexical data. Alekseyenko et al. (2008) extended the model and give applications in genetics. In this paper we extend the inference to handle data missing at random. When trait data are gathered, traits are thinned in a way that depends on b…
▽ More
Nicholls and Gray (2008) describe a phylogenetic model for trait data. They use their model to estimate branching times on Indo-European language trees from lexical data. Alekseyenko et al. (2008) extended the model and give applications in genetics. In this paper we extend the inference to handle data missing at random. When trait data are gathered, traits are thinned in a way that depends on both the trait and missing-data content. Nicholls and Gray (2008) treat missing records as absent traits. Hittite has 12% missing trait records. Its age is poorly predicted in their cross-validation. Our prediction is consistent with the historical record. Nicholls and Gray (2008) dropped seven languages with too much missing data. We fit all twenty four languages in the lexical data of Ringe (2002). In order to model spatial-temporal rate heterogeneity we add a catastrophe process to the model. When a language passes through a catastrophe, many traits change at the same time. We fit the full model in a Bayesian setting, via MCMC.
We validate our fit using Bayes factors to test known age constraints. We reject three of thirty historically attested constraints. Our main result is a unimodel posterior distribution for the age of Proto-Indo-European centered at 8400 years BP with 95% HPD equal 7100-9800 years BP.
△ Less
Submitted 12 August, 2009;
originally announced August 2009.
-
Dated ancestral trees from binary trait data and its application to the diversification of languages
Authors:
Geoff K. Nicholls,
Russell D. Gray
Abstract:
Binary trait data record the presence or absence of distinguishing traits in individuals. We treat the problem of estimating ancestral trees with time depth from binary trait data. Simple analysis of such data is problematic. Each homology class of traits has a unique birth event on the tree, and the birth event of a trait visible at the leaves is biased towards the leaves. We propose a model-ba…
▽ More
Binary trait data record the presence or absence of distinguishing traits in individuals. We treat the problem of estimating ancestral trees with time depth from binary trait data. Simple analysis of such data is problematic. Each homology class of traits has a unique birth event on the tree, and the birth event of a trait visible at the leaves is biased towards the leaves. We propose a model-based analysis of such data, and present an MCMC algorithm that can sample from the resulting posterior distribution. Our model is based on using a birth-death process for the evolution of the elements of sets of traits. Our analysis correctly accounts for the removal of singleton traits, which are commonly discarded in real data sets. We illustrate Bayesian inference for two binary-trait data sets which arise in historical linguistics. The Bayesian approach allows for the incorporation of information from ancestral languages. The marginal prior distribution of the root time is uniform. We present a thorough analysis of the robustness of our results to model mispecification, through analysis of predictive distributions for external data, and fitting data simulated under alternative observation models. The reconstructed ages of tree nodes are relatively robust, whilst posterior probabilities for topology are not reliable.
△ Less
Submitted 12 November, 2007;
originally announced November 2007.