Heavy-Tailed Distributions: Data, Diagnostics, and New Developments

نویسندگان

  • Roger M. Cooke
  • Daan Nieboer
چکیده

This monograph is written for the numerate nonspecialist, and hopes to serve three purposes. First it gathers mathematical material from diverse but related fields of order statistics, records, extreme value theory, majorization, regular variation and subexponentiality. All of these are relevant for understanding fat tails, but they are not, to our knowledge, brought together in a single source for the target readership. Proofs that give insight are included, but for fussy calculations the reader is referred to the excellent sources referenced in the text. Multivariate extremes are not treated. This allows us to present material spread over hundreds of pages in specialist texts in twenty pages. Chapter 5 develops new material on heavy tail diagnostics and gives more mathematical detail. Second, it presents a new measure of obesity. The most popular definitions in terms of regular variation and subexponentiality invoke putative properties that hold at infinity, and this complicates any empirical estimate. Each definition captures some but not all of the intuitions associated with tail heaviness. Chapter 5 studies two candidate indices of tail heaviness based on the tendency of the mean excess plot to collapse as data are aggregated. The probability that the largest value is more than twice the second largest has intuitive appeal but its estimator has very poor accuracy. The Obesity index is defined for a positive random variable X as: Ob(X) = P (X1 +X4 > X2 +X3|X1 ≤ X2 ≤ X3 ≤ X4) , Xi independent copies of X. For empirical distributions, obesity is defined by bootstrapping. This index reasonably captures intuitions of tail heaviness. Among its properties, if α > 1 then Ob(X) < Ob(Xα). However, it does not completely mimic the tail index of regularly varying distributions, or the extreme value index. A Weibull distribution with shape 1/4 is more obese than a Pareto distribution with tail index 1, even though this Pareto has infinite mean and the Weibull’s moments are all finite. Chapter 5 explores properties of the Obesity index. Third and most important, we hope to convince the reader that fat tail phenomena pose real problems; they are really out there and they seriously challenge our usual ways of thinking about historical averages, outliers, trends, regression coefficients and confidence bounds among many other things. Data on flood insurance claims, crop loss claims, hospital discharge bills, precipitation and damages and fatalities from natural catastrophes drive this point home. AMS classification 60-02, 62-02, 60-07.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Fitting Data and Assessing Goodness of t with Stable Distributions

There are now several reliable methods for estimating stable pa rameters from data However little attention has been paid to model veri cation i e how to assess whether the stable parameters esti mated actually do a good job of describing the data We analyze sev eral heavy tailed data sets and demonstrate diagnostics for assessing both univariate and multivariate stable ts to data sets Multivar...

متن کامل

Skew-slash distribution and its application in topics regression

In many issues of statistical modeling, the common assumption is that observations are normally distributed. In many real data applications, however, the true distribution is deviated from the normal. Thus, the main concern of most recent studies on analyzing data is to construct and the use of alternative distributions. In this regard, new classes of distributions such as slash and skew-sla...

متن کامل

نمودار شوهارت ناپارامتری رتبه علامت دار با فاصله نمونه گیری متغیر

Nonparametric control chart based on rank is used for detecting changes in median(mean). In this article ,Signed-rank control chart is considered with variable sampling interval. We compared the performance of Signed-rank with variable sampling interval (VSI-SR) to Signed-rank with Fixed Sampling interval (FSI-SR),the numerical results demonstrated the VSI feature is so useful. Bakir[1] showed ...

متن کامل

Mixture of Normal Mean-Variance of Lindley Distributions

&lrm;Abstract: In this paper, a new mixture modelling using the normal mean-variance mixture of Lindley (NMVL) distribution has been considered. The proposed model is heavy-tailed and multimodal and can be used in dealing with asymmetric data in various theoretic and applied problems. We present a feasible computationally analytical EM algorithm for computing the maximum likelihood estimates. T...

متن کامل

On Bivariate Generalized Exponential-Power Series Class of Distributions

In this paper, we introduce a new class of bivariate distributions by compounding the bivariate generalized exponential and power-series distributions. This&nbsp;new class contains the bivariate generalized exponential-Poisson, bivariate generalized exponential-logarithmic, bivariate generalized exponential-binomial and bivariate generalized exponential-negative binomial distributions as specia...

متن کامل

The Family of Scale-Mixture of Skew-Normal Distributions and Its Application in Bayesian Nonlinear Regression Models

In previous studies on fitting non-linear regression models with the symmetric structure the normality is usually assumed in the analysis of data. This choice may be inappropriate when the distribution of residual terms is asymmetric. Recently, the family of scale-mixture of skew-normal distributions is the main concern of many researchers. This family includes several skewed and heavy-tailed d...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2011