The UNIVARIATE Procedure

Formulas for Fitted Continuous Distributions

The following sections provide information about the families of parametric distributions that you can fit with the HISTOGRAM statement. Properties of these distributions are discussed by Johnson, Kotz, and Balakrishnan (1994, 1995).

Beta Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction left-parenthesis x minus theta right-parenthesis Superscript alpha minus 1 Baseline left-parenthesis sigma plus theta minus x right-parenthesis Superscript beta minus 1 Baseline Over upper B left-parenthesis alpha comma beta right-parenthesis sigma Superscript left-parenthesis alpha plus beta minus 1 right-parenthesis Baseline EndFraction 2nd Column for theta less-than x less-than theta plus sigma 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta or x greater-than-or-equal-to theta plus sigma EndLayout

where upper B left-parenthesis alpha comma beta right-parenthesis equals StartFraction normal upper Gamma left-parenthesis alpha right-parenthesis normal upper Gamma left-parenthesis beta right-parenthesis Over normal upper Gamma left-parenthesis alpha plus beta right-parenthesis EndFraction and

  • theta equals lower threshold parameter (lower endpoint parameter)

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • alpha equals shape parameter left-parenthesis alpha greater-than 0 right-parenthesis

  • beta equals shape parameter left-parenthesis beta greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

Note: This notation is consistent with that of other distributions that you can fit with the HISTOGRAM statement. However, many texts, including Johnson, Kotz, and Balakrishnan (1995), write the beta density function as

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction left-parenthesis x minus a right-parenthesis Superscript p minus 1 Baseline left-parenthesis b minus x right-parenthesis Superscript q minus 1 Baseline Over upper B left-parenthesis p comma q right-parenthesis left-parenthesis b minus a right-parenthesis Superscript p plus q minus 1 Baseline EndFraction 2nd Column for a less-than x less-than b 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to a or x greater-than-or-equal-to b EndLayout

The two parameterizations are related as follows:

  • sigma equals b minus a

  • theta equals a

  • alpha equals p

  • beta equals q

The range of the beta distribution is bounded below by a threshold parameter theta equals a and above by theta plus sigma equals b. If you specify a fitted beta curve by using the BETA option, theta must be less than the minimum data value and theta plus sigma must be greater than the maximum data value. You can specify theta and sigma with the THETA= and SIGMA= beta-options in parentheses after the keyword BETA. By default, sigma equals 1 and theta equals 0. If you specify THETA=EST and SIGMA=EST, maximum likelihood estimates are computed for theta and sigma. However, three- and four-parameter maximum likelihood estimation does not always converge.

In addition, you can specify alpha and beta with the ALPHA= and BETA= beta-options, respectively. By default, the procedure calculates maximum likelihood estimates for alpha and beta. For example, to fit a beta density curve to a set of data bounded below by 32 and above by 212 with maximum likelihood estimates for alpha and beta, use the following statement:

histogram Length / beta(theta=32 sigma=180);

The beta distributions are also referred to as Pearson Type I or II distributions. These include the power function distribution (beta equals 1), the arc sine distribution (alpha equals beta equals one-half), and the generalized arc sine distributions (alpha plus beta equals 1, beta not-equals one-half).

You can use the DATA step function QUANTILE to compute beta quantiles and the DATA step function CDF to compute beta probabilities.

Exponential Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction h v Over sigma EndFraction exp left-parenthesis minus left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis right-parenthesis 2nd Column for x greater-than-or-equal-to theta 2nd Row 1st Column 0 2nd Column for x less-than theta EndLayout

where

  • theta equals threshold parameter

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The threshold parameter theta must be less than or equal to the minimum data value. You can specify theta with the THRESHOLD= exponential-option. By default, theta equals 0. If you specify THETA=EST, a maximum likelihood estimate is computed for theta. In addition, you can specify sigma with the SCALE= exponential-option. By default, the procedure calculates a maximum likelihood estimate for sigma. Note that some authors define the scale parameter as StartFraction 1 Over sigma EndFraction.

The exponential distribution is a special case of both the gamma distribution (with alpha equals 1) and the Weibull distribution (with c equals 1). A related distribution is the extreme value distribution. If upper Y equals exp left-parenthesis negative upper X right-parenthesis has an exponential distribution, then X has an extreme value distribution.

You can use the DATA step function QUANTILE to compute exponential quantiles and the DATA step function CDF to compute exponential probabilities.

Gamma Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction h v Over normal upper Gamma left-parenthesis alpha right-parenthesis sigma EndFraction left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis Superscript alpha minus 1 Baseline exp left-parenthesis minus left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis right-parenthesis 2nd Column for x greater-than theta 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta EndLayout

where

  • theta equals threshold parameter

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • alpha equals shape parameter left-parenthesis alpha greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The threshold parameter theta must be less than the minimum data value. You can specify theta with the THRESHOLD= gamma-option. By default, theta equals 0. If you specify THETA=EST, a maximum likelihood estimate is computed for theta. In addition, you can specify sigma and alpha with the SCALE= and ALPHA= gamma-options. By default, the procedure calculates maximum likelihood estimates for sigma and alpha.

The gamma distributions are also referred to as Pearson Type III distributions, and they include the chi-square, exponential, and Erlang distributions. The probability density function for the chi-square distribution is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartStartFraction 1 OverOver 2 normal upper Gamma left-parenthesis StartFraction nu Over 2 EndFraction right-parenthesis EndEndFraction left-parenthesis StartFraction x Over 2 EndFraction right-parenthesis Superscript StartFraction nu Over 2 EndFraction minus 1 Baseline exp left-parenthesis minus StartFraction x Over 2 EndFraction right-parenthesis 2nd Column for x greater-than 0 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to 0 EndLayout

Notice that this is a gamma distribution with alpha equals StartFraction nu Over 2 EndFraction, sigma equals 2, and theta equals 0. The exponential distribution is a gamma distribution with alpha equals 1, and the Erlang distribution is a gamma distribution with alpha being a positive integer. A related distribution is the Rayleigh distribution. If upper R equals StartFraction max left-parenthesis upper X 1 comma ellipsis comma upper X Subscript n Baseline right-parenthesis Over min left-parenthesis upper X 1 comma ellipsis comma upper X Subscript n Baseline right-parenthesis EndFraction where the upper X Subscript i’s are independent chi Subscript nu Superscript 2 variables, then log upper R is distributed with a chi Subscript nu distribution having a probability density function of

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column left-bracket 2 Superscript StartFraction nu Over 2 EndFraction minus 1 Baseline normal upper Gamma left-parenthesis StartFraction nu Over 2 EndFraction right-parenthesis right-bracket Superscript negative 1 Baseline x Superscript nu minus 1 Baseline exp left-parenthesis minus StartFraction x squared Over 2 EndFraction right-parenthesis 2nd Column for x greater-than 0 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to 0 EndLayout

If nu equals 2, the preceding distribution is referred to as the Rayleigh distribution.

You can use the DATA step function QUANTILE to compute gamma quantiles and the DATA step function CDF to compute gamma probabilities.

Gumbel Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartFraction h v Over sigma EndFraction e Superscript minus left-parenthesis x minus mu right-parenthesis slash sigma Baseline exp left-parenthesis minus e Superscript minus left-parenthesis x minus mu right-parenthesis slash sigma Baseline right-parenthesis

where

  • mu equals location parameter

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

You can specify mu and sigma with the MU= and SIGMA= Gumbel-options, respectively. By default, the procedure calculates maximum likelihood estimates for these parameters.

Note: The Gumbel distribution is also referred to as Type 1 extreme value distribution.

Note: The random variable X has Gumbel (Type 1 extreme value) distribution if and only if e Superscript upper X has Weibull distribution and exp left-parenthesis left-parenthesis upper X minus mu right-parenthesis slash sigma right-parenthesis has standard exponential distribution.

Inverse Gaussian Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v left-parenthesis StartFraction lamda Over 2 pi x cubed EndFraction right-parenthesis Superscript 1 slash 2 Baseline exp left-parenthesis minus StartFraction lamda Over 2 mu squared x EndFraction left-parenthesis x minus mu right-parenthesis squared right-parenthesis 2nd Column for x greater-than 0 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to 0 EndLayout

where

  • mu equals location parameter left-parenthesis mu greater-than 0 right-parenthesis

  • lamda equals shape parameter left-parenthesis lamda greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The location parameter mu has to be greater then zero. You can specify mu with the MU= iGauss-option. In addition, you can specify shape parameter lamda with LAMBDA= iGauss-option. By default, the procedure calculates maximum likelihood estimates for mu and lamda.

Note: The special case where mu equals 1 and lamda equals phi corresponds to the Wald distribution.

You can use the DATA step function QUANTILE to compute inverse Gaussian quantiles and the DATA step function CDF to compute inverse Gaussian probabilities.

Lognormal Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction h v Over sigma StartRoot 2 pi EndRoot left-parenthesis x minus theta right-parenthesis EndFraction exp left-parenthesis minus StartFraction left-parenthesis log left-parenthesis x minus theta right-parenthesis minus zeta right-parenthesis squared Over 2 sigma squared EndFraction right-parenthesis 2nd Column for x greater-than theta 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta EndLayout

where

  • theta equals threshold parameter

  • zeta equals scale parameter left-parenthesis negative normal infinity less-than zeta less-than normal infinity right-parenthesis

  • sigma equals shape parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The threshold parameter theta must be less than the minimum data value. You can specify theta with the THRESHOLD= lognormal-option. By default, theta equals 0. If you specify THETA=EST, a maximum likelihood estimate is computed for theta. You can specify zeta and sigma with the SCALE= and SHAPE= lognormal-options, respectively. By default, the procedure calculates estimates for these parameters as

ModifyingAbove zeta With caret equals StartFraction sigma-summation Underscript i equals 1 Overscript n Endscripts log left-parenthesis x Subscript i Baseline minus theta right-parenthesis Over n EndFraction

and

ModifyingAbove sigma With caret equals StartRoot StartFraction sigma-summation Underscript i equals 1 Overscript n Endscripts left-parenthesis log left-parenthesis x Subscript i Baseline minus theta right-parenthesis minus zeta right-parenthesis squared Over n minus 1 EndFraction EndRoot

Note: The lognormal distribution is also referred to as the upper S Subscript upper L distribution in the Johnson system of distributions.

Note: This book uses sigma to denote the shape parameter of the lognormal distribution, whereas sigma is used to denote the scale parameter of the other distributions. The use of sigma to denote the lognormal shape parameter is based on the fact that StartFraction 1 Over sigma EndFraction left-parenthesis log left-parenthesis upper X minus theta right-parenthesis minus zeta right-parenthesis has a standard normal distribution if X is lognormally distributed. Based on this relationship, you can use the DATA step function PROBIT to compute lognormal quantiles and the DATA step function PROBNORM to compute probabilities.

Normal Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout 1st Row 1st Column StartFraction h v Over sigma StartRoot 2 pi EndRoot EndFraction exp left-parenthesis minus one-half left-parenthesis StartFraction x minus mu Over sigma EndFraction right-parenthesis squared right-parenthesis 2nd Column for negative normal infinity less-than x less-than normal infinity EndLayout

where

  • mu equals mean

  • sigma equals standard deviation left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

You can specify mu and sigma with the MU= and SIGMA= normal-options, respectively. By default, the procedure estimates mu with the sample mean and sigma with the sample standard deviation.

You can use the DATA step function QUANTILE to compute beta quantiles and the DATA step function CDF to compute normal probabilities.

Note: The normal distribution is also referred to as the upper S Subscript upper N distribution in the Johnson system of distributions.

Generalized Pareto Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction h v Over sigma EndFraction left-parenthesis 1 minus alpha left-parenthesis x minus theta right-parenthesis slash sigma right-parenthesis Superscript 1 slash alpha minus 1 Baseline 2nd Column if alpha not-equals 0 2nd Row 1st Column StartFraction h v Over sigma EndFraction exp left-parenthesis negative x slash sigma right-parenthesis 2nd Column if alpha equals 0 EndLayout

where

  • theta equals threshold parameter

  • alpha equals shape parameter

  • sigma equals shape parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The support of the distribution is x greater-than theta for alpha less-than-or-equal-to 0 and theta less-than x less-than sigma slash alpha for alpha greater-than 0.

Note: Special cases of Pareto distribution with alpha equals 0 and alpha equals 1 correspond respectively to the exponential distribution with mean sigma and uniform distribution on the interval left-parenthesis theta comma sigma right-parenthesis.

The threshold parameter theta must be less than the minimum data value. You can specify theta with the THETA= Pareto-option. By default, theta equals 0. You can also specify alpha and sigma with the ALPHA= and SIGMA= Pareto-options,respectively. By default, the procedure calculates maximum likelihood estimates for these parameters.

Note: Maximum likelihood estimation of the parameters works well if alpha less-than one-half, but not otherwise. In this case the estimators are asymptotically normal and asymptotically efficient. The asymptotic normal distribution of the maximum likelihood estimates has mean left-parenthesis alpha comma sigma right-parenthesis and variance-covariance matrix

StartFraction 1 Over n EndFraction Start 2 By 2 Matrix 1st Row 1st Column left-parenthesis 1 minus alpha right-parenthesis squared 2nd Column sigma left-parenthesis 1 minus alpha right-parenthesis 2nd Row 1st Column sigma left-parenthesis 1 minus alpha right-parenthesis 2nd Column 2 sigma squared left-parenthesis 1 minus alpha right-parenthesis EndMatrix period

Note: If no local minimum is found in the region

StartSet alpha less-than 0 comma sigma greater-than 0 EndSet union StartSet 0 less-than alpha less-than-or-equal-to 1 comma sigma slash alpha greater-than max left-parenthesis upper X Subscript i Baseline right-parenthesis EndSet comma

there is no maximum likelihood estimator. More details on how to find maximum likelihood estimators and a suggested algorithm can be found in Grimshaw (1993).

Power Function Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction alpha Over sigma EndFraction left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis Superscript alpha minus 1 Baseline 2nd Column for theta less-than x less-than theta plus sigma 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta or x greater-than-or-equal-to theta plus sigma EndLayout

where

  • theta equals lower threshold parameter (lower endpoint parameter)

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • alpha equals shape parameter left-parenthesis alpha greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

Note: This notation is consistent with that of other distributions that you can fit with the HISTOGRAM statement. However, many texts, including Johnson, Kotz, and Balakrishnan (1995), write the density function of power function distribution as

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction p Over b minus a EndFraction left-parenthesis StartFraction x minus a Over b minus a EndFraction right-parenthesis Superscript p minus 1 Baseline 2nd Column for a less-than x less-than b 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to a or x greater-than-or-equal-to b EndLayout

The two parameterizations are related as follows:

  • sigma equals b minus a

  • theta equals a

  • alpha equals p

Note: The family of power function distributions is subclass of beta distribution with density function

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction left-parenthesis x minus theta right-parenthesis Superscript alpha minus 1 Baseline left-parenthesis sigma plus theta minus x right-parenthesis Superscript beta minus 1 Baseline Over upper B left-parenthesis alpha comma beta right-parenthesis sigma Superscript left-parenthesis alpha plus beta minus 1 right-parenthesis Baseline EndFraction 2nd Column for theta less-than x less-than theta plus sigma 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta or x greater-than-or-equal-to theta plus sigma EndLayout

where upper B left-parenthesis alpha comma beta right-parenthesis equals StartFraction normal upper Gamma left-parenthesis alpha right-parenthesis normal upper Gamma left-parenthesis beta right-parenthesis Over normal upper Gamma left-parenthesis alpha plus beta right-parenthesis EndFraction with parameter beta equals 1. Therefore, all properties and estimation procedures of beta distribution apply.

The range of the power function distribution is bounded below by a threshold parameter theta equals a and above by theta plus sigma equals b. If you specify a fitted power function curve by using the POWER option, theta must be less than the minimum data value and theta plus sigma must be greater than the maximum data value. You can specify theta and sigma with the THETA= and SIGMA= power-options in parentheses after the keyword POWER. By default, sigma equals 1 and theta equals 0. If you specify THETA=EST and SIGMA=EST, maximum likelihood estimates are computed for theta and sigma. However, three-parameter maximum likelihood estimation does not always converge.

In addition, you can specify alpha with the ALPHA= power-option. By default, the procedure calculates maximum likelihood estimate for alpha. For example, to fit a power function density curve to a set of data bounded below by 32 and above by 212 with maximum likelihood estimate for alpha, use the following statement:

histogram Length / power(theta=32 sigma=180);

Rayleigh Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction x minus theta Over sigma squared EndFraction e Superscript minus left-parenthesis x minus theta right-parenthesis squared slash left-parenthesis 2 sigma squared right-parenthesis Baseline 2nd Column for x greater-than-or-equal-to theta 2nd Row 1st Column 0 2nd Column for x less-than theta EndLayout

where

  • theta equals lower threshold parameter (lower endpoint parameter)

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

Note: The Rayleigh distribution is Weibull distribution with density function

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction k Over lamda EndFraction left-parenthesis StartFraction x minus theta Over lamda EndFraction right-parenthesis Superscript k minus 1 Baseline exp left-parenthesis minus left-parenthesis StartFraction x minus theta Over lamda EndFraction right-parenthesis Superscript k Baseline right-parenthesis 2nd Column for x greater-than-or-equal-to theta 2nd Row 1st Column 0 2nd Column for x less-than theta EndLayout

and with shape parameter k equals 2 and scale parameter lamda equals StartRoot 2 EndRoot sigma.

The threshold parameter theta must be less than the minimum data value. You can specify theta with the THETA= Rayleigh-option. By default, theta equals 0. In addition you can specify sigma with the SIGMA= Rayleigh-option. By default, the procedure calculates maximum likelihood estimate for sigma.

For example, to fit a Rayleigh density curve to a set of data bounded below by 32 with maximum likelihood estimate for sigma, use the following statement:

histogram Length / rayleigh(theta=32);

Johnson upper S Subscript upper B Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction delta h v Over sigma StartRoot 2 pi EndRoot EndFraction left-bracket left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis left-parenthesis 1 minus StartFraction x minus theta Over sigma EndFraction right-parenthesis right-bracket Superscript negative 1 Baseline times 2nd Column Blank 2nd Row 1st Column exp left-bracket minus one-half left-parenthesis gamma plus delta log left-parenthesis StartFraction x minus theta Over theta plus sigma minus x EndFraction right-parenthesis right-parenthesis squared right-bracket 2nd Column for theta less-than x less-than theta plus sigma 3rd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta or x greater-than-or-equal-to theta plus sigma EndLayout

where

  • theta equals threshold parameter left-parenthesis negative normal infinity less-than theta less-than normal infinity right-parenthesis

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • delta equals shape parameter left-parenthesis delta greater-than 0 right-parenthesis

  • gamma equals shape parameter left-parenthesis negative normal infinity less-than gamma less-than normal infinity right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The upper S Subscript upper B distribution is bounded below by the parameter theta and above by the value theta plus sigma. The parameter theta must be less than the minimum data value. You can specify theta with the THETA= upper S Subscript upper B-option, or you can request that theta be estimated with the THETA = EST upper S Subscript upper B-option. The default value for theta is zero. The sum theta plus sigma must be greater than the maximum data value. The default value for sigma is one. You can specify sigma with the SIGMA= upper S Subscript upper B-option, or you can request that sigma be estimated with the SIGMA = EST upper S Subscript upper B-option.

By default, the method of percentiles given by Slifker and Shapiro (1980) is used to estimate the parameters. This method is based on four data percentiles, denoted by x Subscript minus 3 z, x Subscript negative z, x Subscript z, and x Subscript 3 z, which correspond to the four equally spaced percentiles of a standard normal distribution, denoted by minus 3 z, negative z, z, and 3 z, under the transformation

z equals gamma plus delta log left-parenthesis StartFraction x minus theta Over theta plus sigma minus x EndFraction right-parenthesis

The default value of z is 0.524. The results of the fit are dependent on the choice of z, and you can specify other values with the FITINTERVAL= option (specified in parentheses after the SB option). If you use the method of percentiles, you should select a value of z that corresponds to percentiles which are critical to your application.

The following values are computed from the data percentiles:

StartLayout 1st Row 1st Column m 2nd Column equals 3rd Column x Subscript 3 z Baseline minus x Subscript z 2nd Row 1st Column n 2nd Column equals 3rd Column x Subscript negative z Baseline minus x Subscript minus 3 z 3rd Row 1st Column p 2nd Column equals 3rd Column x Subscript z Baseline minus x Subscript negative z EndLayout

It was demonstrated by Slifker and Shapiro (1980) that

StartLayout 1st Row 1st Column StartFraction m n Over p squared EndFraction greater-than 1 2nd Column for any upper S Subscript upper U Baseline distribution 2nd Row 1st Column StartFraction m n Over p squared EndFraction less-than 1 2nd Column for any upper S Subscript upper B Baseline distribution 3rd Row 1st Column StartFraction m n Over p squared EndFraction equals 1 2nd Column for any upper S Subscript upper L Baseline left-parenthesis lognormal right-parenthesis distribution EndLayout

A tolerance interval around one is used to discriminate among the three families with this ratio criterion. You can specify the tolerance with the FITTOLERANCE= option (specified in parentheses after the SB option). The default tolerance is 0.01. Assuming that the criterion satisfies the inequality

StartFraction m n Over p squared EndFraction less-than 1 minus tolerance

the parameters of the upper S Subscript upper B distribution are computed using the explicit formulas derived by Slifker and Shapiro (1980).

If you specify FITMETHOD = MOMENTS (in parentheses after the SB option), the method of moments is used to estimate the parameters. If you specify FITMETHOD = MLE (in parentheses after the SB option), the method of maximum likelihood is used to estimate the parameters. Note that maximum likelihood estimates might not always exist. Refer to Bowman and Shenton (1983) for discussion of methods for fitting Johnson distributions.

Johnson upper S Subscript upper U Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column StartFraction delta h v Over sigma StartRoot 2 pi EndRoot EndFraction StartFraction 1 Over StartRoot 1 plus left-parenthesis left-parenthesis x minus theta right-parenthesis slash sigma right-parenthesis squared EndRoot EndFraction times 2nd Column Blank 2nd Row 1st Column exp left-bracket minus one-half left-parenthesis gamma plus delta hyperbolic sine Superscript negative 1 Baseline left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis right-parenthesis squared right-bracket 2nd Column for x greater-than theta 3rd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta EndLayout

where

  • theta equals location parameter left-parenthesis negative normal infinity less-than theta less-than normal infinity right-parenthesis

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • delta equals shape parameter left-parenthesis delta greater-than 0 right-parenthesis

  • gamma equals shape parameter left-parenthesis negative normal infinity less-than gamma less-than normal infinity right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

You can specify the parameters with the THETA=, SIGMA=, DELTA=, and GAMMA= upper S Subscript upper U-options, which are enclosed in parentheses after the SU option. If you do not specify these parameters, they are estimated.

By default, the method of percentiles given by Slifker and Shapiro (1980) is used to estimate the parameters. This method is based on four data percentiles, denoted by x Subscript minus 3 z, x Subscript negative z, x Subscript z, and x Subscript 3 z, which correspond to the four equally spaced percentiles of a standard normal distribution, denoted by minus 3 z, negative z, z, and 3 z, under the transformation

z equals gamma plus delta hyperbolic sine Superscript negative 1 Baseline left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis

The default value of z is 0.524. The results of the fit are dependent on the choice of z, and you can specify other values with the FITINTERVAL= option (specified in parentheses after the SU option). If you use the method of percentiles, you should select a value of z that corresponds to percentiles that are critical to your application.

The following values are computed from the data percentiles:

StartLayout 1st Row 1st Column m 2nd Column equals 3rd Column x Subscript 3 z Baseline minus x Subscript z 2nd Row 1st Column n 2nd Column equals 3rd Column x Subscript negative z Baseline minus x Subscript minus 3 z 3rd Row 1st Column p 2nd Column equals 3rd Column x Subscript z Baseline minus x Subscript negative z EndLayout

It was demonstrated by Slifker and Shapiro (1980) that

StartLayout 1st Row 1st Column StartFraction m n Over p squared EndFraction greater-than 1 2nd Column for any upper S Subscript upper U Baseline distribution 2nd Row 1st Column StartFraction m n Over p squared EndFraction less-than 1 2nd Column for any upper S Subscript upper B Baseline distribution 3rd Row 1st Column StartFraction m n Over p squared EndFraction equals 1 2nd Column for any upper S Subscript upper L Baseline left-parenthesis lognormal right-parenthesis distribution EndLayout

A tolerance interval around one is used to discriminate among the three families with this ratio criterion. You can specify the tolerance with the FITTOLERANCE= option (specified in parentheses after the SU option). The default tolerance is 0.01. Assuming that the criterion satisfies the inequality

StartFraction m n Over p squared EndFraction greater-than 1 plus tolerance

the parameters of the upper S Subscript upper U distribution are computed using the explicit formulas derived by Slifker and Shapiro (1980).

If you specify FITMETHOD = MOMENTS (in parentheses after the SU option), the method of moments is used to estimate the parameters. If you specify FITMETHOD = MLE (in parentheses after the SU option), the method of maximum likelihood is used to estimate the parameters. Note that maximum likelihood estimates do not always exist. Refer to Bowman and Shenton (1983) for discussion of methods for fitting Johnson distributions.

Weibull Distribution

The fitted density function is

p left-parenthesis x right-parenthesis equals StartLayout Enlarged left-brace 1st Row 1st Column h v StartFraction c Over sigma EndFraction left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis Superscript c minus 1 Baseline exp left-parenthesis minus left-parenthesis StartFraction x minus theta Over sigma EndFraction right-parenthesis Superscript c Baseline right-parenthesis 2nd Column for x greater-than theta 2nd Row 1st Column 0 2nd Column for x less-than-or-equal-to theta EndLayout

where

  • theta equals threshold parameter

  • sigma equals scale parameter left-parenthesis sigma greater-than 0 right-parenthesis

  • c equals shape parameter left-parenthesis c greater-than 0 right-parenthesis

  • h equals width of histogram interval

  • v equals vertical scaling factor

and

v equals StartLayout Enlarged left-brace 1st Row 1st Column n 2nd Column the sample size comma for VSCALE equals COUNT 2nd Row 1st Column 100 2nd Column for VSCALE equals PERCENT 3rd Row 1st Column 1 2nd Column for VSCALE equals PROPORTION EndLayout

The threshold parameter theta must be less than the minimum data value. You can specify theta with the THRESHOLD= Weibull-option. By default, theta equals 0. If you specify THETA=EST, a maximum likelihood estimate is computed for theta. You can specify sigma and c with the SCALE= and SHAPE= Weibull-options, respectively. By default, the procedure calculates maximum likelihood estimates for sigma and c.

The exponential distribution is a special case of the Weibull distribution where c equals 1.

You can use the DATA step function QUANTILE to compute Weibull quantiles and the DATA step function CDF to compute Weibull probabilities.

Last updated: April 10, 2023