The SEVSELECT Procedure
A Simple Example of Fitting Predefined Distributions
The simplest way to use PROC SEVSELECT is to fit all the predefined distributions to a set of values and let the procedure identify the best-fitting distribution.
Consider a lognormal distribution, whose probability density function (PDF) f and cumulative distribution function (CDF) F are as follows, respectively, where denotes the CDF of the standard normal distribution:
The following DATA step statements simulate a sample from a lognormal distribution with population parameters and
, and store the sample in the variable
Y of a data set Work.Test_sev1:
/*------------- Simple Lognormal Example -------------*/
data test_sev1(keep=y label='Simple Lognormal Sample');
call streaminit(45678);
label y='Response Variable';
Mu = 1.5;
Sigma = 0.25;
do n = 1 to 100;
y = exp(Mu) * rand('LOGNORMAL')**Sigma;
output;
end;
run;
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 4, Shared Concepts.
You can load the data set Work.Test_sev1 into a data table in your CAS session by using your CAS engine libref with the following DATA step:
data mycas.test_sev1;
set test_sev1;
run;
These statements assume that your CAS engine libref is named mycas, as in the section Using CAS Sessions and CAS Engine Librefs, but you can substitute any appropriately defined CAS engine libref.
The following statements fit all the predefined distribution models to the values of Y and identify the best distribution according to the corrected Akaike’s information criterion (AICC):
proc sevselect data=mycas.test_sev1 crit=aicc
plots(histogram kernel)=all; /*(cdf pdf);*/
loss y;
dist _predefined_;
run;
The PROC SEVSELECT statement specifies the input data table along with the model selection criterion. The PLOTS= option displays comparative plots of the probability density function (PDF) and the cumulative distribution function (CDF) for all candidate distributions. The HISTOGRAM and KERNEL options request that the PDF plot contain histogram and kernel density estimates, respectively, which are nonparametric estimates of the density function. The LOSS statement specifies the variable to be modeled, and the DIST statement with the _PREDEFINED_ keyword specifies that all the predefined distribution models be fitted.
Some of the default output that is displayed by this step is shown in Figure 1 through Figure 5. First, information about the input data table is displayed followed by the "Model Selection" table, as shown in Figure 1. The model selection table displays the convergence status, the value of the selection criterion, and the selection status for each of the candidate models. The Converged column indicates whether the estimation process for a particular distribution model has converged, might have converged, or failed. The Selected column indicates whether a particular distribution has the best fit for the data according to the selection criterion. For this example, the lognormal distribution model is selected, because it has the lowest value for the selection criterion.
Figure 1: Data Table Information and Model Selection Table
| Input Data Table | |
|---|---|
| Name | TEST_SEV1 |
| Caslib | << active caslib >> |
| Model Selection | |||
|---|---|---|---|
| Distribution | Converged | AICC | Selected |
| Burr | Yes | 322.50845 | No |
| Exp | Yes | 508.12287 | No |
| Gamma | Yes | 320.50264 | No |
| Igauss | Yes | 319.61652 | No |
| Logn | Yes | 319.56579 | Yes |
| Pareto | Yes | 510.28172 | No |
| Gpd | Yes | 510.20576 | No |
| Weibull | Yes | 334.82373 | No |
Next, two comparative plots are prepared. These plots enable you to visually verify how the models differ from each other and from the nonparametric estimates. The plot in Figure 2 displays the CDF estimates of all the models and the estimates of the empirical distribution function (EDF). The CDF plot indicates that the Exp (exponential), Pareto, and Gpd (generalized Pareto) distributions are a poor fit as compared to the EDF estimate. The Weibull distribution is also a poor fit, although not as poor as exponential, Pareto, and generalized Pareto. The other four distributions seem to be quite close to each other and to the EDF estimate.
Figure 2: Comparison of EDF and CDF Estimates of the Fitted Models

The plot in Figure 3 displays the PDF estimates of all the models and the nonparametric kernel and histogram estimates. The PDF plot enables better visual comparison between the Burr, gamma, Igauss (inverse Gaussian), and Logn (lognormal) models. The Burr and gamma models differ significantly from the Igauss and Logn models in the central portion of the range of Y values, whereas the latter two models fit the data almost identically. This provides a visual confirmation of the information in the "Model Selection" table in Figure 1, which indicates that the AICC values of the Igauss and Logn models are very close.
Figure 3: Comparison of PDF Estimates of the Fitted Models

Note: When you request plots in a PROC SEVSELECT step that uses more than one worker node, the plots are created by merging empirical quantiles that each worker node computes by using the sample that it uses to compute the EDF-based fit statistics. So the plots are approximate. For more accurate visualization, it is recommended that you request plots only when you run PROC SEVSELECT in a CAS session that uses no more than one worker node.
The comparative plots are followed by the estimation information for each of the candidate models. The information for the lognormal model, which is the best-fitting model, is shown in Figure 4. The first table displays a summary of the distribution. The second table displays the convergence status. This is followed by a summary of the optimization process that indicates the technique used, the number of iterations, the number of times the objective function was evaluated, and the log likelihood that is obtained at the end of the optimization. Because the model with lognormal distribution has converged, PROC SEVSELECT displays its statistics of fit and parameter estimates. The estimates of Mu=1.49605 and Sigma=0.26243 are quite close to the population parameters of Mu=1.5 and Sigma=0.25 from which the sample was generated. The p-value for each estimate indicates the rejection of the null hypothesis that the estimate is 0, implying that both the estimates are significantly different from 0.
Figure 4: Estimation Details for the Lognormal Model
| Model Information | |
|---|---|
| Distribution | Logn |
| Description | Lognormal Distribution |
| Distribution Parameters | 2 |
| Convergence Status |
|---|
| Convergence criterion (GCONV=1E-8) satisfied. |
| Optimization Summary | |
|---|---|
| Optimization Technique | Trust Region |
| Iterations | 2 |
| Function Calls | 8 |
| Log Likelihood | -157.72104 |
| Fit Statistics | |
|---|---|
| -2 Log Likelihood | 315.44208 |
| Akaike's Information Criterion | 319.44208 |
| Corrected Akaike's Information Criterion | 319.56579 |
| Schwarz's Bayesian Information Criterion | 324.65242 |
| Kolmogorov-Smirnov Statistic | 0.50641 |
| Anderson-Darling Statistic | 0.31240 |
| Cramer-von Mises Statistic | 0.04353 |
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Mu | 1 | 1.49605 | 0.02651 | 56.43 | <.0001 |
| Sigma | 1 | 0.26243 | 0.01874 | 14.00 | <.0001 |
The parameter estimates of the Burr distribution are shown in Figure 5. These estimates are used in the next example.
Figure 5: Parameter Estimates for the Burr Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Theta | 1 | 4.62348 | 0.46181 | 10.01 | <.0001 |
| Alpha | 1 | 1.15706 | 0.47493 | 2.44 | 0.0167 |
| Gamma | 1 | 6.41227 | 0.99039 | 6.47 | <.0001 |