The SEVSELECT Procedure

An Example of Modeling Regression Effects

Consider a scenario in which the magnitude of the response variable might be affected by some regressor (exogenous or independent) variables. The SEVSELECT procedure enables you to model the effect of such variables on the distribution of the response variable via an exponential link function. In particular, if you have k random regressor variables denoted by x Subscript j (j equals 1 comma ellipsis comma k), then the distribution of the response variable Y is assumed to have the form

upper Y tilde exp left-parenthesis sigma-summation Underscript j equals 1 Overscript k Endscripts beta Subscript j Baseline x Subscript j Baseline right-parenthesis dot script upper F left-parenthesis normal upper Theta right-parenthesis

where script upper F denotes the distribution of Y with parameters normal upper Theta and beta Subscript j Baseline left-parenthesis j equals 1 comma ellipsis comma k right-parenthesis denote the regression parameters (coefficients).

For the effective distribution of Y to be a valid distribution from the same parametric family as script upper F, it is necessary for script upper F to have a scale parameter. The effective distribution of Y can be written as

upper Y tilde script upper F left-parenthesis theta comma normal upper Omega right-parenthesis

where theta denotes the scale parameter and normal upper Omega denotes the set of nonscale parameters. The scale theta is affected by the regressors as

theta equals theta 0 dot exp left-parenthesis sigma-summation Underscript j equals 1 Overscript k Endscripts beta Subscript j Baseline x Subscript j Baseline right-parenthesis

where theta 0 denotes a base value of the scale parameter.

Given this form of the model, PROC SEVSELECT allows a distribution to be a candidate for modeling regression effects only if it has an untransformed or a log-transformed scale parameter.

All the predefined distributions, except the lognormal distribution, have a direct scale parameter (that is, a parameter that is a scale parameter without any transformation). For the lognormal distribution, the parameter mu is a log-transformed scale parameter. This can be verified by replacing mu with a parameter theta equals e Superscript mu, which results in the following expressions for the PDF f and the CDF F in terms of theta and sigma, respectively, where normal upper Phi denotes the CDF of the standard normal distribution:

f left-parenthesis x semicolon theta comma sigma right-parenthesis equals StartFraction 1 Over x sigma StartRoot 2 pi EndRoot EndFraction e Superscript minus one-half left-parenthesis StartFraction log left-parenthesis x right-parenthesis minus log left-parenthesis theta right-parenthesis Over sigma EndFraction right-parenthesis squared Baseline and upper F left-parenthesis x semicolon theta comma sigma right-parenthesis equals normal upper Phi left-parenthesis StartFraction log left-parenthesis x right-parenthesis minus log left-parenthesis theta right-parenthesis Over sigma EndFraction right-parenthesis

With this parameterization, the PDF satisfies the f left-parenthesis x semicolon theta comma sigma right-parenthesis equals StartFraction 1 Over theta EndFraction f left-parenthesis StartFraction x Over theta EndFraction semicolon 1 comma sigma right-parenthesis condition and the CDF satisfies the upper F left-parenthesis x semicolon theta comma sigma right-parenthesis equals upper F left-parenthesis StartFraction x Over theta EndFraction semicolon 1 comma sigma right-parenthesis condition. This makes theta a scale parameter. Hence, mu equals log left-parenthesis theta right-parenthesis is a log-transformed scale parameter and the lognormal distribution is eligible for modeling regression effects.

The following DATA step simulates a lognormal sample whose scale is decided by the values of the three regressors X1, X2, and X3 as follows:

mu equals log left-parenthesis theta right-parenthesis equals 1 plus 0.75 upper X 1 minus upper X 2 plus 0.25 upper X 3
/*----------- Lognormal Model with Regressors ------------*/
data test_sev3(keep=y x1-x3
               label='A Lognormal Sample Affected by Regressors');
   array x{*} x1-x3;
   array b{4} _TEMPORARY_ (1 0.75 -1 0.25);
   call streaminit(45678);
   label y='Response Influenced by Regressors';
   Sigma = 0.25;
   do n = 1 to 100;
      Mu = b(1); /* log of base value of scale */
      do i = 1 to dim(x);
         x(i) = rand('UNIFORM');
         Mu = Mu + b(i+1) * x(i);
      end;
      y = exp(Mu) * rand('LOGNORMAL')**Sigma;
      output;
   end;
run;

The following DATA step loads the data set Work.Test_sev3 into a data table in your CAS session that is associated with the mycas CAS engine libref:

data mycas.test_sev3;
   set test_sev3;
run;

The following PROC SEVSELECT step fits the lognormal, Burr, and gamma distribution models to these data. The regressors are specified in the SCALEMODEL statement.

proc sevselect data=mycas.test_sev3 crit=aicc print=all;
   loss y;
   scalemodel x1-x3;

   dist logn burr gamma;
run;

Some of the key results that PROC SEVSELECT prepares are shown in Figure 14 through Figure 18. The descriptive statistics of all the variables are shown in Figure 14.

Figure 14: Summary Results for the Regression Example

The SEVSELECT Procedure

Descriptive Statistics for y
Observations100
Observations Used for Estimation100
Minimum1.17863
Maximum6.65269
Mean2.99859
Standard Deviation1.12845

Descriptive Statistics for Regressors
VariableNMinimumMaximumMeanStandard
Deviation
x11000.00051150.979710.516890.28206
x21000.018830.999370.473450.28885
x31000.002550.975580.483010.29709


The comparison of the fit statistics of all the models is shown in Figure 15. It indicates that the lognormal model is the best model according to each of the likelihood-based statistics, whereas the gamma model is the best model according to two of the three EDF-based statistics.

Figure 15: Comparison of Statistics of Fit for the Regression Example

All Fit Statistics
Distribution-2 Log
Likelihood
AICAICCSBCKSADCvM
Logn187.49609*197.49609*198.13439*210.52194*1.9754417.246181.21665
Burr190.69154202.69154203.59476218.322562.0933413.93436*1.28529
Gamma188.91483198.91483199.55313211.940691.94472*15.847871.17617*
Asterisk (*) denotes the best model in the column.


The model information and the convergence results of the lognormal model are shown in Figure 16. The iteration history gives you a summary of how the optimizer is traversing the surface of the log-likelihood function in its attempt to reach the optimum. Both the change in the log likelihood and the maximum gradient of the objective function with respect to any of the parameters typically approach 0 if the optimizer converges.

Figure 16: Convergence Results for the Lognormal Model with Regressors

The SEVSELECT Procedure
 
Logn Distribution

Model Information
DistributionLogn
DescriptionLognormal Distribution
Distribution Parameters2
Regression Parameters3

Convergence Status
Convergence criterion (GCONV=1E-8) satisfied.

Optimization Iteration History
IterFunction
Calls
Log
Likelihood
ChangeMaximum
Gradient
0293.75285 6.16002
1493.74805-0.004810.11031
2693.74805-1.5017E-60.0000338
31093.74805-1.279E-133.1246E-12

Optimization Summary
Optimization TechniqueTrust Region
Iterations3
Function Calls10
Log Likelihood-93.74804701


The final parameter estimates of the lognormal model are shown in Figure 17. All the estimates are significantly different from 0. The estimate that is reported for the parameter Mu is the base value for the log-transformed scale parameter mu. Let x Subscript i Baseline left-parenthesis 1 less-than-or-equal-to i less-than-or-equal-to 3 right-parenthesis denote the observed value for regressor Xi. If the lognormal distribution is chosen to model Y, then the effective value of the parameter mu varies with the observed values of regressors as

mu equals 1.04047 plus 0.65221 x 1 minus 0.91116 x 2 plus 0.16243 x 3

These estimated coefficients are reasonably close to the population parameters (that is, within one or two standard errors).

Figure 17: Parameter Estimates for the Lognormal Model with Regressors

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Mu11.040470.0761413.66<.0001
Sigma10.221770.0160913.78<.0001
x110.652210.081677.99<.0001
x21-0.911160.07946-11.47<.0001
x310.162430.077822.090.0395


The estimates of the gamma distribution model, which is the best model according to a majority of the EDF-based statistics, are shown in Figure 18. The estimate that is reported for the parameter Theta is the base value for the scale parameter theta. If the gamma distribution is chosen to model Y, then the effective value of the scale parameter is theta equals 0.14293 exp left-parenthesis 0.64562 x 1 minus 0.89831 x 2 plus 0.14901 x 3 right-parenthesis.

Figure 18: Parameter Estimates for the Gamma Model with Regressors

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Theta10.142930.023296.14<.0001
Alpha120.377262.932776.95<.0001
x110.645620.082247.85<.0001
x21-0.898310.07962-11.28<.0001
x310.149010.078701.890.0613


Last updated: January 27, 2023