The SEVSELECT Procedure

Example 18.6 Scale Regression with Rich Regression Effects

This example illustrates the use of regression effects that include CLASS variables and interaction effects.

Consider that you, as an actuary at an automobile insurance company, want to evaluate the effect of certain external factors on the distribution of the severity of the losses that your policyholders incur. Such analysis can help you determine the relative differences in premiums that you should charge to policyholders who have different characteristics. Assume that when you collect and record the information about each claim, you also collect and record some key characteristics of the policyholder and the vehicle that is involved in the claim. This example focuses on the following five factors: type of car, safety rating of the car, gender of the policyholder, education level of the policyholder, and annual household income of the policyholder (which can be thought of as a proxy for the luxury level of the car). Let these regressors be recorded in the variables CarType (1: sedan, 2: sport utility vehicle), CarSafety (scaled to be between 0 and 1, the safest being 1), Gender (1: female, 2: male), Education (1: high school graduate, 2: college graduate, 3: advanced degree holder), and Income (scaled by a factor of 1/100,000), respectively. Let the historical data about the severity of each loss be recorded in the LossAmount variable of the mycas.Losses data table in the CAS libref. Let the data table also contain two additional variables, Deductible and Limit, that record the deductible and ground-up loss limit provisions, respectively, of the insurance policy that the policyholder has. The limit on ground-up loss is usually derived from the payment limit that a typical insurance policy states. Deductible serves as the left-truncation variable, and Limit serves as the right-censoring variable.

The following SAS statements simulate an example of the mycas.Losses data table:

proc format casfmtlib='myfmtlib';
   value genderFmt 1='Female'
                   2='Male';
run;

data losses(keep=gender carType education carSafety income
                 lossAmount deductible limit);
   call streaminit(12345);
   array sx{8} _temporary_;
   array sbeta{9} _TEMPORARY_ (5 0.6 0.4 -0.75 -0.3 0.4 0.7 -0.5 -0.3);

   length carType $8 education $16;
   format gender genderFmt.;

   sigma = 0.5;
   do lossEventId=1 to 6000;
      /* Simulate policyholder and vehicle attributes */
      do i=1 to dim(sx);
         sx(i) = 0;
      end;

      if (rand('UNIFORM') < 0.5) then do;
         gender = 1; * female;
         sx(2) = 1;
      end;
      else do;
         gender = 2; * male;
      end;

      if (rand('UNIFORM') < 0.7) then do;
         carInt = 1;
         carType = 'Sedan';
      end;
      else do;
         carInt = 2;
         carType = 'SUV';
         sx(1) = 1;
      end;

      educationLevel = rand('UNIFORM');
      if (educationLevel < 0.5) then do;
         eduInt = 1;
         education = 'High School';
      end;
      else if (educationLevel < 0.85) then do;
         eduInt = 2;
         education = 'College';
         if (carInt=1) then
            sx(8) = 1;
         else
            sx(6) = 1;
      end;
      else do;
         eduInt = 3;
         education = 'AdvancedDegree';
         if (carInt=1) then
            sx(7) = 1;
         else
            sx(5) = 1;
      end;

      carSafety = rand('UNIFORM'); /* scaled to be between 0 & 1 */
      sx(3) = carSafety;

      income = MAX(15000,int(rand('NORMAL', eduInt*30000, 50000)))/100000;
      sx(4) = income;

      /* Simulate lognormal severity */
      Mu = sbeta(1);
      do i=1 to dim(sx);
         Mu = Mu + sx(i) * sbeta(i+1);
      end;
      lossAmount = exp(Mu) * rand('LOGNORMAL')**Sigma;
      loglossAmount = log(lossAmount);

      deductible = lossAmount * rand('UNIFORM');
      if (rand('UNIFORM') < 0.25) then
         limit = lossAmount;

      output;
   end;
run;

data mycas.losses;
   set losses;
run;

The variables CarType, Education, and Gender each contain a known, finite set of discrete values. By specifying such variables as classification variables, you can separately identify the effect of each level of the variable on the severity distribution. For example, you might be interested in finding out how the magnitude of loss for a sport utility vehicle (SUV) differs from that for a sedan. This is an example of a main effect. You might also want to evaluate how the distribution of losses that are incurred by a policyholder with a college degree who drives a SUV differs from that of a policyholder with an advanced degree who drives a sedan. This is an example of an interaction effect. You can include various such types of effects in the scale regression model. For more information about the effect types, see the section Specification and Parameterization of Model Effects in Chapter 3, Shared Concepts.

Analyzing such a rich set of regression effects can help you make more accurate predictions about the losses that a new applicant with certain characteristics might incur when he or she requests insurance for a specific vehicle, which can further help you with ratemaking decisions.

The following PROC SEVSELECT step fits the scale regression model with a lognormal distribution to data in the mycas.Losses data table, and stores the model and parameter estimate information in the mycas.Est data table on the CAS server:

/* Fit scale regression model with different types of regression effects */
proc sevselect data=mycas.losses outest=mycas.est print=all;
   loss lossAmount / lt=deductible rc=limit;
   class carType gender education;
   scalemodel carType gender carSafety income education*carType
              income*gender carSafety*income;
   dist logn;
run;

The SCALEMODEL statement in the preceding PROC SEVSELECT step includes two main effects (carType and gender), two singleton continuous effects (carSafety and income), one interaction effect (education*carType), one continuous-by-class effect (income*gender), and one polynomial continuous effect (carSafety*income).

When you specify a CLASS statement, it is recommended that you observe the "Class Level Information" table. For this example, the table is shown in Output 18.6.1. Note that if you specify BY-group processing, then the class level information might change from one BY group to the next, potentially resulting in a different parameterization for each BY group.

Output 18.6.1: Class Level Information Table

The SEVSELECT Procedure

Class Level Information
ClassLevelsValues
carType2SUV Sedan
gender2Female Male
education3AdvancedDegree College High School


The regression modeling results for the lognormal distribution are shown in Output 18.6.2. The "Initial Parameter Values and Bounds" table is important especially because the preceding PROC SEVSELECT step uses the default GLM parameterization, which is a singular parameterization—that is, it results in some redundant parameters. As shown in the table, the redundant parameters correspond to the last level of each classification variable; this correspondence is a defining characteristic of a GLM parameterization. An alternative would be to use the reference parameterization by specifying the PARAM=REFERENCE option in the CLASS statement, which does not generate redundant parameters for effects that contain CLASS variables and enables you to specify a reference level for each CLASS variable.

Output 18.6.2: Initial Values for the Scale Regression Model with Class and Interaction Effects

Initial Parameter Values and Bounds
ParameterInitial
Value
Lower
Bound
Upper
Bound
Mu4.85228-709.78271709.78271
Sigma0.523481.05367E-8Infty
carType SUV0.54686-709.78271709.78271
carType SedanRedundant-709.78271709.78271
gender Female0.34893-709.78271709.78271
gender MaleRedundant-709.78271709.78271
carSafety-0.63504-709.78271709.78271
income-0.24031-709.78271709.78271
carType SUV * education AdvancedDegree0.32719-709.78271709.78271
carType SUV * education College0.68899-709.78271709.78271
carType SUV * education High SchoolRedundant-709.78271709.78271
carType Sedan * education AdvancedDegree-0.44650-709.78271709.78271
carType Sedan * education College-0.26834-709.78271709.78271
carType Sedan * education High SchoolRedundant-709.78271709.78271
income * gender Female0.00843-709.78271709.78271
income * gender MaleRedundant-709.78271709.78271
carSafety * income-0.04744-709.78271709.78271


The convergence and optimization summary information in Output 18.6.3 indicates that the scale regression model for the lognormal distribution has converged with the default optimization technique in five iterations.

Output 18.6.3: Optimization Summary for the Scale Regression Model with Class and Interaction Effects

Convergence Status
Convergence criterion (GCONV=1E-8) satisfied.

Optimization Summary
Optimization TechniqueTrust Region
Iterations5
Function Calls14
Log Likelihood-19716.25785


The "Parameter Estimates" table in Output 18.6.4 shows the distribution parameter estimates and estimates for various regression effects. You can use the estimates for effects that contain CLASS variables to infer the relative influence of various CLASS variable levels. For example, on average, the magnitude of losses that are incurred by the female drivers is exp left-parenthesis 0.41714 right-parenthesis almost-equals 1.52 times greater than that of male drivers, and an SUV driver with an advanced degree incurs a loss that is on average exp left-parenthesis 0.42454 right-parenthesis slash exp left-parenthesis negative 0.35208 right-parenthesis almost-equals 2.174 times greater than the loss that a college-educated sedan driver incurs. Neither the continuous-by-class effect income*gender nor the polynomial continuous effect carSafety*income is significant in this example.

Output 18.6.4: Parameter Estimates for the Scale Regression with Class and Interaction Effects

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Mu15.123320.03792135.11<.0001
Sigma10.569410.0074476.56<.0001
carType SUV10.646860.0298921.64<.0001
carType Sedan00...
gender Female10.417140.0321312.98<.0001
gender Male00...
carSafety1-0.880780.05532-15.92<.0001
income1-0.380450.05128-7.42<.0001
carType SUV * education AdvancedDegree10.424540.051078.31<.0001
carType SUV * education College10.732280.0373819.59<.0001
carType SUV * education High School00...
carType Sedan * education AdvancedDegree1-0.566850.03525-16.08<.0001
carType Sedan * education College1-0.352080.02618-13.45<.0001
carType Sedan * education High School00...
income * gender Female10.012880.043940.290.7694
income * gender Male00...
carSafety * income10.064990.075470.860.3892


If you want to update the model when new claims data arrive, then you can potentially speed up the estimation process by specifying the OUTEST= data table that is created by the preceding PROC SEVSELECT step as an INEST= data table in a new PROC SEVSELECT step. To illustrate, the following PROC SEVSELECT step refits the model on the same input data as the preceding PROC SEVSELECT, but it uses the mycas.Est data table that is created by that step as an INEST= data table:

/* Refit scale regression model on new data */
proc sevselect data=mycas.losses inest=mycas.est print=all;
   loss lossAmount / lt=deductible rc=limit;
   class carType gender education;
   scalemodel carType gender carSafety income education*carType
              income*gender carSafety*income;
   dist logn;
run;

Because the INEST= data table is used to initialize the distribution and regression parameters, the optimization occurs in very few iterations, as shown in Output 18.6.5.

Output 18.6.5: Optimization Summary for Refitting the Scale Regression Model with the INEST= Option

The SEVSELECT Procedure
 
Logn Distribution

Convergence Status
Convergence criterion (ABSGCONV=0.00001) satisfied.

Optimization Summary
Optimization TechniqueTrust Region
Iterations0
Function Calls4
Log Likelihood-19716.25785


Last updated: April 15, 2021