The CCDM Procedure
Simulation Processes for Collective Risk Simulation Mode
This section describes the simulation processes for the collective risk mode, which is the simulation mode that PROC CCDM uses when you do not specify the SIMULATIONMODE= option or when you specify SIMULATIONMODE=CR. For other simulation modes, see the section Simulation Processes for Pure Premium and Custom Simulation Modes.
PROC CCDM selects a simulation process based on whether you specify external counts or you request that PROC CCDM simulate the counts and on whether or not the severity or frequency models contain regression effects. The following sections describe the processes for the different possibilities.
Simulation with No Regressors and No External Counts
If you specify severity and frequency models that contain no regression effects and if you do not specify externally simulated counts in the EXTERNALCOUNTS statement, then PROC CCDM uses the simulation process that this section describes.
The process is described for one severity distribution, dist. If you specify multiple severity distributions in the SEVERITYMODEL statement, then the process is repeated for each specified distribution.
The following steps are repeated M times to generate a compound distribution sample of size M, where M is the value of the NREPLICATES= option:
Use the frequency model that you specify in the COUNTSTORE= item store or the COUNTMODEL statement to draw a value N from the count distribution, where N denotes the number of loss events that are expected to occur in the time period that is being simulated. Adjust N to conform to the upper limit by setting it equal to
, where
is the value of the MAXCOUNTDRAW= option.
Use the parameter estimates of the severity distribution dist (which are read from the SEVERITYEST= data table or the SEVERITYSTORE= item store that you specify) to draw N severity values,
(
). (For more information about how PROC CCDM draws an individual severity value, see the section Making a Random Draw from a Severity Distribution.)
-
Add the N severity values that are drawn in step 2 to compute one sample point, S, of the compound distribution sample as
Although it is more common to fit a frequency model that contains regressors, the CNTSELECT procedure enables you to fit a frequency model that does not contain regressors. If you do not specify any regressors in the MODEL statement of PROC CNTSELECT, then it fits a model that contains only an intercept.
Simulation with Regressors and No External Counts
If the severity or frequency models contain regression effects and if you do not specify externally simulated counts in the EXTERNALCOUNTS statement, then you must specify a DATA= data table to provide values of the regression variables, which together represent a scenario for which you want to simulate the CDM. In this case, PROC CCDM uses the simulation process that this section describes.
The process is described for one severity distribution. If you specify multiple severity distributions in the SEVERITYMODEL statement, then the process is repeated for each specified distribution.
Note that you are performing scenario analysis when regression effects are present. Let K denote the number of observations that form the scenario. This is the number of observations either in the current BY group or in the entire DATA= data table if you do not specify the BY statement. If , then you are modeling the scenario for a group of entities. If K = 1, then you are modeling the scenario for one entity.
The following steps are repeated M times to generate a compound distribution sample of size M, where M is the value of the NREPLICATES= option:
For each observation k (
), draw a count
from the frequency model, where
denotes the number of loss events that are expected to be generated by entity k in the time period that is being simulated. If you specify the COUNTSTORE= item store, then draw
from the frequency model whose parameters are determined by the frequency regressors in observation k. If you specify the COUNTMODEL statement, the frequency model does not depend on the regressors, so draw
by using the information in the COUNTMODEL statement. Adjust
to conform to the upper limit by setting it equal to
, where
is the value of the MAXCOUNTDRAW= option.
Add counts from all observations to compute
, which is the total number of loss events that are expected to occur in the time period that is being simulated.
Make N random draws from the severity distribution and add them to generate one sample point of the compound distribution sample. If you specify a scale regression model for the severity distribution, then use the values of severity regressors in the kth observation to compute the scale parameter of the severity distribution and make
severity draws from that distribution. (For more information about how PROC CCDM makes a random severity draw, see the section Making a Random Draw from a Severity Distribution.)
If you specify the BY statement, then a separate sample of size M is created for each BY group in the DATA= data table.
Illustration of Aggregate Loss Simulation Process
As an illustration of the simulation process, consider a very simple example of analyzing the distribution of an aggregate loss that is incurred by a set of policyholders of an automobile insurance company in a period of one year. It is postulated that the frequency and severity distributions depend on three variables: Age, Gender (1: female, 2: male), and CarType (1: sedan, 2: sport utility vehicle). So these variables are used as regressors while you fit the count model and severity scale regression model by using the CNTSELECT and SEVSELECT procedures, respectively. Now, consider that you want to use the fitted frequency and severity models to estimate the distribution of the aggregate loss that is incurred by a set of five policyholders. The characteristics of the five policyholders are encoded in a data table named mycas.Scenario that has the following contents:
Obs age gender carType
1 30 2 1
2 25 1 2
3 45 2 2
4 33 1 1
5 50 1 1
The column Obs contains the observation number. It is shown only for the purpose of illustration. It does not need to be present in the data table. The following PROC CCDM step simulates the scenario in the data table mycas.Scenario:
proc ccdm data=mycas.scenario
severityest=<severity parameter estimates data table>
countstore=<count model store> nreplicates=<sample size>;
severitymodel <severity distribution name(s)>;
run;
The following process generates a sample from the aggregate loss distribution for the scenario in the data table mycas.Scenario:
-
Use the values
Age=30,Gender=2, andCarType=1 in the first observation to draw a count from the count distribution. Let that count be 2. Repeat the process for the remaining four observations. Let the counts be as shown in theCountcolumn in the following table:Obs age gender carType count 1 30 2 1 2 2 25 1 2 1 3 45 2 2 2 4 33 1 1 3 5 50 1 1 0The
Countcolumn is shown for illustration only; it is not added as a variable to the DATA= data table. The simulated counts from all the observations are added together to get a value of N = 8. This means that for this particular sample point, you expect a total of eight loss events in a year from these five policyholders.
-
For the first observation, the scale parameter of the severity distribution is computed by using the values
Age=30,Gender=2, andCarType=1. That value of the scale parameter is used together with estimates of the other parameters from the SEVERITYEST= data table to make two draws from the severity distribution. Each draw simulates the magnitude of the loss that is expected from the first policyholder. The process is repeated for the remaining four policyholders. The fifth policyholder does not generate any loss event for this particular sample point, so no severity draws are made by using the fifth observation. Let the severity draws, rounded to integers for convenience, be as shown in the_SEV_column in the following table:Obs age gender carType count _sev_ 1 30 2 1 2 350 2100 2 25 1 2 1 4500 3 45 2 2 2 700 4300 4 33 1 1 3 600 1500 950 5 50 1 1 0The
_SEV_column is shown for illustration only; it is not added as a variable to the DATA= data table.PROC CCDM adds the severity values of the eight draws to compute an aggregate loss value of 15,000. After recording this amount in the sample, the process returns to step 1 to compute the next point in the aggregate loss sample. For example, in the second iteration, the count distribution of each policyholder might generate one loss event, for a total of five loss events, and the five severity draws from the severity distributions that govern each policyholder might add up to 5,000. Then, the value of 5,000 is recorded as the second point in the aggregate loss sample. The process continues until M aggregate loss sample points are simulated, where the M is the value that you specify in the NREPLICATES= option.
Simulation with External Counts
If you specify externally simulated counts by using the EXTERNALCOUNTS statement, then each replication in the input data table represents the loss events that are generated by an entity. An entity can be an individual or organization for which you want to estimate the compound distribution. If an entity has any characteristics that are used as external factors (regressors) in developing the severity scale regression model, then you must specify the values of those factors in the DATA= data table. If you specify the ID= variable, then multiple observations for the same replication ID represent different entities in a group for which you are simulating the CDM.
This section describes the simulation process that PROC CCDM uses in the presence of externally simulated counts.
The process is described for one severity distribution. If you specify multiple severity distributions in the SEVERITYMODEL statement, then the process is repeated for each specified distribution.
Let there be M distinct replications in the current BY group of the DATA= data table or in the entire DATA= data table if you do not specify the BY statement. A replication is identified by either the observation number or the value of the ID= variable that you specify in the EXTERNALCOUNTS statement.
For each of the M values of the replication identifier, the following steps are executed R times, where R is the value of the NREPLICATES= option:
If there are K (
) observations for the current value of the replication identifier, then for each observation k (
), set
equal to the value of the COUNT= variable in that observation. Adjust
to conform to the upper limit by setting it equal to
, where
is the value of the MAXCOUNTDRAW= option.
Compute the total number of losses N across all observations of the replication as
.
Make N random draws from the severity distribution and add them to generate one sample point of the compound distribution sample. If you specify a scale regression model for the severity distribution, then use the values of severity regressors in the kth observation to compute the scale parameter of the severity distribution and make
severity draws from that distribution. (For more information about how PROC CCDM makes a random severity draw, see the section Making a Random Draw from a Severity Distribution.)
This process generates a compound distribution sample of size . If you specify the BY statement, then a separate sample of size
is created for each BY group in the DATA= data table.
Illustration of the Simulation Process with External Counts
To illustrate the simulation process, consider the following simple example. In this example, your severity model does not contain any regressors. An example that uses a severity scale regression model is illustrated later. Assume that you have made 10 random draws from an external count model and recorded them in the variable ExtCount of a data table named mycas.Counts1, as follows:
Obs extCount
1 3
2 2
3 0
4 1
5 3
6 4
7 1
8 2
9 0
10 5
Because the data table does not contain an ID= variable, the observation number that is shown in the Obs column acts as the replicate identifier. The following PROC CCDM step simulates an aggregate loss sample by using the data table mycas.Counts1:
proc ccdm data=mycas.counts1 nreplicates=5
severityest=<severity parameter estimates data table>;
severitymodel <severity distribution name(s)>;
externalcounts count=extCount;
run;
The simulation process works as follows:
For the first replication, which is associated with the first observation, three severity values are drawn from the severity distribution by using the parameter estimates that you specify in the SEVERITYEST= data table. If the severity values are 150, 500, and 320, then their sum of 970 is recorded as the first point of the aggregate loss sample. Because the value of the NREPLICATES= option is 5, this process of drawing three severity values and adding them together to form a point of the aggregate loss sample is repeated four more times to generate a total of five sample points that correspond to the first observation.
For the second replication, two severity values are drawn from the severity distribution. If the severity values are 450 and 100, then their sum of 550 is recorded as a point of the aggregate loss sample. This process of drawing two severity values and adding them together to form a point of the aggregate loss sample is repeated four more times to generate a total of five sample points that correspond to the second observation.
The process continues until all the replications, which are observations in this case, are exhausted.
The process results in an aggregate loss sample of size 50, which is equal to the number of replications in the data table (10) multiplied by the value of the NREPLICATES= option (5).
Now, consider an example in which the severity models in the SEVERITYEST= data table are scale regression models. In this case, the severity distribution that is used for drawing the severity value is determined by the values of regressors in the observation that is being processed. Consider that you want to simulate the aggregate loss that one policyholder incurs and you have recorded, in the variable ExtCount, the results of 10 random draws from an external count model. The DATA= data table has the following contents:
Obs age gender carType extCount
1 30 2 1 5
2 30 2 1 2
3 30 2 1 0
4 30 2 1 1
5 30 2 1 3
6 30 2 1 4
7 30 2 1 1
8 30 2 1 2
9 30 2 1 0
10 30 2 1 5
The simulation process in this case is the same as the process in the previous case of no regressors, except that the severity distribution that PROC CCDM uses for drawing the severity values has a scale parameter that is determined by the values of the regressors Age, Gender, and CarType in the observation that is being processed. In this particular example, all observations have the same value for all regressors, indicating that you are modeling a scenario in which the characteristics of the policyholder do not change during the time for which you have simulated the number of events. You can also model a scenario in which the characteristics of the policyholder change by recording those changes in the values of the appropriate regressors.
Extending this example further, consider that you want to analyze the distribution of the aggregate loss that is incurred by a group of policyholders, as in the example in the section Illustration of Aggregate Loss Simulation Process. Let the data table mycas.Counts2 record multiple replications of the number of losses that each policyholder might generate. The contents of the data table mycas.Counts2 are as follows:
Obs replicateId age gender carType extCount
1 1 30 2 1 2
2 1 25 1 2 1
3 1 45 2 2 3
4 1 33 1 1 5
5 1 50 1 1 1
6 2 30 2 1 3
7 2 25 1 2 2
8 1 45 2 2 0
9 2 33 1 1 4
10 2 50 1 1 1
The variable ReplicateId records the identifier for the replication. Each replication contains multiple observations, such that each observation represents one of the policyholders that you are analyzing. For simplicity, only the first two replications are shown here.
The following PROC CCDM step simulates an aggregate loss sample by using the data table mycas.Counts2:
proc ccdm data=mycas.counts2 nreplicates=3
severityest=<severity parameter estimates data table>;
severitymodel <severity distribution name(s)>;
externalcounts count=extCount id=replicateId;
output out=aggloss samplevar=totalLoss;
run;
When you specify an ID= variable in the EXTERNALCOUNTS statement, the DATA= data table must be partitioned by the ID= variable as noted in the description of the ID= option.
The simulation process works as follows:
-
First, the five observations of the first replication (
ReplicateId=1) are analyzed. For the first observation (Obs=1), the scale parameter of the severity distribution is computed by using the valuesAge=30,Gender=2, andCarType=1. That value of the scale parameter is used together with estimates of the other parameters from the SEVERITYEST= data table to make two draws from the severity distribution. Next, the regressor values of the second observation are used to compute the scale parameter of the severity distribution, which is used to make one severity draw. The process continues such that the regressor values in the third, fourth, and fifth observations are used to select which severity distribution to make three, five, and one draws from, respectively. Let the severity values that are drawn from the observations of this replication be as shown in the_SEV_column in the following table, where the_SEV_column is shown for illustration only; it is not added as a variable to the DATA= data table.Obs replicateId age gender carType extCount _sev_ 1 1 30 2 1 2 700 500 2 1 25 1 2 1 5000 3 1 45 2 2 3 900 1400 300 4 1 33 1 1 5 350 2000 150 800 600 5 1 50 1 1 1 250The values of all 12 severity draws are added together to compute and record the value of 12,950 as the first point of the aggregate loss sample. Because you specify NREPLICATES=3 in the PROC CCDM step, this process of making 12 severity draws from the respective observations is repeated two more times to generate a total of three sample points for the first replication.
The five observations of the second replication (
ReplicateId=2) are analyzed next to draw three, two, four, and one severity values from the severity distributions, with scale parameters that are determined by the regressor values in the sixth, seventh, ninth, and tenth observations, respectively. The 10 severity values are added together to form a point of the aggregate loss sample. This process of making 10 severity draws from the respective observations is repeated two more times to generate a total of three sample points for the second replication.
If your data table mycas.Counts2 contains 10,000 distinct values of ReplicateId, then 30,000 observations are written to the data table mycas.AggLoss that you specify in the OUTPUT statement of the preceding PROC CCDM step. Because you specify SAMPLEVAR=TotalLoss in the OUTPUT statement, the aggregate loss sample is available in the TotalLoss column of the data table mycas.AggLoss.
Making a Random Draw from a Severity Distribution
PROC CCDM uses a simple process to make a random draw from the severity distribution. It first makes a draw from the uniform distribution. If this value is denoted by p, then the corresponding severity draw is computed as , where F denotes the cumulative distribution function (CDF) of the severity distribution.
For a severity distribution dist, PROC CCDM computes the inverse of the CDF (that is, ) by evaluating the dist_QUANTILE function. If PROC CCDM does not find the definition of the dist_QUANTILE function in the DEFINITIONSTABLE= table or in the libraries that you specify in the CMPLIB= system option, it uses a numerical inversion algorithm to invert the CDF function at p. This algorithm yields an approximate inverse and can be slow because it makes multiple calls to the dist_CDF or dist_LOGCDF function of the distribution. So it is recommended that you define the dist_QUANTILE function whenever possible for any custom severity distribution that you define for use with PROC CCDM. For more information about defining custom severity distributions, see the section Defining a Severity Distribution Model with the FCMP Procedure in Chapter 23, The SEVSELECT Procedure. The dist_QUANTILE function is already defined for all severity distributions in the library
Sashelp.Svrtdist.
To evaluate the dist_QUANTILE or dist_CDF function, PROC CCDM uses the distribution parameters that it reads from the SEVERITYEST= data table or the SEVERITYSTORE= item store. If the severity model contains regression effects, then PROC CCDM computes the scale parameter of the severity distribution by using the base value of the scale parameter and the values of the regression effects.
Making a Random Draw from a Truncated Severity Distribution
PROC CCDM assumes that the severity distribution that you specify in the SEVERITYEST= data table or the SEVERITYSTORE= item store is an untruncated distribution—that is, a distribution that has the full support. For severity distributions that have a scale parameter, full support is , and for severity distributions that have a log-scale parameter, full support is
. If you want PROC CCDM to make draws from a truncated distribution, which is especially useful in insurance applications, then you can specify the LEFTTRUNCATION= or RIGHTTRUNCATION= options in the PROC CCDM statement.
Let L and R denote the left-truncation threshold and right-truncation limit that you specify in the LEFTTRUNCATION= and RIGHTTRUNCATION= options, respectively. Let F and denote the cumulative distribution functions (CDF) of the untruncated and truncated severity distributions, respectively. Not specifying the LEFTTRUNCATION= option is equivalent to specifying
—that is,
. Not specifying the RIGHTTRUNCATION= option is equivalent to specifying
—that is,
.
PROC CCDM makes a random draw from the truncated severity distribution by using the following steps:
Draw a value
from the uniform distribution. This is the CDF of the truncated distribution.
-
Compute p, which represents the CDF of the untruncated distribution, as
Compute the severity value x by inverting the CDF of the untruncated distribution as
. The value of x that is computed in this manner is essentially a random draw from
—that is,
. The inversion of the untruncated CDF is described in the section Making a Random Draw from a Severity Distribution.