The CCDM Procedure
Simulation of Adjusted Compound Distribution Sample
If you specify a list of D adjusted severity symbols in the ADJUSTEDSEVERITY= option and specify programming statements that compute those symbols, then a separate compound distribution sample is generated for each adjusted severity symbol. The set of all programming statements that you specify in the PROC CCDM step is called the severity adjustment program.
Formally, the severity adjustment program is expected to implement D adjustment functions for all D adjusted severity symbols. The adjustment function for the dth symbol is denoted by (
). The function
uses the unadjusted severity value,
along with other values, some of which depend on the simulation mode, to compute and return an adjusted severity value,
. The information that is available to compute
depends on the simulation mode as follows:
-
For the collective risk simulation mode, if N denotes the number of loss events that are simulated for the current replication of the simulation process, then for the severity draw,
, of the jth loss event (
), the adjusted severity value for the dth symbol is
where
is the aggregate unadjusted loss,
is the aggregate value of the dth adjusted severity symbol before the jth severity draw that generates
, and
denotes the vector of values of the stochastic symbols that you specify in the SIMULATEDSYMBOL statement. The initial values of both S and
are set to 0—that is,
and
(
). Prior to executing your adjustment program, PROC CCDM makes a random draw for each stochastic symbol by using that symbol’s probability distribution and parameter values. The aggregate adjusted loss for the replication is
, which is denoted by
for simplicity and is defined as
-
For pure premium and custom simulation modes, there is only one severity draw for every frequency draw, so the index j is always equal to 1.
For pure premium mode, the adjusted severity value depends only on the unadjusted severity value and stochastic symbols as
and the aggregate adjusted loss for the replication is
.
For custom simulation mode, you are free to combine the values of frequency, severity, and stochastic symbols in any manner to compute the adjusted severity value as
and PROC CCDM sets the aggregate adjusted loss equal to the adjusted severity value as
.
In your severity adjustment program, you can use the following keywords as placeholder symbols for the input arguments of the function :
Your severity adjustment program must assign the output of the dth function to the dth symbol that you specify in the ADJUSTEDSEVERITY= option in the PROC CCDM statement. PROC CCDM uses the final assigned value of the dth symbol as the value of
.
You can use most DATA step statements and functions in your program. The DATA step file and the data table I/O statements (for example, INPUT, FILE, SET, and MERGE) are not available. However, some functionality of the PUT statement is supported. For more information, see the section "PROC FCMP and DATA Step Differences" in Base SAS Procedures Guide.
The simulation process that generates the aggregate adjusted loss sample is identical to the process that is described in the section Simulation with Regressors and No External Counts or the section Simulation with External Counts, except that after making each of the N severity draws, PROC CCDM executes your severity adjustment program to compute the adjusted severity symbols (). All the N adjusted severity values are added together to compute
, which forms one point of the dth aggregate adjusted loss sample. The process is illustrated using an example in the section Illustration of the Aggregate Adjusted Loss Simulation Process.
Using Severity Adjustment Variables
If you do not specify the DATA= data table, then your ability to adjust the severity value is limited, because you can use only the current severity draw, sums of unadjusted and adjusted severity draws that are made before the current draw, and some constant numbers to encode your adjustment policy. That is sufficient if you want to estimate the distribution of aggregate adjusted loss for only one entity. However, if you are simulating a scenario that contains more than one entity, then it might be more useful if the adjustment policy depends on factors that are specific to each entity that you are simulating. To do that, you must specify the DATA= data table and encode such factors as adjustment variables in the DATA= data table. If A denotes the set of values of the adjustment variables, then each adjustment function accepts the set A as one of its inputs, in addition to the other input values that it accepts according to the simulation mode. PROC CCDM reads the values of adjustment variables from the DATA= data table and supplies the set of those values (A) to your severity adjustment program. For an invocation of each
with an unadjusted severity value of
, the values in set A are read from the same observation that is used to simulate
.
All adjustment variables that you use in your program must be present in the DATA= data table. You must not use any keyword for a placeholder symbol as a name of any variable in the DATA= data table, whether the variable is a severity adjustment variable or a regressor in the frequency or severity model. Further, the following restrictions apply to the adjustment variables in your severity adjustment program:
You can use only numeric-valued variables. This restriction also implies that you cannot use SAS functions or call routines that require character-valued arguments, unless you pass those arguments as constant (literal) strings or characters.
You cannot use functions that create lagged versions of a variable. If you need lagged versions, then you can use a DATA step before the PROC CCDM step to add those versions to the input data table.
The use of adjustment variables is illustrated using an example in the section Illustration of the Aggregate Adjusted Loss Simulation Process.
Aggregate Adjusted Loss Simulation for a Multi-entity Scenario
If you are simulating a scenario that consists of multiple entities, then you can use some additional pieces of information in your severity adjustment program. Let the scenario consist of K entities, and let denote the number of loss events that are incurred by the kth entity (
) in the current iteration of the simulation process. The simulation mode determines the processing of multiple entities as follows:
-
For the collective risk simulation mode, each value of
is adjusted to conform to the upper limit that is imposed by the value of the MAXCOUNTDRAW= option. The total number of severity draws that need to be made is
. The dth aggregate adjusted loss is now defined as
where
is an adjusted severity value of the dth symbol for the kth entity in the jth draw (
). The form of the adjustment function
that computes
is
where
is the value of the jth draw of unadjusted severity for the kth entity;
and
are the aggregate unadjusted loss and the aggregate adjusted loss, respectively, for the kth entity before
is generated;
denotes the vector of values of the jth draw of stochastic symbols that you specify in the SIMULATEDSYMBOL statement for the kth entity; and
is the set of values of the severity adjustment variables that PROC CCDM reads from the DATA= data table for the kth entity.
The index n (
) keeps track of the total number of severity draws, across all entities, that are made before
is generated. So
and
are the aggregate unadjusted loss and aggregate adjusted loss, respectively, for all the entities that are processed before
is generated. Note that
and
include the
draws that are made for the kth entity before
is generated.
The initial values of all types of aggregate losses are set to 0. In other words,
; for all values of d,
; for all values of k,
; and for all combinations of k and d,
.
In your severity adjustment program, you can use the following two additional placeholder keywords:
The previously described placeholder symbols _SEVSUM_ and _ADJSEVSUM_ represent
and
, respectively. If you have only one entity in the scenario (K = 1), then the values of _SEVSUMFOROBS_ and _ADJSEVSUMFOROBSd_ are identical to the values of _SEVSUM_ and _ADJSEVSUMd_, respectively.
-
For pure premium and custom simulation modes, there is one frequency draw per entity and only one severity draw for every frequency draw, so the index j is always equal to 1. PROC CCDM draws the frequency value
for each entity k and uses it according to the simulation mode as follows:
-
For the pure premium simulation mode, the form of the adjustment function
that computes
is
and PROC CCDM computes the aggregate adjusted loss as
-
For the custom simulation mode, you are free to combine the values of frequency, severity, stochastic symbols, and the adjustment variables in any manner to compute the adjusted severity value
. The form of the adjustment function
that computes
is
and PROC CCDM computes the aggregate adjusted loss as
For both the pure premium and custom simulation modes, your adjustment program can use the previously described placeholder symbols _SEVSUM_ and _ADJSEVSUM_ to represent
and
, respectively.
-
Your severity adjustment program must implement all D adjustment functions and assign the output of the dth function to the dth adjusted severity symbol that you specify in the ADJUSTEDSEVERITY= option. PROC CCDM executes the severity adjustment program for each jth severity value that it draws for each kth entity, and it uses the final value of the dth symbol as the value of
.
There is one caveat when a scenario consists of more than one entity () and when you use any of the symbols for cumulative severity values (_SEVSUM_, _ADJSEVSUMd_, _SEVSUMFOROBS_, or _ADJSEVSUMFOROBSd_) in your severity adjustment program. In this case, to make the simulation realistic, it is important to randomize the order of N severity draws across K entities. For more information, see the section Randomizing the Order of Severity Draws across Observations of a Scenario.
Illustration of the Aggregate Adjusted Loss Simulation Process
This section continues the example in the section Simulation with Regressors and No External Counts to illustrate the simulation of aggregate adjusted loss for the collective risk simulation mode.
Recall that the earlier example simulates a scenario that consists of five policyholders. Assume that you want to compute the distribution of the aggregate amount paid to all the policyholders in a year, where the payment for each loss is determined by a deductible and a per-payment limit. To begin with, you must record the deductible and limit information in the input DATA= data table. The following table shows the DATA= data table from the earlier example, extended to include two variables, Deductible and Limit:
Obs age gender carType deductible limit
1 30 2 1 250 5000
2 25 1 2 500 3000
3 45 2 2 100 2000
4 33 1 1 500 5000
5 50 1 1 200 2000
The variables Deductible and Limit are referred to as severity adjustment variables, because you need to use them to compute the adjusted severity. Let AmountPaid represent the value of adjusted severity that you are interested in. Further, let the following SAS programming statements encode your logic of computing the value of AmountPaid:
amountPaid = MAX(_sev_ - deductible, 0);
amountPaid = MIN(amountPaid, MAX(limit - _adjsevsumforobs_, 0));
PROC CCDM supplies your program with values of the placeholder symbols _SEV_ and _ADJSEVSUMFOROBS_, which represent the value of the current unadjusted severity draw and the sum of adjusted severity values from the previous draws, respectively, for the observation that is being processed. The use of _ADJSEVSUMFOROBS_ helps you ensure that the payment that is made to a particular policyholder in a year does not exceed the limit that is recorded in the variable Limit.
To simulate a sample for the aggregate of AmountPaid, you need to submit a PROC CCDM step whose structure is like the following:
proc ccdm data=<data table name> adjustedseverity=amountPaid
severityest=<severity parameter estimates data table>
countstore=<count model store>;
severitymodel <severity distribution name(s)>;
amountPaid = MAX(_sev_ - deductible, 0);
amountPaid = MIN(amountPaid, MAX(limit - _adjsevsumforobs_, 0));
run;
The simulation process of one replication that generates one point of the aggregate loss sample and the corresponding point of the aggregate adjusted loss sample is as follows:
-
Use the values
Age=30,Gender=2, andCarType=1 in the first observation to draw a count from the count distribution. Let that count be 3. Repeat the process for the remaining four observations. Let the counts be as shown in theCountcolumn in the following table:Obs age gender carType deductible limit count 1 30 2 1 250 5000 2 2 25 1 2 500 3000 1 3 45 2 2 100 2000 2 4 33 1 1 500 5000 3 5 50 1 1 200 2000 0The
Countcolumn is shown for illustration only; it is not added as a variable to the DATA= data table. The simulated counts from all the observations are added together to get a value of N = 8. This means that for this particular replication, you expect a total of eight loss events in a year from these five policyholders.
-
For the first observation, the scale parameter of the severity distribution is computed by using the values
Age=30,Gender=2, andCarType=1. That value of the scale parameter is used together with estimates of the other parameters from the SEVERITYEST= data table to make two draws from the severity distribution. The process is repeated for the remaining four policyholders. The fifth policyholder does not generate any loss event for this particular replication, so no severity draws are made by using the fifth observation. Let the severity draws, rounded to integers for convenience, be as shown in the_SEV_column in the following table:Obs age gender carType deductible limit count _sev_ 1 30 2 1 250 5000 2 350 2100 2 25 1 2 500 3000 1 4500 3 45 2 2 100 2000 2 700 4300 4 33 1 1 200 5000 3 600 1500 950 5 50 1 1 200 2000 0The
_SEV_column is shown for illustration only; it is not added as a variable to the DATA= data table. The sample point for the aggregate unadjusted loss is computed by adding together the severity values of eight draws, which gives an aggregate loss value of 15,000. The aggregate unadjusted loss is also referred to as the ground-up loss.For each severity draw, your severity adjustment program is executed to compute the adjusted severity, which is the value of
AmountPaidin this case. For the draws in the preceding table, the values ofAmountPaidare as follows:Obs deductible limit _sev_ _adjsevsumforobs_ amountPaid 1 250 5000 350 0 100 1 250 5000 2100 100 1850 2 500 3000 4500 0 3000 3 100 2000 700 0 600 3 100 2000 4300 600 1400 4 200 5000 600 0 400 4 200 5000 1500 400 1300 4 200 5000 950 1700 750The adjusted severity values are added together to compute the cumulative payment value of 9,400, which forms the first sample point for the aggregate adjusted loss.
After recording the aggregate unadjusted and aggregate adjusted loss values in their respective samples, the process returns to step 1 to compute the next sample point unless the specified number of sample points have been simulated.
In this particular example, you can verify that the order in which the eight loss events are simulated does not affect the aggregate adjusted loss. As a simple example, consider the following order of draws, which is different from the consecutive order that the preceding table uses:
Obs deductible limit _sev_ _adjsevsumforobs_ amountPaid 4 200 5000 600 0 400 3 100 2000 4300 0 2000 1 250 5000 350 0 100 3 100 2000 700 2000 0 4 200 5000 950 400 750 1 250 5000 2100 100 1850 2 500 3000 4500 0 3000 4 200 5000 1500 1150 1300Although the payments that are made for individual loss events differ, the aggregate adjusted loss is still 9,400.
However, in general, when you use a cumulative severity value such as _ADJSEVSUMFOROBS_ in your program, the order in which the draws are processed affects the final value of aggregate adjusted loss. For more information, see the sections Randomizing the Order of Severity Draws across Observations of a Scenario and Illustration of the Need to Randomize the Order of Severity Draws.
Randomizing the Order of Severity Draws across Observations of a Scenario
If you specify a scenario that consists of more than one entity, then it is assumed that each entity generates its loss events independently from the other entities. In other words, the time at which the loss event of one entity is generated or recorded is independent of the time at which the loss event of another entity is generated or recorded. To honor this assumption of independence among entities, the order in which those entities are processed must be randomized across K entities such that no entity is preferred over another. The K entities are represented by K observations of the scenario in the DATA= data table. If you specify external counts, the K observations correspond to the observations that have the same replication identifier value. If you do not specify the external counts, then the K observations correspond to all the observations in the BY group, or to all the observations in the entire DATA= data table if you do not specify the BY statement.
The randomization process depends on the simulation mode:
-
For the collective risk simulation mode, if entity k generates
loss events, where
is adjusted to conform to the upper limit that is imposed by the value of the MAXCOUNTDRAW= option, then the total number of loss events for a group of K entities is
. To simulate the aggregate loss for this group, N severity draws are made and aggregated to compute one point of the compound distribution sample. To honor the assumption of independence among entities, the order of those N severity draws must be randomized across K entities such that no entity is preferred over another.
PROC CCDM implements the randomization process over N severity draws as follows: First, it chooses one of the K observations at random, say the kth observation, and draws one severity value (
) from the severity distribution that is implied by that observation. It then supplies the severity value, along with the values of the severity adjustment variables (
) and random draws of simulated symbols (
), to your severity adjustment program to compute the adjusted severity values (
) and updates the respective aggregate adjusted losses (
). Next, it chooses another observation at random, draws one severity value from its implied severity distribution, and repeats the analysis of computing
and updating
. In each step, PROC CCDM increments by 1 the total number of events that are simulated for the selected observation k. When it simulates all
events for an observation k, it retires that observation and continues the process with the remaining observations until it makes a total of N severity draws.
For pure premium and custom simulation modes, there is one frequency and severity draw per entity. To honor the assumption of independence among entities, PROC CCDM randomizes the order in which it processes the entities as follows: First, it chooses one of the K observations at random, say the kth observation. It draws one frequency value and one severity value from the count and severity distributions, respectively, that are implied by the kth observation. It then supplies those frequency (
) and severity (
) values, along with the values of the severity adjustment variables (
) and random draws of simulated symbols (
), to your severity adjustment program to compute the adjusted severity values (
) and updates the respective aggregate adjusted losses (
). Next, PROC CCDM retires the processed observation, chooses another observation at random from the remaining
observations, and repeats the analysis of computing
and updating
. The process continues until all K observations are exhausted.
The randomization process is especially important in the context of simulating an adjusted compound distribution sample when your severity adjustment program uses the aggregate adjusted severity to adjust the next severity value. For an illustration of the need to randomize in such cases for the collective risk simulation mode, see the next section.
Illustration of the Need to Randomize the Order of Severity Draws
This section uses the example in the section Illustration of the Aggregate Adjusted Loss Simulation Process, but with the following PROC CCDM step:
proc ccdm data=<data table name> adjustedseverity=amountPaid
severityest=<severity parameter estimates data table>
countstore=<count model store>;
severitymodel <severity distribution name(s)>;
if (_adjsevsum_ > 15000) then
amountPaid = 0;
else do;
penaltyFactor = MIN(3, 15000/(15000 - _adjsevsum_));
amountPaid = MAX(0, _sev_ - deductible * penaltyFactor);
end;
run;
The severity adjustment statements in the preceding code compute the value of AmountPaid by using the collective risk simulation mode and the following provisions in the insurance policy:
There is a limit of 15,000 on the total amount that can be paid in a year to the group of policyholders that is being simulated. The amount of payment for each loss event depends on the total amount of payments before that loss event.
The penalty for incurring more losses is imposed in the form of an increased deductible. In particular, the deductible is increased by the ratio of the maximum cumulative payment (15,000) to the amount that remains available to pay for future losses in the year. The factor by which the deductible can be raised has a limit of three.
This example illustrates only step 3 of the simulation process, where randomization is done. It assumes that step 2 of the simulation process is identical to step 2 of the example in the section Illustration of the Aggregate Adjusted Loss Simulation Process. At the beginning of step 3, let the severity draws from all the observations be as shown in the _SEV_ column in the following table:
Obs age gender carType deductible count _sev_
1 30 2 1 250 2 350 2100
2 25 1 2 500 1 4500
3 45 2 2 100 2 700 4300
4 33 1 1 200 3 600 1500 950
5 50 1 1 200 0
If the order of these eight draws is not randomized, then all the severity draws for the first observation are adjusted before all the severity draws of the second observation, and so on. The execution of the severity adjustment program leads to the following sequence of values for AmountPaid:
Obs deductible _sev_ _adjsevsum_ penaltyFactor amountPaid
1 250 350 0 1 100
1 250 2100 100 1.0067 1848.32
2 500 4500 1948.32 1.1493 3925.36
3 100 700 5873.68 1.6436 535.64
3 100 4300 6409.32 1.7461 4125.39
4 200 600 10534.72 3 0
4 200 1500 10534.72 3 900
4 200 950 11434.72 3 350
The preceding sequence of simulating loss events results in a cumulative payment of 11,784.72.
If the sequence of draws is randomized over observations, then the computation of the cumulative payment might proceed as follows for one instance of randomization:
Obs deductible _sev_ _adjsevsum_ penaltyFactor amountPaid
2 500 4500 0 1 4000
1 250 350 4000 1.3636 9.09
3 100 700 4009.09 1.3648 563.52
4 200 950 4572.61 1.4385 662.30
4 200 1500 5234.91 1.5361 1192.78
1 250 2100 6427.69 1.7498 1662.54
4 200 600 8090.24 2.1708 165.83
3 100 4300 8256.07 2.2242 4077.58
In this example, a policyholder is identified by the value in the Obs column. As the table indicates, PROC CCDM randomizes the order of loss events not only across policyholders but also across the loss events that a given policyholder incurs. The particular sequence of loss events that is shown in the table results in a cumulative payment of 12,333.65. This differs from the cumulative payment that results from the previously considered nonrandomized sequence of loss events, which tends to penalize the fourth policyholder by always processing her payments after all other payments, with a possibility of underestimating the total amount paid. This comparison not only illustrates that the order of randomization affects the aggregate adjusted loss sample but also corroborates the arguments about the importance of order randomization that are made at the beginning of the section Randomizing the Order of Severity Draws across Observations of a Scenario.