CCDM Procedure

Estimating a Simple Compound Distribution Model

This example illustrates the simplest use of PROC CCDM. Assume that you are an insurance company that has used its historical data about the number of losses per year and the severity of each loss to determine that the Poisson distribution is the best distribution for the loss frequency and that the gamma distribution is the best distribution for the severity of each loss. Now, you want to estimate the distribution of an aggregate loss to determine the worst-case loss that your policyholders can incur in a year. In other words, you want to estimate the compound distribution of upper S equals sigma-summation Underscript i equals 1 Overscript upper N Endscripts upper X Subscript i, where the loss frequency, N, follows the fitted Poisson distribution and the severity of each loss event, upper X Subscript i, follows the fitted gamma distribution.

To illustrate, let the historical count and severity data be stored in the data sets Work.ClaimCount and Work.ClaimSev, respectively. The following two DATA steps simulate a frequency sample from a Poisson distribution and a severity sample from the gamma distribution, respectively:

/* Simulate data for an intercept-only Poisson count model */
data claimcount(keep=numLosses);
   call streaminit(12345);
   label numLosses='Number of Loss Events in a Year';
   Lambda = 2;
   do n = 1 to 500;
      numLosses = rand('POISSON',Lambda);
      output;
   end;
run;

/* Simulate data for a gamma severity model */
data claimsev(keep=lossValue);
   call streaminit(67890);
   label lossValue='Severity of a Loss Event';
   Theta = 1000;
   Alpha = 2;
   do n = 1 to 500;
      lossValue = quantile('Gamma', rand('UNIFORM'), Alpha, Theta);
      output;
   end;
run;

You can load the data sets Work.ClaimCount and Work.ClaimSev into data tables in your CAS session by using your CAS engine libref with the following DATA steps:

/* Load the data */
data mylib.claimcount;
   set claimcount;
run;
data mylib.claimsev;
   set claimsev;
run;

These statements assume that your CAS engine libref is named mylib, as in the section Using CAS Sessions and CAS Engine Librefs, but you can substitute any appropriately defined CAS engine libref.

To create model parameter estimates in a format that PROC CCDM can use, you need to use the following PROC CNTSELECT and PROC SEVSELECT steps to fit and store the parameter estimates of the frequency and severity models:

/* Fit an intercept-only Poisson count model and
   write estimates to an item store */
proc cntselect data=mylib.claimcount store=mylib.countStorePoisson;
   model numLosses= / dist=poisson;
run;

/* Fit severity models and write estimates to a data table */
proc sevselect data=mylib.claimsev criterion=aicc
     outest=mylib.sevest covout;
   loss lossValue;
   dist _predefined_;
run;

The STORE= option in the PROC CNTSELECT statement saves the count model information, including the parameter estimates, in the item store mylib.CountStorePoisson. An item store contains the model information in a binary format that cannot be modified after it is created. You can examine the contents of an item store that is created by a PROC CNTSELECT step by using the VIEWSTORE statement in another PROC CNTSELECT step. (For more information, see Chapter 10, CNTSELECT Procedure.)

The OUTEST= option in the PROC SEVSELECT statement stores the estimates of all the fitted severity models in the data table mylib.SevEst. You can verify that the best severity model that the PROC SEVSELECT step chooses is the gamma distribution model.

You can now submit the following PROC CCDM step to simulate an aggregate loss sample of size 10,000 by specifying the count model’s item store in the COUNTSTORE= option and the severity model’s data table of estimates in the SEVERITYEST= option:

/* Simulate and estimate Poisson-gamma compound distribution model */
proc ccdm countstore=mylib.countStorePoisson severityest=mylib.sevest
          seed=13579 nreplicates=10000 plots=(edf density)
          print=(summarystatistics percentiles);
   severitymodel gamma;
   output out=mylib.aggregateLossSample samplevar=aggloss;
   outsum out=mylib.aggregateLossSummary mean stddev skewness kurtosis
          p01 p05 p95 p995=var pctlpts=90 97.5;
run;

The SEVERITYMODEL statement requests that an aggregate sample be generated by compounding only the gamma distribution and the frequency distribution. Specifying the SEED= value helps you get an identical sample each time you execute this step, provided that you use the same execution environment. In the single-machine mode of execution, the execution environment is the combination of the operating environment and the number of threads that are used for execution. In distributed computing mode, the execution environment is the combination of the operating environment, the number of worker nodes, and the number of threads that are used for execution on each worker node.

Upon completion, PROC CCDM creates the two output data tables that you specify in the OUT= options in the OUTPUT and OUTSUM statements. The data table mylib.AggregateLossSample contains 10,000 observations such that the value of the variable AggLoss in each observation represents one possible aggregate loss value that you can expect to see in one year. Together, the set of the 10,000 values of AggLoss represents one sample of compound distribution. PROC CCDM uses this sample to compute the empirical estimates of various summary statistics and percentiles of the compound distribution. The data table mylib.AggregateLossSummary contains the estimates of mean, standard deviation, skewness, and kurtosis that you specify in the OUTSUM statement. It also contains the estimates of the 1st, 5th, 90th, 95th, 97.5th, and 99.5th percentiles that you specify in the OUTSUM statement. The value-at-risk (VaR) is an aggregate loss value such that there is a very low probability that an observed aggregate loss value exceeds the VaR. One of the commonly used probability levels to define VaR is 0.005, which makes the 99.5th percentile an empirical estimate of VaR. Hence, the OUTSUM statement of this example stores the 99.5th percentile in a variable named VaR. VaR is one of the widely used measures of worst-case risk.

Some of the default output and some of the output that you have requested by specifying the PRINT= option are shown in Figure 1.

Figure 1: Information, Summary Statistics, and Percentiles of the Poisson-Gamma Compound Distribution

The CCDM Procedure
 
Severity Model: Gamma
Count Model: Poisson

Compound Distribution Information
Severity ModelGamma Distribution
Count ModelPoisson Model in Item Store COUNTSTOREPOISSON

The CCDM Procedure
 
Severity Model: Gamma
Count Model: Poisson

Sample Summary Statistics
Mean4025.3Median3360.3
Standard Deviation3443.7Interquartile Range4569.8
Variance11859027.2Minimum0
Skewness1.14279Maximum25106.0
Kurtosis1.71709Sample Size10000

Sample Percentiles
PercentileValue
10
50
251345.4
503360.3
755915.1
908691.7
9510602.8
97.512294.1
9914720.1
99.516055.7
Percentile Method = 5


The "Sample Summary Statistics" table indicates that for the given parameter estimates of the Poisson frequency and gamma severity models, you can expect to see a mean aggregate loss and a median aggregate loss in a year. The "Sample Percentiles" table indicates that there is a 0.5% chance that the aggregate loss exceeds the 99.5th percentile, which is the VaR estimate, and a 2.5% chance that the aggregate loss exceeds the 97.5th percentile. These summary statistics and percentile estimates provide a quantitative picture of the compound distribution.

You can also visually analyze the compound distribution by examining the plots that PROC CCDM creates when you specify the PLOTS= option. The first plot in Figure 2 shows the empirical distribution function (EDF), which is a nonparametric estimate of the cumulative distribution function (CDF). The second plot shows the histogram and the kernel density estimate, which are nonparametric estimates of the probability density function (PDF).

Figure 2: Nonparametric CDF and PDF Plots of the Poisson-Gamma Compound Distribution

Nonparametric CDF and PDF Plots of the Poisson-Gamma Compound Distribution
External File:images/cdmgs1o1g1.png


The plots confirm the right skew that is indicated by the estimate of skewness in Figure 1 and a relatively fat tail, which is indicated by comparing the maximum and the 99.5th percentiles in Figure 1.

Last updated: July 09, 2026