The CCDM Procedure
PROC CCDM Statement
PROC CCDM options;
The PROC CCDM statement invokes the procedure. You can specify the following options, which are listed in alphabetical order.
-
ADJUSTEDSEVERITY=symbol-name
ADJSEV=symbol-name
ADJUSTEDSEVERITY=(symbol-name <…symbol-name>)
ADJSEV=(symbol-name <…symbol-name>) names one or more symbols that represent an adjusted severity value in the SAS programming statements that you specify. If you specify more than one symbol-name, then separate them with spaces and enclose them in parentheses. Each symbol-name is a SAS name that conforms to the naming conventions of a SAS variable. For more information, see the section Programming Statements.
- COUNTSTORE=CAS-libref.data-table
-
names the input data table that contains all the information about the frequency (count) model in the form of an item store. The CNTSELECT procedure generates this item store data table when you use the STORE= option.
CAS-libref.data-table is a two-level name, where CAS-libref refers to the
casliband session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The CAS-libref must be identical to the CAS-libref that you specify in the DATA= option.The exogenous variables in the frequency model, if any, are deduced from this item store. The DATA= data table must contain all those variables.
If you specify a BY statement in the PROC CNTSELECT step that creates the COUNTSTORE= item store, then you must specify an identical BY statement in the PROC CCDM step.
You must specify this option if you do not specify the COUNTMODEL or EXTERNALCOUNTS statement. This option is ignored if you specify the EXTERNALCOUNTS statement, because PROC CCDM does not need to simulate frequency counts internally when you specify externally simulated counts.
- DATA=CAS-libref.data-table
-
names the input data table for PROC CCDM to use. CAS-libref.data-table is a two-level name, where
- CAS-libref
refers to a collection of information that is defined in the LIBNAME statement and includes the
caslib, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about CAS-libref, see the section Using CAS Sessions and CAS Engine Librefs.- data-table
specifies the name of the input data table.
The DATA= data table specifies information about the scenario for which you want to estimate the aggregate loss distribution. It is expected to contain the values of regression variables in the frequency or severity model and severity adjustment variables that you use in the programming statements.
-
IGNORECOV
IGNOREPARMCOVARIANCE -
specifies that the covariance estimates be ignored for perturbation analysis. If you specify this option, then even if the covariance estimates of severity or frequency model parameters are valid, they are not used for perturbation analysis, and only standard errors are used. This enables you to force independence among parameters. However, perturbation analysis based on covariance estimates is more accurate than perturbation analysis based on the standard errors alone.
For more information about parameter perturbation analysis, see the section Parameter Perturbation Analysis.
-
LEFTTRUNCATION=number
LTRUNC=number -
specifies the left-truncation threshold for truncating a severity distribution.
Specifying this option, and optionally the RIGHTTRUNCATION= option, enables you to simulate from a truncated severity distribution. For more information, see the section Making a Random Draw from a Truncated Severity Distribution.
-
MAXCOUNTDRAW=number
MAXCOUNT=number -
specifies an upper limit on the number of loss events (count) that is used for simulating one aggregate loss sample point. If number is equal to
, then any count greater than
is assumed to be equal to
, and only
severity draws are made to compute one point in the aggregate loss sample.
If you specify this option and also specify the COUNTSTORE= item store or the COUNTMODEL statement, then the limit is applied to each count that PROC CCDM randomly draws from the count distribution. Any count draw larger than number is replaced by number.
If you specify this option and also specify the EXTERNALCOUNTS statement, then the limit is applied to each observation in the DATA= data table, and any value of the COUNT= variable larger than number is replaced by number.
If you do not specify this option, then a default value of 1,000 is used.
If you specify a number significantly larger than 1,000, then PROC CCDM might take a very long time to complete the simulation, especially when some counts are closer to the limit.
- NOPRINT
turns off all displayed output. If you specify this option, then PROC CCDM ignores any value that you specify for the PRINT= option.
-
NPERTURBEDSAMPLES=number
NPERTURB=number -
performs parameter perturbation analysis, where number determines how many perturbed parameter sets are generated. A separate full sample is simulated for each set of perturbed parameter values. The summary statistics and percentiles are computed for each such perturbed sample, and their values are aggregated across the samples to compute the mean and standard deviation of each summary statistic and percentile.
The parameter perturbation analysis makes random draws of parameter values from a multivariate normal distribution if the covariance estimates of the parameters are available. For the multivariate normal distribution of severity model parameters, PROC CCDM attempts to read the covariance estimates from the SEVERITYEST= data table or the SEVERITYSTORE= item store. For the multivariate normal distribution of count model parameters, PROC CCDM attempts to read the covariance estimates from the item store that is specified in the COUNTSTORE= option. If covariance estimates are not available or valid, or if you specify the IGNORECOV option, then for each parameter, PROC CCDM makes a random draw from the univariate normal distribution whose mean and standard deviation are equal to that parameter’s point estimate and standard error, respectively. The standard errors for severity model parameters are read from either the SEVERITYEST= data table or the SEVERITYSTORE= item store. The standard errors for count model parameters are either read from the COUNTSTORE= item store or you must specify them in the COUNTMODEL statement. If neither covariance nor standard error estimates are available for any severity or frequency model parameter, then perturbation analysis is not conducted.
If you specify the PRINT=ALL or PRINT=PERTURBSUMMARY option, then a summary of the perturbation analysis is printed for the core summary statistics and the percentiles of the aggregate loss distribution. If you specify the OUTSUM statement, then the requested summary statistics are written to the OUTSUM= data table for each perturbed sample. You can also optionally request that each perturbed sample be written in its entirety to the OUT= data table by specifying the PERTURBOUT option in the OUTPUT statement.
For more information about parameter perturbation analysis, see the section Parameter Perturbation Analysis.
-
NREPLICATES=number
NREP=number -
specifies a number that controls the size of the compound distribution sample that PROC CCDM simulates. The number is interpreted differently based on whether you specify the EXTERNALCOUNTS statement.
If you do not specify the EXTERNALCOUNTS statement, then the sample size is equal to number. If you do not specify this option, then a default value of 100,000 is used.
If you specify the EXTERNALCOUNTS statement, then the number of replicates that you specify in the DATA= data table is multiplied by number to get the total size of the compound distribution sample. If you do not specify this option, then a default value of 1 is used.
- ONEADJSEVCOMPAT
-
specifies that the output of the case of single adjusted severity symbol be prepared by using the same method as used by the versions of the CCDM procedure prior to SAS Econometrics 8.5. In particular, specifying this option causes the following output-related changes, where
symdenotes the name of the single adjusted severity symbol that you specify in the ADJUSTEDSEVERITY= option:If you specify the OUTPUT statement, then the default name of the adjusted severity variable in the OUTSAMPLE= table is _AGGADJSEV_ instead of
sym.If you specify the OUTSUM statement, then the default value in the _SAMPLEVAR_ column of the OUTSUM= table is _AGGADJSEV_ instead of
symfor the rows that correspond to the aggregate adjusted severity.The ODS path for the adjusted severity group’s output table or plot uses AggregateAdjustedLoss instead of
sym.The label for the adjusted severity group in the table of contents of a SAS client is "Aggregate Adjusted Loss" instead of "Aggregate Adjusted Loss
sym."The title of the SummaryStatistics ODS table is "Adjusted Sample Summary Statistics" instead of "Summary Statistics for
sym."The title of the PerturbedSummary ODS table is "Adjusted Sample Perturbation Analysis" instead of "Sample Perturbation Analysis for
sym."The title of the Percentiles ODS table is "Adjusted Sample Percentiles" instead of "Percentiles for
sym."The title of the PerturbedPctlSummary ODS table is "Adjusted Sample Percentile Perturbation Analysis" instead of "Percentile Perturbation Analysis for
sym."The titles of plots use just "Adjusted Sample" instead of "Adjusted Sample
sym."The X-axis variable label in plots is just "Aggregate Adjusted Loss" instead of "Aggregate Adjusted Loss
sym."
When you do not specify this option, the style of the output for the case of a single adjusted severity symbol is no different from the style of the output for the case of more than one adjusted severity symbol.
- OUTDRAW=CAS-libref.data-table
-
specifies the output data table to contain the individual severity random draws. CAS-libref.data-table is a two-level name, where CAS-libref refers to the
casliband session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The CAS-libref must be identical to the CAS-libref that you specify in the DATA= option.For more information about the variables in the OUTDRAW= data table, see the section OUTDRAW= Data Table.
- PCTLDEF=percentile-method
specifies the method of computing the percentiles of the compound distribution. The percentile-method can be 1, 2, 3, 4, or 5. By default, PCTLDEF=5. For more information, see the description of the PCTLDEF= option in the UNIVARIATE procedure in the Base SAS Procedures Guide: Statistical Procedures.
-
PERTURBMODE=mode
PERTURBMETHOD=mode -
specifies a method to generate perturbed samples.
This option has no effect if you do not specify the NPERTURBEDSAMPLES= option or if you specify the EXTERNALCOUNTS statement.
Let P and W denote the values of NPETURBEDSAMPLES= option and number of worker nodes in the CAS session that the procedure is running in, respectively. Also, let N denote the value of the NREPLICATES= option.
You can specify one of the following modes:
- AUTO | SYNC | 1
-
generates perturbed samples in either distributed or partitioned manner depending on the values of P and W as follows:
If
, then each perturbed sample is distributed across all W workers, such that each worker generates a subsample of approximately
points of each perturbed aggregate loss sample. If you specify the PERTURBOUT option in the OUTPUT statement, then observations that have the same value of the _DRAWID_ variable are distributed across all worker nodes. The controller node needs to gather the subsamples from all the worker nodes in order to compute the percentiles, and perturbed samples are generated sequentially for each severity model within each BY group.
If
, then perturbed samples are generated in a partitioned manner such that each worker node generates approximately
perturbed samples, each of size N. If you specify the PERTURBOUT option in the OUTPUT statement, then observations that have the same value of the _DRAWID_ variable are available on only one of the worker nodes. There is no need for communicating any sample data to the controller node, because each worker node has enough data to compute the percentiles of each perturbed sample. The controller node gathers only the summary statistics and percentiles of perturbed samples from all worker nodes. Further, each worker node can simulate its assigned
perturbed samples independently of other worker nodes. Together, these two aspects make the perturbation analysis more efficient.
- DIST | SYNCDIST | 2
always generates perturbed samples in distributed manner irrespective of the values of P and W. This option’s behavior is equivalent to the behavior of the AUTO option for the
case.
It is recommended that you specify the AUTO method. By default, PERTURBMETHOD=DIST to ensure that the current release of PROC CCDM produces, by default, the same perturbation results as releases prior to SAS Econometrics 8.4.
-
PLOTS <(global-plot-options)> =plot-request-option
PLOTS <(global-plot-options)> =(plot-request-option …plot-request-option) -
specifies the desired graphical output.
By default, the CCDM procedure produces no graphical output.
When you request plots by using this option, PROC CCDM creates a plot data table on the server. This table contains an estimate of the empirical distribution function (EDF) for a carefully selected subsample of the simulated aggregate loss sample. PROC CCDM brings the plot data table back to the client machine where your SAS session is running and uses it to create all the plots that you specify.
You can specify the following global-plot-options:
- ONLY
turns off the default graphical output and creates only the requested plots.
- NOBS=number
-
specifies the approximate number of observations in the plot data table. The larger the number, the more accurate are the plots, but also the larger the size of the data that need to be brought back to the client machine. A larger plot data table not only requires more time and resources to prepare but can also cost you additional money if the data movement on the client-server connection is metered.
If you specify a nonzero number, then it must be greater than 10. If you do not specify this option or if you specify NOBS=0, then a default value of 5,000 is used.
-
NADJSEVTOPLOT=k
NADJSEV=k -
specifies the number of aggregate adjusted loss samples for which to create the requested plots. PROC CCDM creates plots for the first k adjusted severity symbols that you specify in the ADJUSTEDSEVERITY= option.
This option is ignored if you do not specify the ADJUSTEDSEVERITY= option and the programming statements to adjust the simulated severity values.
If you specify m symbols in the ADJUSTEDSEVERITY= option and if k
, then PROC CCDM creates the requested plots for all adjusted severity symbols.
Note that PROC CCDM prepares a separate plot data table for each of the k aggregate adjusted loss samples, so the amount of data that PROC CCDM needs to bring back to the client machine increases linearly with the value of k.
By default, PROC CCDM does not create plots for aggregate adjusted loss samples—that is, NADJSEVTOPLOT=0 by default.
You control which plots to create for each aggregate loss sample by specifying the plot-request-options. If you specify more than one plot-request-option, then separate them with spaces and enclose them in parentheses. You can specify the following plot-request-options:
- ALL
displays all the graphical output.
-
CONDITIONALDENSITY (conditional-density-plot-options)
CONDPDF (conditional-density-plot-options) -
creates a group of plots of the conditional density functions estimates. The group contains at most three plots, each conditional on the value of the aggregate loss being in one of the three regions that are defined by the quantiles that you specify in the following conditional-density-plot-options:
- LEFTQ=number
-
specifies the quantile in the range (0,1) that marks the end of the left-tail region. If you specify a value of l for number, then the left-tail region is defined as the set of values that are less than or equal to
, where
is the lth quantile. For the left-tail region, nonparametric estimates of the conditional probability density function
are plotted. The value of
is estimated by the
th percentile of the simulated compound distribution sample.
If you do not specify this option or you specify a missing value for this option, then the left-tail region is not plotted.
- RIGHTQ=number
-
specifies the quantile in the range (0,1) that marks the beginning of the right-tail region. If you specify a value of r for number, then the right-tail region is defined as the set of values that are greater than
, where
is the rth quantile. For the right-tail region, nonparametric estimates of the conditional probability density function
are plotted. The value of
is estimated by the
th percentile of the simulated compound distribution sample.
If you do not specify this option or you specify a missing value for this option, then the right-tail region is not plotted.
You must specify a nonmissing value for at least one of the preceding two options. For the region between the LEFTQ= and RIGHTQ= quantiles, which is referred to as the central or body region, nonparametric estimates of the conditional probability density function
are plotted. If you do not specify a LEFTQ= value, then
is assumed to be 0. If you do not specify a RIGHTQ= value, then
is assumed to be
.
- DENSITY
creates a plot of the nonparametric estimates of the probability density function (in particular, histogram and kernel density estimates) of the compound distribution.
- EDF <(edf-plot-option)>
-
creates a plot of the nonparametric estimates of the cumulative distribution function of the compound distribution.
You can request that the confidence interval be plotted by specifying the following edf-plot-option:
- NONE
displays none of the graphical output. If you specify this option, then it overrides all other plot request options. The default graphical output is also suppressed.
-
PRINT <(global-display-option)> =display-option
PRINT <(global-display-option)> =(display-option …display-option) -
specifies the desired displayed output. If you specify more than one display-option, then separate them with spaces and enclose them in parentheses.
You can specify the following global-display-option:
- ONLY
turns off the default displayed output and displays only the requested output.
You can specify the following display-options:
- ALL
displays all the output.
- NONE
displays no output. If you specify this option, then it overrides all other display options. The default displayed output is also suppressed.
- PERCENTILES
displays the percentiles of the compound distribution sample. This includes all the predefined percentiles and percentiles that you request in the OUTSUM statement.
- PERTURBSUMMARY
displays the mean and standard deviation of the summary statistics and percentiles that are taken across all P samples that the procedure produces by perturbing the model parameters, where P is the value of the NPERTURBEDSAMPLES= option that you specify in the PROC CCDM statement.
- SUMMARYSTATISTICS | SUMSTAT
displays the summary statistics of the compound distribution sample.
If you do not specify the PRINT= option or the ONLY global-display-option, then the procedure behaves as if you have specified PRINT=(SUMMARYSTATISTICS).
-
RIGHTTRUNCATION=number
RTRUNC=number -
specifies the right-truncation limit for truncating a severity distribution.
Specifying this option, and optionally the LEFTTRUNCATION= option, enables you to simulate from a truncated severity distribution. For more information, see the section Making a Random Draw from a Truncated Severity Distribution.
- SEED=number
-
specifies the integer to use as the seed in generating the pseudorandom numbers that are used for simulating severity and frequency values.
If you do not specify the seed, or if number is negative or 0, then PROC CCDM uses the time of day from the computer’s clock as the seed.
- SEVERITYEST=CAS-libref.data-table
-
names the input data table that contains the parameter estimates for the severity models. CAS-libref.data-table is a two-level name, where CAS-libref refers to the
casliband session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The CAS-libref must be identical to the CAS-libref that you specify in the DATA= option. This data table must be an OUTEST= data table that the SEVSELECT procedure produces.The names of the regression variables in the scale regression models, if any, are deduced from this data table. If the severity models are scale regression models, then the DATA= data table must contain all the regressors in the scale regression models.
To ensure that PROC CCDM correctly matches the values of regressors and the values of regression parameter estimates, you need to ensure that names of the regressors in the DATA= data table match the names of the regressors that you specify in the SCALEMODEL statement of the PROC SEVSELECT step that fits the severity models.
If you specify a BY statement in the PROC SEVSELECT step that creates the SEVERITYEST= data table, then you must specify an identical BY statement in the PROC CCDM step.
-
SEVERITYSTORE=CAS-libref.data-table
SEVSTORE=CAS-libref.data-table -
specifies the input data table that contains the context and estimates of the severity models in the form of an item store. Specifying the OUTSTORE= option in a PROC SEVSELECT step creates this item store data table.
CAS-libref.data-table is a two-level name, where CAS-libref refers to the
casliband session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The CAS-libref must be identical to the CAS-libref that you specify in the DATA= option.If your severity models contain classification or interaction effects, then you need to use this option instead of the SEVERITYEST= option to specify the estimates of the severity models. If you specify this option, you cannot specify the SEVERITYEST= option.
If you specify a BY statement in the PROC SEVSELECT step that creates the SEVERITYSTORE= item store, then you must specify an identical BY statement in the PROC CCDM step.
- SIMULATIONMODE=mode
-
specifies the mode of simulating aggregate loss. You can specify one of the following modes:
-
COLLECTIVERISK
CR -
specifies the collective risk mode. Aggregate loss is defined as
, where
denotes a continuous random variable that represents the severity of one loss event and N denotes the discrete random variable that represents the number of loss events you expect to see in a particular time period. The severity random variables
that are associated with all loss events are assumed to be iid. For each random frequency draw N, PROC CCDM makes N random draws from the severity distribution and adds those severity draws to compute one sample point in the aggregate loss sample.
If you specify this mode (or choose it by default) and if you specify the frequency distribution by using the COUNTMODEL statement, then the count distribution must be a discrete distribution.
-
CUSTOM
CSTM -
specifies the custom mode. In this mode, you use programming statements to combine the severity and frequency values according to your business logic in order to compute the aggregate loss. If you specify this mode, then you must also use the ADJUSTEDSEVERITY= option to specify an aggregate loss symbol that your programming statements compute. For each random frequency draw N, PROC CCDM makes one random draw, X, from the severity distribution and sets the keyword symbols _FREQ_ and _SEV_ to N and X, respectively. It then executes your programming statements to compute one sample point in the aggregate loss sample.
This mode is especially useful if you also use the SIMULATEDSYMBOL statement to specify some stochastic variables that you have externally estimated to follow certain parametric probability distributions. For example, the following statements simulate a sample of the total paid loss
paidLossas, where N denotes the average frequency of loss events per exposure and follows a lognormal distribution with
and
, E denotes the estimate of the total exposure in a particular time period and follows a normal distribution
, X denotes the average severity of a loss event, and C denotes the random change in losses from one time period to next and follows a normal distribution
:
proc ccdm severitystore=mycas.sevstore simulationmode=custom adjustedseverity=paidLoss; countmodel lognormal(3,0.5); simulatedsymbol E ~ normal(1000, 75) C ~ normal(500, 10); paidLoss = _freq_ * E * _sev_ + C; run; -
PUREPREMIUM
PP specifies the pure premium mode. Aggregate loss is defined as
, where N denotes the average frequency of loss events you expect to see in a particular time period and X denotes a continuous random variable that represents the average severity of the loss events in that time period. For each random frequency draw N, PROC CCDM makes one random draw, X, from the severity distribution and multiplies N and X to compute one sample point in the aggregate loss sample.
By default, SIMULATIONMODE=COLLECTIVERISK.
-
COLLECTIVERISK
- TRUNCATEZEROS
-
truncates (removes) zero-valued aggregate loss values from the aggregate loss sample.
The value of the aggregate loss is 0 when the sum of randomly drawn counts for all observations is 0 for the internally simulated counts or when the total count for a replication is 0 in the external counts data table.
If you omit this option, then PROC CCDM keeps the zero-valued aggregate losses in the output sample.
- VARDEF=divisor
-
specifies the divisor to use in the calculation of variance, standard deviation, kurtosis, and skewness of the compound distribution sample. If the sample size is N, then you can specify one of the following values for the divisor:
By default, VARDEF=DF. For more information, see the section Descriptive Statistics.