The PARTITION Procedure
Getting Started: PARTITION Procedure
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This example performs stratified partitioning of 5,960 fictitious mortgages, with the BY variable BAD used as the stratum. The input data table mycas.hmeq includes information about the fictitious mortgages. Each observation represents an applicant for a home equity loan, and all applicants have an existing mortgage.
You can load the sampsio.hmeq data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step. This DATA step assumes that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
data mycas.hmeq;
set sampsio.hmeq;
run;
The following statements perform the partitioning:
proc partition data=mycas.hmeq samppct=10 samppct2=20 seed=1234 partind;
by BAD;
output out=mycas.out1 copyvars=(BAD loan derog mortdue value yoj
delinq clage ninq clno debtinc);
run;
The SAMPPCT=10 option requests that 10% of the input data be included in the training partition, and the SAMPPCT2=20 option requests that 20% of the input data be included in the testing partition. The SEED= option specifies 1234 as the random seed to be used in the partitioning process. The PARTIND option requests that the output data table, mycas.out1, include an indicator that shows whether each observation is selected to a partition (1 for training or 2 for testing) or not (0). The binary BY variable BAD indicates whether an applicant eventually defaulted or was ever seriously delinquent. The BY statement triggers stratified sampling, which enables you to sample each subpopulation in the BY variable (stratum) independently. The OUTPUT statement creates a new data table to contain the variables from the input data table that are listed in the COPYVARS= option and the partition indicator. The displayed output includes a frequency table (Figure 1) that shows the frequency of observations in each level of BAD.
Figure 1: Frequency Information Table
| Stratified Sampling Frequency | ||||
|---|---|---|---|---|
| Index | BAD | Number of Obs | Sample Size 1 | Sample Size 2 |
| 0 | 0 | 4771 | 477 | 954 |
| 1 | 1 | 1189 | 119 | 238 |