Sampling and Partitioning Action Set

k-Fold Partitioning

This section contains PROC CAS code.

Note: Input data must be accessible in your CAS session, either as a CAS table or as a transient-scope table. A CAS table has a two-level name: the first level is your CAS engine libref, and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 2, Shared Concepts. A transient-scope table is called directly from the action and exists in memory for the duration of the action. For more information about accessing data, see SAS Viya: System Programming Guide. For more information about PROC CAS and programming in CASL, see SAS Cloud Analytic Services: CASL Programmer’s Guide and SAS Cloud Analytic Services: CASL Reference.

This example performs k-fold partitioning of 5,960 fictitious mortgages, with the BY variable bad used as the stratum. The input data table mycas.hmeq includes information about fictitious mortgages. Each observation represents an applicant for a home equity loan, and all applicants have an existing mortgage.

You can load the sampsio.hmeq data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step:

data mycas.hmeq;
   set sampsio.hmeq;
run;

This DATA step assumes that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

The following statements load the sampling action set and then use the kfold action to partition the mycas.hmeq data table for each stratum:

proc cas;
   loadactionset "sampling";
   action kfold result=r/table={name="hmeq",groupby={"BAD"}}
      k=10  seed=123
      output={casout={name="out",replace="TRUE"},
              copyvars={"bad","job","reason","loan","value","delinq","derog"},
              foldname='myfold'};
   run;
   print r.KFOLDFreq; run;
quit;

proc print data=mycas.out(obs=20);
run;

The table parameter names the input data table to be analyzed. The groupby subparameter in the table parameter names the variables to be used for stratification. The k parameter requests that the input data be partitioned to 10 folds. The seed parameter specifies 123 as the random seed to be used in the partitioning process.

The output parameter requests that the sampled data be stored in a table named mycas.out, the copyvars parameter lists the variables to be copied from mycas.hmeq to mycas.out, and the foldname parameter renames the generated column that is named _Fold_ by default to myfold.

Output 22.1.1 shows the fold frequency information in each level of stratification, as indicated by the groupby variable BAD in the mycas.hmeq data table.

Output 22.1.1: Frequency Information Table

KFOLDFreq: Results from sampling.kfold

k-Fold Partition Frequency
IndexBADNumber
of Obs
Fold Size 1Fold Size 2Fold Size 3Fold Size 4Fold Size 5Fold Size 6Fold Size 7Fold Size 8Fold Size 9Fold Size 10
004771478477477477477477477477477477
111189119118119119119119119119119119


Output 22.1.2 shows the first 20 output sample observations in mycas.out; the myfold column shows which fold the observation belongs to.

Output 22.1.2: Sample Output with Fold Indicator

ObsBADJOBREASONLOANVALUEDELINQDEROGmyfold
11OtherHomeImp110039025005
21  1500...8
31OtherHomeImp180057037233
41SalesHomeImp200062250007
51OtherHomeImp200055000001
61OtherHomeImp220034687107
71OtherHomeImp230040150003
81ProfExeHomeImp240073395015
91OtherHomeImp2400171800010
101 HomeImp250020200008
110ProfExeHomeImp250078600006
121ProfExeDebtCon2900113000016
131OtherHomeImp290067996032
141OtherHomeImp300020300003
151OtherHomeImp3000193500003
161OtherHomeImp3000141000010
170MgrHomeImp3000715002.1
180  310070400..1
191OtherHomeImp320040834006
201MgrHomeImp3200.2.2


k-Fold Partitioning

This section contains Lua code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the hmeq data to the comma-separated-value (CSV) file hmeq.csv and then use the following code to load the CSV file into CAS:

s:loadtable{casLib="casuser", path="hmeq.csv"}

For more information about coding in Lua, see Getting Started with SAS Viya for Lua and SAS Viya: System Programming Guide.

The following code loads the sampling action set and then performs k-fold partitioning on the hmeq data table:

s:loadactionset{actionset="sampling"}
s:kfold{table={name='hmeq',groupby={'bad'}},
            k=10, seed=123,
             outputTables={names={KFOLDFreq='kfoldfreq'}},
             output={casout={name="out", replace="TRUE"},
                     copyvars={"bad","job","reason","loan","value","delinq","derog"}}}
s:fetch{table={name="out"},to=20}

k-Fold Partitioning

This section contains Python code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the hmeq data to the comma-separated-value (CSV) file hmeq.csv and then use the following code to load the CSV file into CAS:

s.upload_file('hmeq.csv')

For more information about coding in Python, see Getting Started with SAS Viya for Python and SAS Viya: System Programming Guide.

The following code loads the sampling action set and then performs k-fold partitioning on the hmeq data table:

s.loadactionset(actionset="sampling")
s.kfold(display={"names":"KFOLDFreq"},
              output={"casOut":{"name":"out", "replace":True}, "copyVars":"ALL"},
              k=10, seed=123,
              table={"name":"hmeq", "groupBy":{"bad"}},
              outputTables={"names":{"KFOLDFreq"},"replace":True})

kfold_out=s.CASTable('out')
print(kfold_out.fetch(to=20))


k-Fold Partitioning

This section contains R code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the hmeq data to the comma-separated-value (CSV) file hmeq.csv and then use the following code to load the CSV file into CAS:

m <- cas.read.csv(s, "hmeq.csv", casOut=list(name="hmeq"))

For more information about coding in R, see Getting Started with SAS Viya for R and SAS Viya: System Programming Guide.

The following code loads the sampling action set and then performs k-fold partitioning on the hmeq data table:

cas.read.csv(s,
          "hmeq.csv",
          header = TRUE,
          casOut = list(name = "hmeq", replace = TRUE))

loadActionSet(s, 'sampling')
cas.sampling.stratified(s,
           table        = list(name = "hmeq", groupby = "bad"),
           k            = 10,
           seed         = 123,
           partind      = TRUE,
           seed         = 10,
           output       = list(casout   = list(name = "out", replace = "TRUE"),
                               copyvars = list("job", "reason", "loan", "value",
                                               "delinq", "derog")),
           outputTables = list(names = "KFOLDFreq", replace = TRUE)
)
KFOLDFreq <- defCasTable(s, "KFOLDFreq")
KFOLDFreq
out3 <- defCasTable(s, "out")
out3
Last updated: September 13, 2022