Active Learning Action Set

Analyzing German Credit Data Using the uncertaintySampling Action

This section contains PROC CAS code.

Note: Input data must be accessible in your CAS session, either as a CAS table or as a transient-scope table. A CAS table has a two-level name: the first level is your CAS engine libref, and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 2, Shared Concepts (SAS Visual Data Mining and Machine Learning: Procedures). A transient-scope table is called directly from the action and exists in memory for the duration of the action. For more information about accessing data, see SAS Viya: System Programming Guide. For more information about PROC CAS and programming in CASL, see SAS Cloud Analytic Services: CASL Programmer’s Guide and SAS Cloud Analytic Services: CASL Reference.

This example trains the model by using German Credit Benchmark data, which are available in the sampsio.dmagecr data set. This data set contains 1,000 observations, each of which contains an applicant’s information, including the applicant’s credit rating (GOOD or BAD). The binary target is named GOOD_BAD. Other input variables are Checking, Duration, History, and so on.

The contents of the sampsio.dmagecr data set are described at http://support.sas.com/documentation/cdl/en/emgs/59885/HTML/default/a001026918.htm.

The following DATA step creates the ID variable alucIndex and loads the sampsio.dmagecr data set into your CAS session:

data mycas.dmagecr;
   set sampsio.dmagecr;
   alucIndex = _n_;
run;

These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

The following code trains a random forest model and predicts the propensity of the binary target GOOD_BAD:

proc forest data=mycas.dmagecr outmodel=mycas.forest_model seed=12345;
   input age amount checking coapp depends duration employed existcr foreign
   history housing installp job marital other property resident savings
   telephon/ level = interval;
   target good_bad /level=nominal;
   output out=mycas.score_at_runtime copyvars=(_ALL_);
run;

The following code loads the uncertaintySampling action from the activeLearn action set to perform active learning based on the output from the random forest model that you trained in the previous step:

proc cas;
   action activeLearn.uncertaintySampling  /
   table={name='score_at_runtime'},
   topK = 183,
   id='alucIndex',
   inputs = {'age' 'amount' 'checking' 'coapp' 'depends' 'duration' 'employed'
             'existcr' 'foreign' 'history' 'housing' 'installp' 'job' 'marital'
             'other' 'property' 'resident'  'savings' 'telephon' }
   selectQuery = {method='uncertainty', includeAllData=TRUE, probVar = 'P_good_badgood'}
   output = {casout= {name='queryout', replace='true'} copyvars='ALL'}
   target = 'good_bad';
run;
quit;

The table parameter names the input data table to be analyzed. The inputs parameter lists all the input variables to be used in the training. The target parameter specifies the variable to predict.

Output 3.1.1 displays the "Model Information" table, which shows the method that you specify in the selectQuery parameter as well as the value in the topK parameter.

Output 3.1.1: Model Information

Results from activeLearn.uncertaintySampling

Model Information
Query TypeUncertainty
topK183


Output 3.1.2 displays the "Number of Observations" table. This table shows that all 100 observations in the data set are used in the analysis.

Output 3.1.2: Number of Observations

Results from activeLearn.uncertaintySampling

Number of Observations
Observations typesN
Number of Observations Read1000
Number of Observations Used1000
Number of Labeled Observations1000
Number of Unlabeled Observations0
Number of Labeled Observations Used1000
Number of Unlabeled Observations Used0


Analyzing German Credit Data Using the uncertaintySampling Action

This section contains Lua code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:

s:loadtable{casLib="casuser", path="dmagecr.csv"}

For more information about coding in Lua, see Getting Started with SAS Viya for Lua and SAS Viya: System Programming Guide.

The following DATA step creates an ID variable alucIndex and load the sampsio.dmagecr data set into your CAS session:

s:dataStep_runCode{single='no', code=[[
data dmagecr;
   set dmagecr;
   alucIndex = _n_;
run;
]] }

The following code trains random forest model using the German Credit Benchmark data set.

s:loadactionset{actionset='decisionTree'}
r=s:decisionTree_forestTrain {
     table     = {name    = "dmagecr"},
     seed      = 12345,
     casOut    = {name    = "forest_model"},
     inputs    = {{name   = "age"},
                  {name   = "amount"},
                  {name   = "checking"},
                  {name   = "coapp"},
                  {name   = "depends"},
                  {name   = "duration"},
                  {name   = "employed"},
                  {name   = "existcr"},
                  {name   = "foreign"},
                  {name   = "history"},
                  {name   = "housing"},
                  {name   = "installp"},
                  {name   = "job"},
                  {name   = "marital"},
                  {name   = "other"},
                  {name   = "property"},
                  {name   = "resident"},
                  {name   = "savings"},
                  {name   = "telephon"}},
     target = "good_bad"  }

The following code scores the table dmagecr and predicts the propensity of the binary target GOOD_BAD

s:loadactionset{actionset='decisionTree'}
r=s:decisionTree_forestScore {
   table     = {name    = "dmagecr"},
   casOut    = {name    = "score_at_runtime"},
   copyVars  = {'age','amount','checking','coapp','depends','duration','employed','existcr',
                'foreign','history','housing','installp','job','marital','other','property',
                'resident','savings','telephon','alucIndex','good_bad'},
   modelTable = {name    = "forest_model"},
   encodeName = true }

The following code loads the uncertaintySampling action in the activeLearn action set to perform active learning based on the output from the random forest model:

s:loadactionset{actionset='activeLearn'}
r=s:uncertaintySampling{
     table={name    = "score_at_runtime"},
     topK = 183,
     id ="alucIndex",
     inputs = {'age','amount','checking','coapp','depends','duration','employed','existcr',
               'foreign','history','housing','installp','job','marital','other','property',
               'resident','savings','telephon'},
     selectQuery = {method="uncertainty", includeAllData=true,
                    probVar ="P_good_badgood"},
     output = {casout= {name="queryout", replace=true}, copyvars="ALL"},
     target = "good_bad"  }

The following commands display the tables that are produced by this action call:

print(r.ModelInfo)
print(r.Nobs)

For more information about the results of this analysis, see the CASL version of this example.

Analyzing German Credit Data Using the uncertaintySampling Action

This section contains Python code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be accessible to the CAS server. One way to do this is to convert the dmagecr data to the comma-separated-value (CSV) file dmagecr.csv and then use the following code to load the CSV file into CAS:

s.upload_file('dmagecr.csv')

For more information about coding in Python, see Getting Started with SAS Viya for Python and SAS Viya: System Programming Guide.

The following DATA step creates an ID variable alucIndex and load the sampsio.dmagecr data set into your CAS session:

s.dataStep.runCode( single='no', code="\
data dmagecr;\
   set dmagecr;\
   alucIndex = _n_;\
run;  " )

The following code calls the forestTrain and trains random forest model using the German Credit Benchmark data set.

s.loadactionset("decisionTree")
out=s.forestTrain(
     table = {"name":"dmagecr"},
     seed=12345,
     casOut = {"name":"forest_model"},
     inputs = {"age", "amount", "checking", "coapp", "depends", "duration",
     "employed", "existcr", "foreign", "history", "housing", "installp", "job",
     "marital", "other", "property", "resident", "savings", "telephon"},
     target = "good_bad"
)

The following code scores the table dmagecr and predicts the propensity of the binary target GOOD_BAD

s.loadactionset("decisionTree")
  out=s.forestScore(
  table = {"name":"dmagecr"},
  casOut = {"name":"score_at_runtime"},
  copyVars = {"age", "amount", "checking", "coapp", "depends", "duration",
  "employed", "existcr", "foreign", "history", "housing", "installp", "job",
  "marital", "other", "property", "resident", "savings", "telephon","alucIndex",
  "good_bad"},
  encodeName=True,
  modelTable = {"name":"forest_model"}
)

The following code loads the uncertaintySampling action in the activeLearn action set to perform active learning based on the output from the random forest model:

s.loadactionset("activeLearn")
out=s.uncertaintySampling(
     table = {"name":"score_at_runtime"},
     topK  = 183,
     id = "alucIndex",
     inputs = {"age", "amount", "checking", "coapp", "depends", "duration",
     "employed", "existcr", "foreign", "history", "housing", "installp", "job",
     "marital", "other", "property", "resident", "savings", "telephon"},
     selectQuery = {"method":"uncertainty", "includeAllData":"True",
                    "probVar" : "P_good_badgood"},
     output = {"casout": {"name":"queryout", "replace":"True"}, "copyvars":"ALL"} ,
     target = "good_bad"
)

The following commands display the tables that are produced by this action call:

print(out.ModelInfo)
print(out.Nobs)

For more information about the results of this analysis, see the CASL version of this example.

Analyzing German Credit Data Using the uncertaintySampling Action

This example is not available for the R programming language.

Last updated: September 15, 2022