Bayesian Network Classifier Action Set: Syntax

Provides actions for performing classification using Bayesian network models

bnet Action

Bayesian Network Classifier Action.

CASL Syntax

bayesianNetClassifier.bnet <result=results> <status=rc> /
alpha={double-1 <, double-2, ...>}
attributes={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
bestModel=TRUE | FALSE
code={
casOut={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=TRUE | FALSE,
fmtWdth=integer,
iProb=TRUE | FALSE,
indentSize=integer,
intoCutPt=double,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
pCatAll=TRUE | FALSE,
tabForm=TRUE | FALSE
}
codeGroup="string"
diagnostics={
eyecatcher="string"
}
display={
caseSensitive=TRUE | FALSE,
exclude=TRUE | FALSE,
excludeAll=TRUE | FALSE,
keyIsPath=TRUE | FALSE,
names={"string-1" <, "string-2", ...>},
pathType="LABEL" | "NAME",
traceNames=TRUE | FALSE
}
freq="string"
id={"variable-name-1" <, "variable-name-2", ...>}
inputs={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
maxParents=integer
miAlpha=double
nominals={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
numBin=integer
outNetwork={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
output={
required parameter casOut={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | {"variable-name-1" <, "variable-name-2", ...>},
role="string"
}
outputTables={
groupByVarsRaw=TRUE | FALSE,
includeAll=TRUE | FALSE,
names={"string-1" <, "string-2", ...>} | {key-1={casouttable-1} <, key-2={casouttable-2}, ...>},
repeated=TRUE | FALSE,
replace=TRUE | FALSE
}
parenting={"BESTONE", "BESTSET"}
partByFrac={
seed=integer,
test=double,
validate=double
}
partByVar={
required parameter name="variable-name",
test="string",
train="string",
validate="string"
}
preScreening={"ONE", "ZERO"}
printtarget=TRUE | FALSE
resident=TRUE | FALSE
saveState={
caslib="string",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE
}
structures={"MB", "NAIVE", "PC", "TAN"}
required parameter table={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
}
target="string"
varSelect={"ONE", "THREE", "TWO", "ZERO"}
;

Parameter Descriptions

alpha={double-1 <, double-2, ...>}

specifies the significance level for independence tests by using chi-square or G-square statistics. If you want to choose the best model among several, you can specify up to five numbers, separated by spaces. If you specify multiple numbers but you do not specify the value True for the bestModel parameter, the action uses the first number and ignores the remaining numbers.

RequirementThe specified values must be unique.

attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

bestModel=TRUE | FALSE

when set to True, selects the best model.

DefaultFALSE

code={aircodegen}

For more information about specifying the code parameter, see the common aircodegen parameter (Appendix A: Common Parameters).

codeGroup="string"

Code Group

diagnostics={_diagnostics}

eyecatcher="string"

specifies a quoted string that will be prefixed to any messages that are associated with this action invocation.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

freq="string"

specifies the frequency variable.

id={"variable-name-1" <, "variable-name-2", ...>}

specifies the variables to copy to the generated table.

indepTest="ALL" | "CHIGSQUARE" | "CHISQUARE" | "GSQUARE" | "MI"

specifies the method for independence tests.

DefaultCHIGSQUARE
ALL

uses the chi-square statistic, the G-square statistic, and the normalized mutual information for independence tests. A variable is independent of the target if the p-values of both the chi-square and the G-square statistics are greater than the value of the alpha parameter and the normalized mutual information is less than the value of the miAlpha parameter.

CHIGSQUARE

uses both the chi-square and G-square statistics for independence tests. A variable is independent of the target if the p-values of both the chi-square and G-square statistics are greater than the value of the alpha parameter.

CHISQUARE

uses the chi-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

GSQUARE

uses the G-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

MI

uses the normalized mutual information for independence tests. A variable is independent of the target if the normalized mutual information is less than the value of the miAlpha parameter.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies variables to use for analysis.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

maxParents=integer

specifies the maximum number of parents allowed for each node in the network.

Default5
Range1–16

miAlpha=double

specifies the significance level for independence tests that use mutual information.

Default0.05
Range0–1

missingInt="IGNORE" | "IMPUTE"

specifies how to handle missing values for interval variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the interval variables.

IMPUTE

replaces the missing values in any interval variable by the mean of the variable.

missingNom="IGNORE" | "IMPUTE" | "LEVEL"

specifies how to handle missing values for nominal variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the nominal variables.

IMPUTE

replaces the missing values in any nominal variable by the mode of the variable.

LEVEL

treats the missing values in any nominal variable as a separate level of the variable.

nominals={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies nominal variables to use for analysis.

For more information about specifying the nominals parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

numBin=integer

specifies the binning number for interval variables.

Default5
Range2–1024

outNetwork={casouttable}

specifies the name of the output table for the network structure and the probability distributions.

For more information about specifying the outNetwork parameter, see the common casouttable parameter (Appendix A: Common Parameters).

output={BnetOutputStatement}

creates an output table to contain the predicted target values of the input table.

* casOut={casouttable}

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | {"variable-name-1" <, "variable-name-2", ...>}

specifies a list of one or more variables to be copied from the input table to the output table. You can alternatively specify the value ALL, ALL_MODEL, or ALL_NUMERIC, which respectively copies all variables, all variables used in the modeling, or all numeric variables from the input table to the output table.

role="string"

renames the generated column _ROLE_ in the output data table to the specified role name.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

parenting={"BESTONE", "BESTSET"}

specifies the structure learning methods. If you want the action to choose between the two methods, you can specify both BESTONE and BESTSET and also specify the value True for the bestModel parameter. If you specify both methods but you do not specify the value True for the bestModel parameter, the action uses the first specified method and ignores the other.

DefaultBESTSET
BESTONE

uses a greedy approach to determine the parents of each node; that is, for each node, the best candidate is added as a parent of the node in each iteration.

BESTSET

determines the best set of variables among possible candidate sets as the parents of each node; that is, instead of adding one variable in an iteration, the action tests multiple sets of variables together and chooses the best set as the parents of the node.

partByFrac={partByFracStatement}

The partByFracStatement value can be one or more of the following:
seed=integer

specifies the seed to use in the random number generator that is used for partitioning the data.

Default0
test=double

randomly assigns the specified proportion of observations in the input table to the testing role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Range0–1
validate=double

randomly assigns the specified proportion of observations in the input table to the validation role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Aliasvalid
Range0–1

partByVar={partByVarStatement}

Long formpartByVar={name="variable-name"}
Shortcut formpartByVar="variable-name"
The partByVarStatement value can be one or more of the following:
* name="variable-name"

names the variable in the input table whose values are used to assign rows to each observation.

test="string"

specifies the formatted value of the variable that is used to assign observations to the testing role.

train="string"

specifies the formatted value of the variable that is used to assign observations to the training role. If you do not specify the train parameter, then all observations whose roles are not determined by the test and validate parameters are assigned to training.

validate="string"

specifies the formatted value of the variable that is used to assign observations to the validation role.

Aliasvalid

preScreening={"ONE", "ZERO"}

specifies the initial screening for the input variables. If you want the action to choose the best model with or without prescreening, you can specify {"ZERO","ONE"} or {"ONE","ZERO"} for the parameter and also specify the value True for the bestModel parameter. If you specify both ONE and ZERO but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the other.

DefaultONE
RequirementThe specified values must be unique.
ONE

uses only the input variables that are dependent on the target.

ZERO

uses all the input variables.

printtarget=TRUE | FALSE

when set to True, generates names for the predicted target variable and the predicted probability variables.

DefaultFALSE

resident=TRUE | FALSE

DefaultTRUE

saveState={casouttable}

specifies the table in which to save the model for future scoring.

Long formsaveState={name="table-name"}
Shortcut formsaveState="table-name"
caslib="string"

specifies the name of the caslib to use.

name="table-name"

specifies the name to associate with the table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE

structures={"MB", "NAIVE", "PC", "TAN"}

specifies the network structure types. Together with the maxParents parameter, this parameter determines which network structure the action learns from the training data. If you want the action to choose the best structure among several structures, you can specify multiple values in any combination, separated by spaces, and also specify the value True for the bestModel parameter. If you specify multiple structures but you do not specify the value True for the bestModel parameter, the first value that you specify is used and the rest are ignored.

Aliasstructure
DefaultPC
RequirementThe specified values must be unique.
MB

learns the Markov blanket of the target variable. The Markov blanket includes the parents, the children, and the other parents of the children. After learning the Markov blanket, the action further determines the parents of the target, the links from the parents to the children, and the links among the children. When you specify the value MB for the structure parameter, the action learns the Markov blanket regardless of the values of the preScreening and the varSelect parameters.

NAIVE

learns a naive Bayesian network structure (that is, the target has a direct link to each input variable). If you specify the value 1 for maxParents, the structure being trained is a naive Bayesian network. If you specify a value greater than 1 for maxParents, the structure is a Bayesian network-augmented naive Bayesian network.

PC

learns the parent-child Bayesian network structure. PC structure differs from NAIVE structure in that some input variables could be learned as the parents of the target variable. In addition, links from the parents to the children and among the children are also possible in PC.

TAN

learns the tree-augmented naive Bayesian network structure. The TAN structure includes a direct link from the target to each input variable plus a tree structure among the input variables.

* table={castable}

specifies the settings for an input table.

Long formtable={name="table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

target="string"

specifies the target variable to use for analysis.

varSelect={"ONE", "THREE", "TWO", "ZERO"}

specifies how input variables are selected beyond prescreening. If you specify the value "ONE", "TWO", or "THREE", the action automatically tests each input variable for unconditional independence of the target regardless of the value of the preScreening parameter. If no variables are left at a particular variable selection level, the action rolls back to the previous level. For example, if you specify "THREE" and there are no variables in the Markov blanket of the target, the action uses the variables from the previous level "TWO". If you want to choose the best model among different levels of variable selection, you can specify any combination of values for this parameter and also specify the value True for the bestModel parameter. If you specify multiple values for the varSelect parameter but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the remaining values.

DefaultONE
RequirementThe specified values must be unique.
ONE

tests each input variable for conditional independence of the target variable given any other input variable. This type of selection uses only the variables that are conditionally dependent on the target given any other input variable.

THREE

determines the Markov blanket of the target variable and uses only the variables in the Markov blanket.

TWO

tests each input variable further for conditional independence of the target variable given any subset of other input variables. This type of selection uses only the variables that are conditionally dependent on the target given any subset of other input variables.

ZERO

uses all input variables that remain after the initial screening is performed as specified in the PRESCREENING= parameter.

bnet Action

Bayesian Network Classifier Action.

Lua Syntax

results, info = s:bayesianNetClassifier_bnet{
alpha={double-1 <, double-2, ...>},
attributes={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
bestModel=true | false,
code={
casOut={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=true | false,
fmtWdth=integer,
iProb=true | false,
indentSize=integer,
intoCutPt=double,
labelId=integer,
lineSize=integer,
noTrim=true | false,
pCatAll=true | false,
tabForm=true | false
},
codeGroup="string",
diagnostics={
eyecatcher="string"
},
display={
caseSensitive=true | false,
exclude=true | false,
excludeAll=true | false,
keyIsPath=true | false,
names={"string-1" <, "string-2", ...>},
pathType="LABEL" | "NAME",
traceNames=true | false
},
freq="string",
id={"variable-name-1" <, "variable-name-2", ...>},
inputs={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
maxParents=integer,
miAlpha=double,
nominals={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
numBin=integer,
outNetwork={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
output={
required parameter casOut={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | {"variable-name-1" <, "variable-name-2", ...>},
role="string"
},
outputTables={
groupByVarsRaw=true | false,
includeAll=true | false,
names={"string-1" <, "string-2", ...>} | {key-1={casouttable-1} <, key-2={casouttable-2}, ...>},
repeated=true | false,
replace=true | false
},
parenting={"BESTONE", "BESTSET"},
partByFrac={
seed=integer,
test=double,
validate=double
},
partByVar={
required parameter name="variable-name",
test="string",
train="string",
validate="string"
},
preScreening={"ONE", "ZERO"},
printtarget=true | false,
resident=true | false,
saveState={
caslib="string",
name="table-name",
promote=true | false,
replace=true | false
},
structures={"MB", "NAIVE", "PC", "TAN"},
required parameter table={
caslib="string",
computedOnDemand=true | false,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
},
target="string",
varSelect={"ONE", "THREE", "TWO", "ZERO"}
}

Parameter Descriptions

alpha={double-1 <, double-2, ...>}

specifies the significance level for independence tests by using chi-square or G-square statistics. If you want to choose the best model among several, you can specify up to five numbers, separated by spaces. If you specify multiple numbers but you do not specify the value True for the bestModel parameter, the action uses the first number and ignores the remaining numbers.

RequirementThe specified values must be unique.

attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

bestModel=true | false

when set to True, selects the best model.

Defaultfalse

code={aircodegen}

For more information about specifying the code parameter, see the common aircodegen parameter (Appendix A: Common Parameters).

codeGroup="string"

Code Group

diagnostics={_diagnostics}

eyecatcher="string"

specifies a quoted string that will be prefixed to any messages that are associated with this action invocation.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

freq="string"

specifies the frequency variable.

id={"variable-name-1" <, "variable-name-2", ...>}

specifies the variables to copy to the generated table.

indepTest="ALL" | "CHIGSQUARE" | "CHISQUARE" | "GSQUARE" | "MI"

specifies the method for independence tests.

DefaultCHIGSQUARE
ALL

uses the chi-square statistic, the G-square statistic, and the normalized mutual information for independence tests. A variable is independent of the target if the p-values of both the chi-square and the G-square statistics are greater than the value of the alpha parameter and the normalized mutual information is less than the value of the miAlpha parameter.

CHIGSQUARE

uses both the chi-square and G-square statistics for independence tests. A variable is independent of the target if the p-values of both the chi-square and G-square statistics are greater than the value of the alpha parameter.

CHISQUARE

uses the chi-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

GSQUARE

uses the G-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

MI

uses the normalized mutual information for independence tests. A variable is independent of the target if the normalized mutual information is less than the value of the miAlpha parameter.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies variables to use for analysis.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

maxParents=integer

specifies the maximum number of parents allowed for each node in the network.

Default5
Range1–16

miAlpha=double

specifies the significance level for independence tests that use mutual information.

Default0.05
Range0–1

missingInt="IGNORE" | "IMPUTE"

specifies how to handle missing values for interval variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the interval variables.

IMPUTE

replaces the missing values in any interval variable by the mean of the variable.

missingNom="IGNORE" | "IMPUTE" | "LEVEL"

specifies how to handle missing values for nominal variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the nominal variables.

IMPUTE

replaces the missing values in any nominal variable by the mode of the variable.

LEVEL

treats the missing values in any nominal variable as a separate level of the variable.

nominals={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies nominal variables to use for analysis.

For more information about specifying the nominals parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

numBin=integer

specifies the binning number for interval variables.

Default5
Range2–1024

outNetwork={casouttable}

specifies the name of the output table for the network structure and the probability distributions.

For more information about specifying the outNetwork parameter, see the common casouttable parameter (Appendix A: Common Parameters).

output={BnetOutputStatement}

creates an output table to contain the predicted target values of the input table.

* casOut={casouttable}

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | {"variable-name-1" <, "variable-name-2", ...>}

specifies a list of one or more variables to be copied from the input table to the output table. You can alternatively specify the value ALL, ALL_MODEL, or ALL_NUMERIC, which respectively copies all variables, all variables used in the modeling, or all numeric variables from the input table to the output table.

role="string"

renames the generated column _ROLE_ in the output data table to the specified role name.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

parenting={"BESTONE", "BESTSET"}

specifies the structure learning methods. If you want the action to choose between the two methods, you can specify both BESTONE and BESTSET and also specify the value True for the bestModel parameter. If you specify both methods but you do not specify the value True for the bestModel parameter, the action uses the first specified method and ignores the other.

DefaultBESTSET
BESTONE

uses a greedy approach to determine the parents of each node; that is, for each node, the best candidate is added as a parent of the node in each iteration.

BESTSET

determines the best set of variables among possible candidate sets as the parents of each node; that is, instead of adding one variable in an iteration, the action tests multiple sets of variables together and chooses the best set as the parents of the node.

partByFrac={partByFracStatement}

The partByFracStatement value can be one or more of the following:
seed=integer

specifies the seed to use in the random number generator that is used for partitioning the data.

Default0
test=double

randomly assigns the specified proportion of observations in the input table to the testing role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Range0–1
validate=double

randomly assigns the specified proportion of observations in the input table to the validation role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Aliasvalid
Range0–1

partByVar={partByVarStatement}

Long formpartByVar={name="variable-name"}
Shortcut formpartByVar="variable-name"
The partByVarStatement value can be one or more of the following:
* name="variable-name"

names the variable in the input table whose values are used to assign rows to each observation.

test="string"

specifies the formatted value of the variable that is used to assign observations to the testing role.

train="string"

specifies the formatted value of the variable that is used to assign observations to the training role. If you do not specify the train parameter, then all observations whose roles are not determined by the test and validate parameters are assigned to training.

validate="string"

specifies the formatted value of the variable that is used to assign observations to the validation role.

Aliasvalid

preScreening={"ONE", "ZERO"}

specifies the initial screening for the input variables. If you want the action to choose the best model with or without prescreening, you can specify {"ZERO","ONE"} or {"ONE","ZERO"} for the parameter and also specify the value True for the bestModel parameter. If you specify both ONE and ZERO but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the other.

DefaultONE
RequirementThe specified values must be unique.
ONE

uses only the input variables that are dependent on the target.

ZERO

uses all the input variables.

printtarget=true | false

when set to True, generates names for the predicted target variable and the predicted probability variables.

Defaultfalse

resident=true | false

Defaulttrue

saveState={casouttable}

specifies the table in which to save the model for future scoring.

Long formsaveState={name="table-name"}
Shortcut formsaveState="table-name"
caslib="string"

specifies the name of the caslib to use.

name="table-name"

specifies the name to associate with the table.

promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse

structures={"MB", "NAIVE", "PC", "TAN"}

specifies the network structure types. Together with the maxParents parameter, this parameter determines which network structure the action learns from the training data. If you want the action to choose the best structure among several structures, you can specify multiple values in any combination, separated by spaces, and also specify the value True for the bestModel parameter. If you specify multiple structures but you do not specify the value True for the bestModel parameter, the first value that you specify is used and the rest are ignored.

Aliasstructure
DefaultPC
RequirementThe specified values must be unique.
MB

learns the Markov blanket of the target variable. The Markov blanket includes the parents, the children, and the other parents of the children. After learning the Markov blanket, the action further determines the parents of the target, the links from the parents to the children, and the links among the children. When you specify the value MB for the structure parameter, the action learns the Markov blanket regardless of the values of the preScreening and the varSelect parameters.

NAIVE

learns a naive Bayesian network structure (that is, the target has a direct link to each input variable). If you specify the value 1 for maxParents, the structure being trained is a naive Bayesian network. If you specify a value greater than 1 for maxParents, the structure is a Bayesian network-augmented naive Bayesian network.

PC

learns the parent-child Bayesian network structure. PC structure differs from NAIVE structure in that some input variables could be learned as the parents of the target variable. In addition, links from the parents to the children and among the children are also possible in PC.

TAN

learns the tree-augmented naive Bayesian network structure. The TAN structure includes a direct link from the target to each input variable plus a tree structure among the input variables.

* table={castable}

specifies the settings for an input table.

Long formtable={name="table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

target="string"

specifies the target variable to use for analysis.

varSelect={"ONE", "THREE", "TWO", "ZERO"}

specifies how input variables are selected beyond prescreening. If you specify the value "ONE", "TWO", or "THREE", the action automatically tests each input variable for unconditional independence of the target regardless of the value of the preScreening parameter. If no variables are left at a particular variable selection level, the action rolls back to the previous level. For example, if you specify "THREE" and there are no variables in the Markov blanket of the target, the action uses the variables from the previous level "TWO". If you want to choose the best model among different levels of variable selection, you can specify any combination of values for this parameter and also specify the value True for the bestModel parameter. If you specify multiple values for the varSelect parameter but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the remaining values.

DefaultONE
RequirementThe specified values must be unique.
ONE

tests each input variable for conditional independence of the target variable given any other input variable. This type of selection uses only the variables that are conditionally dependent on the target given any other input variable.

THREE

determines the Markov blanket of the target variable and uses only the variables in the Markov blanket.

TWO

tests each input variable further for conditional independence of the target variable given any subset of other input variables. This type of selection uses only the variables that are conditionally dependent on the target given any subset of other input variables.

ZERO

uses all input variables that remain after the initial screening is performed as specified in the PRESCREENING= parameter.

bnet Action

Bayesian Network Classifier Action.

Python Syntax

results= s.bayesianNetClassifier.bnet(
alpha=[double-1 <, double-2, ...>],
attributes=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
bestModel=True | False,
code={
"casOut":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"comment":True | False,
"fmtWdth":integer,
"iProb":True | False,
"indentSize":integer,
"intoCutPt":double,
"labelId":integer,
"lineSize":integer,
"noTrim":True | False,
"pCatAll":True | False,
"tabForm":True | False
},
codeGroup="string",
diagnostics={
"eyecatcher":"string"
},
display={
"caseSensitive":True | False,
"exclude":True | False,
"excludeAll":True | False,
"keyIsPath":True | False,
"names":["string-1" <, "string-2", ...>],
"pathType":"LABEL" | "NAME",
"traceNames":True | False
},
freq="string",
id=["variable-name-1" <, "variable-name-2", ...>],
inputs=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
maxParents=integer,
miAlpha=double,
nominals=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
numBin=integer,
outNetwork={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
output={
required parameter "casOut":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"copyVars":"ALL" | "ALL_MODEL" | "ALL_NUMERIC" | ["variable-name-1" <, "variable-name-2", ...>],
"role":"string"
},
outputTables={
"groupByVarsRaw":True | False,
"includeAll":True | False,
"names":["string-1" <, "string-2", ...>] | {"key-1":{casouttable-1} <, "key-2":{casouttable-2}, ...>},
"repeated":True | False,
"replace":True | False
},
parenting=["BESTONE", "BESTSET"],
partByFrac={
"seed":integer,
"test":double,
"validate":double
},
partByVar={
required parameter "name":"variable-name",
"test":"string",
"train":"string",
"validate":"string"
},
preScreening=["ONE", "ZERO"],
printtarget=True | False,
resident=True | False,
saveState={
"caslib":"string",
"name":"table-name",
"promote":True | False,
"replace":True | False
},
structures=["MB", "NAIVE", "PC", "TAN"],
required parameter table={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"importOptions":{"fileType":"AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression"
},
target="string",
varSelect=["ONE", "THREE", "TWO", "ZERO"]
)

Parameter Descriptions

alpha=[double-1 <, double-2, ...>]

specifies the significance level for independence tests by using chi-square or G-square statistics. If you want to choose the best model among several, you can specify up to five numbers, separated by spaces. If you specify multiple numbers but you do not specify the value True for the bestModel parameter, the action uses the first number and ignores the remaining numbers.

RequirementThe specified values must be unique.

attributes=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

bestModel=True | False

when set to True, selects the best model.

DefaultFalse

code={aircodegen}

For more information about specifying the code parameter, see the common aircodegen parameter (Appendix A: Common Parameters).

codeGroup="string"

Code Group

diagnostics={_diagnostics}

"eyecatcher":"string"

specifies a quoted string that will be prefixed to any messages that are associated with this action invocation.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

freq="string"

specifies the frequency variable.

id=["variable-name-1" <, "variable-name-2", ...>]

specifies the variables to copy to the generated table.

indepTest="ALL" | "CHIGSQUARE" | "CHISQUARE" | "GSQUARE" | "MI"

specifies the method for independence tests.

DefaultCHIGSQUARE
ALL

uses the chi-square statistic, the G-square statistic, and the normalized mutual information for independence tests. A variable is independent of the target if the p-values of both the chi-square and the G-square statistics are greater than the value of the alpha parameter and the normalized mutual information is less than the value of the miAlpha parameter.

CHIGSQUARE

uses both the chi-square and G-square statistics for independence tests. A variable is independent of the target if the p-values of both the chi-square and G-square statistics are greater than the value of the alpha parameter.

CHISQUARE

uses the chi-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

GSQUARE

uses the G-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

MI

uses the normalized mutual information for independence tests. A variable is independent of the target if the normalized mutual information is less than the value of the miAlpha parameter.

inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies variables to use for analysis.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

maxParents=integer

specifies the maximum number of parents allowed for each node in the network.

Default5
Range1–16

miAlpha=double

specifies the significance level for independence tests that use mutual information.

Default0.05
Range0–1

missingInt="IGNORE" | "IMPUTE"

specifies how to handle missing values for interval variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the interval variables.

IMPUTE

replaces the missing values in any interval variable by the mean of the variable.

missingNom="IGNORE" | "IMPUTE" | "LEVEL"

specifies how to handle missing values for nominal variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the nominal variables.

IMPUTE

replaces the missing values in any nominal variable by the mode of the variable.

LEVEL

treats the missing values in any nominal variable as a separate level of the variable.

nominals=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies nominal variables to use for analysis.

For more information about specifying the nominals parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

numBin=integer

specifies the binning number for interval variables.

Default5
Range2–1024

outNetwork={casouttable}

specifies the name of the output table for the network structure and the probability distributions.

For more information about specifying the outNetwork parameter, see the common casouttable parameter (Appendix A: Common Parameters).

output={BnetOutputStatement}

creates an output table to contain the predicted target values of the input table.

* "casOut":{casouttable}

specifies the settings for an output table.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"copyVars":"ALL" | "ALL_MODEL" | "ALL_NUMERIC" | ["variable-name-1" <, "variable-name-2", ...>]

specifies a list of one or more variables to be copied from the input table to the output table. You can alternatively specify the value ALL, ALL_MODEL, or ALL_NUMERIC, which respectively copies all variables, all variables used in the modeling, or all numeric variables from the input table to the output table.

"role":"string"

renames the generated column _ROLE_ in the output data table to the specified role name.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

parenting=["BESTONE", "BESTSET"]

specifies the structure learning methods. If you want the action to choose between the two methods, you can specify both BESTONE and BESTSET and also specify the value True for the bestModel parameter. If you specify both methods but you do not specify the value True for the bestModel parameter, the action uses the first specified method and ignores the other.

DefaultBESTSET
BESTONE

uses a greedy approach to determine the parents of each node; that is, for each node, the best candidate is added as a parent of the node in each iteration.

BESTSET

determines the best set of variables among possible candidate sets as the parents of each node; that is, instead of adding one variable in an iteration, the action tests multiple sets of variables together and chooses the best set as the parents of the node.

partByFrac={partByFracStatement}

The partByFracStatement value can be one or more of the following:
"seed":integer

specifies the seed to use in the random number generator that is used for partitioning the data.

Default0
"test":double

randomly assigns the specified proportion of observations in the input table to the testing role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Range0–1
"validate":double

randomly assigns the specified proportion of observations in the input table to the validation role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Aliasvalid
Range0–1

partByVar={partByVarStatement}

Long formpartByVar={"name":"variable-name"}
Shortcut formpartByVar="variable-name"
The partByVarStatement value can be one or more of the following:
* "name":"variable-name"

names the variable in the input table whose values are used to assign rows to each observation.

"test":"string"

specifies the formatted value of the variable that is used to assign observations to the testing role.

"train":"string"

specifies the formatted value of the variable that is used to assign observations to the training role. If you do not specify the train parameter, then all observations whose roles are not determined by the test and validate parameters are assigned to training.

"validate":"string"

specifies the formatted value of the variable that is used to assign observations to the validation role.

Aliasvalid

preScreening=["ONE", "ZERO"]

specifies the initial screening for the input variables. If you want the action to choose the best model with or without prescreening, you can specify {"ZERO","ONE"} or {"ONE","ZERO"} for the parameter and also specify the value True for the bestModel parameter. If you specify both ONE and ZERO but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the other.

DefaultONE
RequirementThe specified values must be unique.
ONE

uses only the input variables that are dependent on the target.

ZERO

uses all the input variables.

printtarget=True | False

when set to True, generates names for the predicted target variable and the predicted probability variables.

DefaultFalse

resident=True | False

DefaultTrue

saveState={casouttable}

specifies the table in which to save the model for future scoring.

Long formsaveState={"name":"table-name"}
Shortcut formsaveState="table-name"
"caslib":"string"

specifies the name of the caslib to use.

"name":"table-name"

specifies the name to associate with the table.

"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse

structures=["MB", "NAIVE", "PC", "TAN"]

specifies the network structure types. Together with the maxParents parameter, this parameter determines which network structure the action learns from the training data. If you want the action to choose the best structure among several structures, you can specify multiple values in any combination, separated by spaces, and also specify the value True for the bestModel parameter. If you specify multiple structures but you do not specify the value True for the bestModel parameter, the first value that you specify is used and the rest are ignored.

Aliasstructure
DefaultPC
RequirementThe specified values must be unique.
MB

learns the Markov blanket of the target variable. The Markov blanket includes the parents, the children, and the other parents of the children. After learning the Markov blanket, the action further determines the parents of the target, the links from the parents to the children, and the links among the children. When you specify the value MB for the structure parameter, the action learns the Markov blanket regardless of the values of the preScreening and the varSelect parameters.

NAIVE

learns a naive Bayesian network structure (that is, the target has a direct link to each input variable). If you specify the value 1 for maxParents, the structure being trained is a naive Bayesian network. If you specify a value greater than 1 for maxParents, the structure is a Bayesian network-augmented naive Bayesian network.

PC

learns the parent-child Bayesian network structure. PC structure differs from NAIVE structure in that some input variables could be learned as the parents of the target variable. In addition, links from the parents to the children and among the children are also possible in PC.

TAN

learns the tree-augmented naive Bayesian network structure. The TAN structure includes a direct link from the target to each input variable plus a tree structure among the input variables.

* table={castable}

specifies the settings for an input table.

Long formtable={"name":"table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
"caslib":"string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasoptions, dataSource
"groupBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the variables to use for grouping results.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"groupByMode":"NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the table to use.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"where":"where-expression"

specifies an expression for subsetting the input data.

target="string"

specifies the target variable to use for analysis.

varSelect=["ONE", "THREE", "TWO", "ZERO"]

specifies how input variables are selected beyond prescreening. If you specify the value "ONE", "TWO", or "THREE", the action automatically tests each input variable for unconditional independence of the target regardless of the value of the preScreening parameter. If no variables are left at a particular variable selection level, the action rolls back to the previous level. For example, if you specify "THREE" and there are no variables in the Markov blanket of the target, the action uses the variables from the previous level "TWO". If you want to choose the best model among different levels of variable selection, you can specify any combination of values for this parameter and also specify the value True for the bestModel parameter. If you specify multiple values for the varSelect parameter but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the remaining values.

DefaultONE
RequirementThe specified values must be unique.
ONE

tests each input variable for conditional independence of the target variable given any other input variable. This type of selection uses only the variables that are conditionally dependent on the target given any other input variable.

THREE

determines the Markov blanket of the target variable and uses only the variables in the Markov blanket.

TWO

tests each input variable further for conditional independence of the target variable given any subset of other input variables. This type of selection uses only the variables that are conditionally dependent on the target given any subset of other input variables.

ZERO

uses all input variables that remain after the initial screening is performed as specified in the PRESCREENING= parameter.

bnet Action

Bayesian Network Classifier Action.

R Syntax

results <– cas.bayesianNetClassifier.bnet(s,
alpha=list(double-1 <, double-2, ...>),
attributes=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
bestModel=TRUE | FALSE,
code=list(
casOut=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
comment=TRUE | FALSE,
fmtWdth=integer,
iProb=TRUE | FALSE,
indentSize=integer,
intoCutPt=double,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
pCatAll=TRUE | FALSE,
tabForm=TRUE | FALSE
),
codeGroup="string",
diagnostics=list(
eyecatcher="string"
),
display=list(
caseSensitive=TRUE | FALSE,
exclude=TRUE | FALSE,
excludeAll=TRUE | FALSE,
keyIsPath=TRUE | FALSE,
names=list("string-1" <, "string-2", ...>),
pathType="LABEL" | "NAME",
traceNames=TRUE | FALSE
),
freq="string",
id=list("variable-name-1" <, "variable-name-2", ...>),
inputs=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
maxParents=integer,
miAlpha=double,
nominals=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
numBin=integer,
outNetwork=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
output=list(
required parameter casOut=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | list("variable-name-1" <, "variable-name-2", ...>),
role="string"
),
outputTables=list(
groupByVarsRaw=TRUE | FALSE,
includeAll=TRUE | FALSE,
names=list("string-1" <, "string-2", ...>) | list(key-1=list(casouttable-1) <, key-2=list(casouttable-2), ...>),
repeated=TRUE | FALSE,
replace=TRUE | FALSE
),
parenting=list("BESTONE", "BESTSET"),
partByFrac=list(
seed=integer,
test=double,
validate=double
),
partByVar=list(
required parameter name="variable-name",
test="string",
train="string",
validate="string"
),
preScreening=list("ONE", "ZERO"),
printtarget=TRUE | FALSE,
resident=TRUE | FALSE,
saveState=list(
caslib="string",
name="table-name",
promote=TRUE | FALSE,
replace=TRUE | FALSE
),
structures=list("MB", "NAIVE", "PC", "TAN"),
required parameter table=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression"
),
target="string",
varSelect=list("ONE", "THREE", "TWO", "ZERO")
)

Parameter Descriptions

alpha=list(double-1 <, double-2, ...>)

specifies the significance level for independence tests by using chi-square or G-square statistics. If you want to choose the best model among several, you can specify up to five numbers, separated by spaces. If you specify multiple numbers but you do not specify the value True for the bestModel parameter, the action uses the first number and ignores the remaining numbers.

RequirementThe specified values must be unique.

attributes=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

bestModel=TRUE | FALSE

when set to True, selects the best model.

DefaultFALSE

code=list(aircodegen)

For more information about specifying the code parameter, see the common aircodegen parameter (Appendix A: Common Parameters).

codeGroup="string"

Code Group

diagnostics=list(_diagnostics)

eyecatcher="string"

specifies a quoted string that will be prefixed to any messages that are associated with this action invocation.

display=list(displayTables)

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

freq="string"

specifies the frequency variable.

id=list("variable-name-1" <, "variable-name-2", ...>)

specifies the variables to copy to the generated table.

indepTest="ALL" | "CHIGSQUARE" | "CHISQUARE" | "GSQUARE" | "MI"

specifies the method for independence tests.

DefaultCHIGSQUARE
ALL

uses the chi-square statistic, the G-square statistic, and the normalized mutual information for independence tests. A variable is independent of the target if the p-values of both the chi-square and the G-square statistics are greater than the value of the alpha parameter and the normalized mutual information is less than the value of the miAlpha parameter.

CHIGSQUARE

uses both the chi-square and G-square statistics for independence tests. A variable is independent of the target if the p-values of both the chi-square and G-square statistics are greater than the value of the alpha parameter.

CHISQUARE

uses the chi-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

GSQUARE

uses the G-square statistic for independence tests. A variable is independent of the target if the p-value of the statistic is greater than the value of the alpha parameter.

MI

uses the normalized mutual information for independence tests. A variable is independent of the target if the normalized mutual information is less than the value of the miAlpha parameter.

inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies variables to use for analysis.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

maxParents=integer

specifies the maximum number of parents allowed for each node in the network.

Default5
Range1–16

miAlpha=double

specifies the significance level for independence tests that use mutual information.

Default0.05
Range0–1

missingInt="IGNORE" | "IMPUTE"

specifies how to handle missing values for interval variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the interval variables.

IMPUTE

replaces the missing values in any interval variable by the mean of the variable.

missingNom="IGNORE" | "IMPUTE" | "LEVEL"

specifies how to handle missing values for nominal variables.

DefaultIGNORE
IGNORE

ignores the observations that have missing values in any of the nominal variables.

IMPUTE

replaces the missing values in any nominal variable by the mode of the variable.

LEVEL

treats the missing values in any nominal variable as a separate level of the variable.

nominals=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies nominal variables to use for analysis.

For more information about specifying the nominals parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

numBin=integer

specifies the binning number for interval variables.

Default5
Range2–1024

outNetwork=list(casouttable)

specifies the name of the output table for the network structure and the probability distributions.

For more information about specifying the outNetwork parameter, see the common casouttable parameter (Appendix A: Common Parameters).

output=list(BnetOutputStatement)

creates an output table to contain the predicted target values of the input table.

* casOut=list(casouttable)

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

copyVars="ALL" | "ALL_MODEL" | "ALL_NUMERIC" | list("variable-name-1" <, "variable-name-2", ...>)

specifies a list of one or more variables to be copied from the input table to the output table. You can alternatively specify the value ALL, ALL_MODEL, or ALL_NUMERIC, which respectively copies all variables, all variables used in the modeling, or all numeric variables from the input table to the output table.

role="string"

renames the generated column _ROLE_ in the output data table to the specified role name.

outputTables=list(outputTables)

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

parenting=list("BESTONE", "BESTSET")

specifies the structure learning methods. If you want the action to choose between the two methods, you can specify both BESTONE and BESTSET and also specify the value True for the bestModel parameter. If you specify both methods but you do not specify the value True for the bestModel parameter, the action uses the first specified method and ignores the other.

DefaultBESTSET
BESTONE

uses a greedy approach to determine the parents of each node; that is, for each node, the best candidate is added as a parent of the node in each iteration.

BESTSET

determines the best set of variables among possible candidate sets as the parents of each node; that is, instead of adding one variable in an iteration, the action tests multiple sets of variables together and chooses the best set as the parents of the node.

partByFrac=list(partByFracStatement)

The partByFracStatement value can be one or more of the following:
seed=integer

specifies the seed to use in the random number generator that is used for partitioning the data.

Default0
test=double

randomly assigns the specified proportion of observations in the input table to the testing role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Range0–1
validate=double

randomly assigns the specified proportion of observations in the input table to the validation role. The sum of the fractions that are specified in the test and validate parameters must be less than 1.

Aliasvalid
Range0–1

partByVar=list(partByVarStatement)

Long formpartByVar=list(name="variable-name")
Shortcut formpartByVar="variable-name"
The partByVarStatement value can be one or more of the following:
* name="variable-name"

names the variable in the input table whose values are used to assign rows to each observation.

test="string"

specifies the formatted value of the variable that is used to assign observations to the testing role.

train="string"

specifies the formatted value of the variable that is used to assign observations to the training role. If you do not specify the train parameter, then all observations whose roles are not determined by the test and validate parameters are assigned to training.

validate="string"

specifies the formatted value of the variable that is used to assign observations to the validation role.

Aliasvalid

preScreening=list("ONE", "ZERO")

specifies the initial screening for the input variables. If you want the action to choose the best model with or without prescreening, you can specify {"ZERO","ONE"} or {"ONE","ZERO"} for the parameter and also specify the value True for the bestModel parameter. If you specify both ONE and ZERO but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the other.

DefaultONE
RequirementThe specified values must be unique.
ONE

uses only the input variables that are dependent on the target.

ZERO

uses all the input variables.

printtarget=TRUE | FALSE

when set to True, generates names for the predicted target variable and the predicted probability variables.

DefaultFALSE

resident=TRUE | FALSE

DefaultTRUE

saveState=list(casouttable)

specifies the table in which to save the model for future scoring.

Long formsaveState=list(name="table-name")
Shortcut formsaveState="table-name"
caslib="string"

specifies the name of the caslib to use.

name="table-name"

specifies the name to associate with the table.

promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE

structures=list("MB", "NAIVE", "PC", "TAN")

specifies the network structure types. Together with the maxParents parameter, this parameter determines which network structure the action learns from the training data. If you want the action to choose the best structure among several structures, you can specify multiple values in any combination, separated by spaces, and also specify the value True for the bestModel parameter. If you specify multiple structures but you do not specify the value True for the bestModel parameter, the first value that you specify is used and the rest are ignored.

Aliasstructure
DefaultPC
RequirementThe specified values must be unique.
MB

learns the Markov blanket of the target variable. The Markov blanket includes the parents, the children, and the other parents of the children. After learning the Markov blanket, the action further determines the parents of the target, the links from the parents to the children, and the links among the children. When you specify the value MB for the structure parameter, the action learns the Markov blanket regardless of the values of the preScreening and the varSelect parameters.

NAIVE

learns a naive Bayesian network structure (that is, the target has a direct link to each input variable). If you specify the value 1 for maxParents, the structure being trained is a naive Bayesian network. If you specify a value greater than 1 for maxParents, the structure is a Bayesian network-augmented naive Bayesian network.

PC

learns the parent-child Bayesian network structure. PC structure differs from NAIVE structure in that some input variables could be learned as the parents of the target variable. In addition, links from the parents to the children and among the children are also possible in PC.

TAN

learns the tree-augmented naive Bayesian network structure. The TAN structure includes a direct link from the target to each input variable plus a tree structure among the input variables.

* table=list(castable)

specifies the settings for an input table.

Long formtable=list(name="table-name")
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasoptions, dataSource
groupBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

target="string"

specifies the target variable to use for analysis.

varSelect=list("ONE", "THREE", "TWO", "ZERO")

specifies how input variables are selected beyond prescreening. If you specify the value "ONE", "TWO", or "THREE", the action automatically tests each input variable for unconditional independence of the target regardless of the value of the preScreening parameter. If no variables are left at a particular variable selection level, the action rolls back to the previous level. For example, if you specify "THREE" and there are no variables in the Markov blanket of the target, the action uses the variables from the previous level "TWO". If you want to choose the best model among different levels of variable selection, you can specify any combination of values for this parameter and also specify the value True for the bestModel parameter. If you specify multiple values for the varSelect parameter but you do not specify the value True for the bestModel parameter, the action uses the first specified value and ignores the remaining values.

DefaultONE
RequirementThe specified values must be unique.
ONE

tests each input variable for conditional independence of the target variable given any other input variable. This type of selection uses only the variables that are conditionally dependent on the target given any other input variable.

THREE

determines the Markov blanket of the target variable and uses only the variables in the Markov blanket.

TWO

tests each input variable further for conditional independence of the target variable given any subset of other input variables. This type of selection uses only the variables that are conditionally dependent on the target given any subset of other input variables.

ZERO

uses all input variables that remain after the initial screening is performed as specified in the PRESCREENING= parameter.

Last updated: June 07, 2018