Decision Tree Action Set: Syntax
Provides actions for modeling and scoring with decision trees, forests, and gradient boosting
dtreeSplit Action
Splits decision tree nodes.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametermodelTable |
— |
specifies the table containing the model. |
|
required parametertable |
— |
specifies the settings for an input table. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the table to store the decision tree model in. When not specified, a random name is generated. | |
|
casOut |
requests that the action produce SAS score code. Specify additional parameters. | |
|
— |
specifies the table to store the generated aStore model. |
Parameter Descriptions
alpha=double
specifies the value to use for minimal cost-complexity pruning for regression trees.
| Minimum value | 0 |
|---|
applyRowOrder=TRUE | FALSE
Specifies that you wish the action use a prespecified row ordering. This requires using the orderby and groupby parameters on a preliminary table.partition action call.
| Alias | reproducibleRowOrder |
|---|---|
| Default | FALSE |
attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies temporary attributes, such as a format, to apply to input variables.
For more information about specifying the attributes parameter, see the common casinvardesc parameter.
| Aliases | attribute |
|---|---|
| attrs | |
| attr | |
| varAttrs |
binOrder=TRUE | FALSE
by default, the bin order is preserved for numeric variables. When set to False, the bin order is ignored for numeric variables.
| Default | TRUE |
|---|
bonferroni=TRUE | FALSE
when set to True, specifies to perform a Bonferroni correction when the split criterion uses a chi-square statistic or CHAID.
| Default | FALSE |
|---|
casOut={casouttable}
specifies the table to store the decision tree model in. When not specified, a random name is generated.
For more information about specifying the casOut parameter, see the common casouttable parameter.
cfLev=double
specifies the aggressiveness of tree pruning according to the C4.5 algorithm.
| Default | 0.25 |
|---|
code={codegen}
requests that the action produce SAS score code. Specify additional parameters.
For more information about specifying the code parameter, see the common codegen parameter.
crit="CHAID" | "CHISQUARE" | "FTEST" | "GAIN" | "GAINRATIO" | "GINI" | "VARIANCE"
specifies the split criterion for each tree node.
encodeName=TRUE | FALSE
specifies whether to encode the variable names such as predicted probabilities of a binary or nominal target in the generated casout table. The predicted probabilities are named with the prefix P_ instead of _DT_P_.
| Default | FALSE |
|---|
freq="variable-name"
specifies a numeric variable that contains the frequency of occurrence of each observation.
greedy=TRUE | FALSE
by default, a greedy search or exhaustive search is used to determine the best split for each variable of each tree node. When set to False, a fast and efficient algorithm that is based on clustering is applied. Setting this parameter to False is recommended for variables with high cardinality.
| Default | TRUE |
|---|
includeMissing=TRUE | FALSE
by default, observations with missing values are included. When set to False, observations with missing values for the analysis variables are excluded.
| Default | TRUE |
|---|
* inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the input variables to use in the analysis.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | input |
|---|
leafSize=integer
specifies the minimum number of observations on each node.
| Default | 5 |
|---|---|
| Minimum value | 1 |
maxBranch=integer
specifies the maximum number of children (branches) allowed for each level of the tree.
| Default | 2 |
|---|---|
| Minimum value | 1 |
maxLevel=integer
specifies the maximum number of the tree level.
| Default | 6 |
|---|---|
| Minimum value | 1 |
mergeBin=TRUE | FALSE
by default, when the largest value in one bin matches the lowest value in a neighboring bin, the values are merged into the lower bin. When set to False, the action does not try to merge bins.
| Default | TRUE |
|---|
minGain=double
specifies the minimum value to use to validate a splitting point when the criteria is not chi-square or CHAID.
| Minimum value | 0 |
|---|
minUseInSearch=integer
specifies a threshold for utilizing missing values in the split search when the missing parameter is set to USEINSEARCH. If the number of observations in which the splitting variable has missing values in a node is greater than or equal to the specified value, then the action initiates the USEINSEARCH policy. Otherwise, the missing values are assigned to a popular branch.
| Default | 1 |
|---|
missing="BRANCH" | "MACSMALL" | "POPULAR" | "SIMILAR" | "USEINSEARCH"
specifies the missing policy to handle missing values.
| Default | USEINSEARCH |
|---|
MACSMALL
specifies to treat the missing values for numeric variables as the smallest machine value and to treat missing values for nominal variables as a separate level.
modelId="string"
specifies the model ID variable name to use when generating SAS score code. By default, _DT_ is prefixed to the target variable name and _ is appended.
* modelTable={castable}
specifies the table containing the model.
| Long form | modelTable={name="table-name"} |
|---|---|
| Shortcut form | modelTable="table-name" |
| Alias | model |
|---|
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | FALSE |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
|---|
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable={groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
nBins=integer
specifies the number of bins to use for numeric variables in the calculation of the decision tree.
| Default | 50 |
|---|---|
| Minimum value | 1 |
nodeId={integer-1 <, integer-2, ...>}
specifies the leaf node IDs to split. When not specified or the list is not valid then all leaf nodes in the tree model are split.
nominals={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the nominal input variables to use in the analysis.
For more information about specifying the nominals parameter, see the common casinvardesc parameter.
| Alias | nominal |
|---|
nominalSearch={tkcasdt_nomSearchOpts}
specifies the method for finding a split on a nominal input.
| Alias | nomSearch |
|---|
The tkcasdt_nomSearchOpts value can be one or more of the following:
handling="CLASSIC" | "ENHANCED"
maxCategories=64-bit-integer
specifies the maximum number of levels for a splitting rule to include.
| Aliases | maxCats |
|---|---|
| maxLevels | |
| maxValues | |
| cluster | |
| minCardCluster | |
| Default | 128 |
| Minimum value | 0 |
shrinkage=double
specifies how much weight to give the category average in the sort method.
| Default | 10 |
|---|---|
| Minimum value | 0 |
sort=64-bit-integer
specifies the minimum cardinality of an input to use the sort method.
| Alias | minCardSort |
|---|---|
| Default | 10 |
| Minimum value | 0 |
sortBy="COUNT" | "TARGET"
noSplit=TRUE | FALSE
runs the action without splitting nodes or growing the tree. If you specify this parameter, then new columns appear in the output model table that you specify in the casOut parameter. These columns contain statistics about the data specified in the table parameter with respect to the tree model that you specify in the modelTable parameter. You should use this parameter if you want statistics about validation data.
| Default | FALSE |
|---|
prune=TRUE | FALSE
specify true to use a C4.5 pruning method for classification trees or minimal cost-complexity pruning for regression trees.
| Default | FALSE |
|---|
pVal=double
specifies the maximum P-value to use in chi-square and CHAID split criteria to validate a splitting point.
| Default | 1 |
|---|---|
| Range | 0–1 |
quantileBin=TRUE | FALSE
specifies bin boundaries at quantiles of numerical inputs instead of bins of equal width.
| Aliases | qbin |
|---|---|
| qtbin | |
| Default | TRUE |
saveState={casouttable}
specifies the table to store the generated aStore model.
For more information about specifying the saveState parameter, see the common casouttable parameter.
splitOnce=TRUE | FALSE
when set to True, the analysis variable only appears once for each path from the tree root node to the leaf node.
| Default | FALSE |
|---|
stat=TRUE | FALSE
specifies whether to include the variable splitting information for each node in the decision tree.
| Default | FALSE |
|---|
* table={castable}
specifies the settings for an input table.
| Long form | table={name="table-name"} |
|---|---|
| Shortcut form | table="table-name" |
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | FALSE |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
|---|
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable={groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
* target="variable-name"
specifies the target or response variable for training. If the variable is numeric, but not specified in the nominal= parameter and nbinstarget= is not specified, then a regression tree is trained.
updateBin=TRUE | FALSE
when set to True, the bin information for all the numerical analysis variables is computed on each node that is used for a split. All the splits make use of the updated bin information. Note that this changes the bin information significantly compared to the bin information that is used in the input model. As a result, this might bias the splits that occur later.
| Default | FALSE |
|---|
userDefinedSplit={userDefinedSplit}
indicates that you will manually define the split of a single node as specified in the nodeId parameter. The userDefinedSplit parameter contains sub-parameters needed for defining this split.
| Long form | userDefinedSplit={splitVar="variable-name"} |
|---|---|
| Shortcut form | userDefinedSplit="variable-name" |
The userDefinedSplit value can be one or more of the following:
intervalSplit={double-1 <, double-2, ...>}
is a list of interval split points. The number of branches is the number of numbers listed plus one. The branches are numbered with the 0 branch being all values less than the least value. Each other branch is numbered in increasing order.
missingBranch=64-bit-integer
indicates to which branch missing values should be assigned. By default, missing values are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
nominalSplit={any-list-or-data-type-1 <, any-list-or-data-type-2, ...>}
is a list of string lists that contain nominal split values. The number of branches is the number of string lists. The branches are ordered in which they are given to the action, with the first branch being numbered as branch 0.
splitVar="variable-name"
indicates which variable the decision tree is to use for splitting.
unseenBranch=64-bit-integer
indicates to which branch unseen levels should be assigned. By default, unseen levels are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
varImp=TRUE | FALSE
specifies whether the variable importance information is generated. The importance value is determined by the total Gini reduction.
| Default | FALSE |
|---|
dtreeSplit Action
Splits decision tree nodes.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametermodelTable |
— |
specifies the table containing the model. |
|
required parametertable |
— |
specifies the settings for an input table. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the table to store the decision tree model in. When not specified, a random name is generated. | |
|
casOut |
requests that the action produce SAS score code. Specify additional parameters. | |
|
— |
specifies the table to store the generated aStore model. |
Parameter Descriptions
alpha=double
specifies the value to use for minimal cost-complexity pruning for regression trees.
| Minimum value | 0 |
|---|
applyRowOrder=true | false
Specifies that you wish the action use a prespecified row ordering. This requires using the orderby and groupby parameters on a preliminary table.partition action call.
| Alias | reproducibleRowOrder |
|---|---|
| Default | false |
attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies temporary attributes, such as a format, to apply to input variables.
For more information about specifying the attributes parameter, see the common casinvardesc parameter.
| Aliases | attribute |
|---|---|
| attrs | |
| attr | |
| varAttrs |
binOrder=true | false
by default, the bin order is preserved for numeric variables. When set to False, the bin order is ignored for numeric variables.
| Default | true |
|---|
bonferroni=true | false
when set to True, specifies to perform a Bonferroni correction when the split criterion uses a chi-square statistic or CHAID.
| Default | false |
|---|
casOut={casouttable}
specifies the table to store the decision tree model in. When not specified, a random name is generated.
For more information about specifying the casOut parameter, see the common casouttable parameter.
cfLev=double
specifies the aggressiveness of tree pruning according to the C4.5 algorithm.
| Default | 0.25 |
|---|
code={codegen}
requests that the action produce SAS score code. Specify additional parameters.
For more information about specifying the code parameter, see the common codegen parameter.
crit="CHAID" | "CHISQUARE" | "FTEST" | "GAIN" | "GAINRATIO" | "GINI" | "VARIANCE"
specifies the split criterion for each tree node.
encodeName=true | false
specifies whether to encode the variable names such as predicted probabilities of a binary or nominal target in the generated casout table. The predicted probabilities are named with the prefix P_ instead of _DT_P_.
| Default | false |
|---|
freq="variable-name"
specifies a numeric variable that contains the frequency of occurrence of each observation.
greedy=true | false
by default, a greedy search or exhaustive search is used to determine the best split for each variable of each tree node. When set to False, a fast and efficient algorithm that is based on clustering is applied. Setting this parameter to False is recommended for variables with high cardinality.
| Default | true |
|---|
includeMissing=true | false
by default, observations with missing values are included. When set to False, observations with missing values for the analysis variables are excluded.
| Default | true |
|---|
* inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the input variables to use in the analysis.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | input |
|---|
leafSize=integer
specifies the minimum number of observations on each node.
| Default | 5 |
|---|---|
| Minimum value | 1 |
maxBranch=integer
specifies the maximum number of children (branches) allowed for each level of the tree.
| Default | 2 |
|---|---|
| Minimum value | 1 |
maxLevel=integer
specifies the maximum number of the tree level.
| Default | 6 |
|---|---|
| Minimum value | 1 |
mergeBin=true | false
by default, when the largest value in one bin matches the lowest value in a neighboring bin, the values are merged into the lower bin. When set to False, the action does not try to merge bins.
| Default | true |
|---|
minGain=double
specifies the minimum value to use to validate a splitting point when the criteria is not chi-square or CHAID.
| Minimum value | 0 |
|---|
minUseInSearch=integer
specifies a threshold for utilizing missing values in the split search when the missing parameter is set to USEINSEARCH. If the number of observations in which the splitting variable has missing values in a node is greater than or equal to the specified value, then the action initiates the USEINSEARCH policy. Otherwise, the missing values are assigned to a popular branch.
| Default | 1 |
|---|
missing="BRANCH" | "MACSMALL" | "POPULAR" | "SIMILAR" | "USEINSEARCH"
specifies the missing policy to handle missing values.
| Default | USEINSEARCH |
|---|
MACSMALL
specifies to treat the missing values for numeric variables as the smallest machine value and to treat missing values for nominal variables as a separate level.
modelId="string"
specifies the model ID variable name to use when generating SAS score code. By default, _DT_ is prefixed to the target variable name and _ is appended.
* modelTable={castable}
specifies the table containing the model.
| Long form | modelTable={name="table-name"} |
|---|---|
| Shortcut form | modelTable="table-name" |
| Alias | model |
|---|
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=true | false
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | false |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=true | false
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | false |
|---|
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable={groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
nBins=integer
specifies the number of bins to use for numeric variables in the calculation of the decision tree.
| Default | 50 |
|---|---|
| Minimum value | 1 |
nodeId={integer-1 <, integer-2, ...>}
specifies the leaf node IDs to split. When not specified or the list is not valid then all leaf nodes in the tree model are split.
nominals={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the nominal input variables to use in the analysis.
For more information about specifying the nominals parameter, see the common casinvardesc parameter.
| Alias | nominal |
|---|
nominalSearch={tkcasdt_nomSearchOpts}
specifies the method for finding a split on a nominal input.
| Alias | nomSearch |
|---|
The tkcasdt_nomSearchOpts value can be one or more of the following:
handling="CLASSIC" | "ENHANCED"
maxCategories=64-bit-integer
specifies the maximum number of levels for a splitting rule to include.
| Aliases | maxCats |
|---|---|
| maxLevels | |
| maxValues | |
| cluster | |
| minCardCluster | |
| Default | 128 |
| Minimum value | 0 |
shrinkage=double
specifies how much weight to give the category average in the sort method.
| Default | 10 |
|---|---|
| Minimum value | 0 |
sort=64-bit-integer
specifies the minimum cardinality of an input to use the sort method.
| Alias | minCardSort |
|---|---|
| Default | 10 |
| Minimum value | 0 |
sortBy="COUNT" | "TARGET"
noSplit=true | false
runs the action without splitting nodes or growing the tree. If you specify this parameter, then new columns appear in the output model table that you specify in the casOut parameter. These columns contain statistics about the data specified in the table parameter with respect to the tree model that you specify in the modelTable parameter. You should use this parameter if you want statistics about validation data.
| Default | false |
|---|
prune=true | false
specify true to use a C4.5 pruning method for classification trees or minimal cost-complexity pruning for regression trees.
| Default | false |
|---|
pVal=double
specifies the maximum P-value to use in chi-square and CHAID split criteria to validate a splitting point.
| Default | 1 |
|---|---|
| Range | 0–1 |
quantileBin=true | false
specifies bin boundaries at quantiles of numerical inputs instead of bins of equal width.
| Aliases | qbin |
|---|---|
| qtbin | |
| Default | true |
saveState={casouttable}
specifies the table to store the generated aStore model.
For more information about specifying the saveState parameter, see the common casouttable parameter.
splitOnce=true | false
when set to True, the analysis variable only appears once for each path from the tree root node to the leaf node.
| Default | false |
|---|
stat=true | false
specifies whether to include the variable splitting information for each node in the decision tree.
| Default | false |
|---|
* table={castable}
specifies the settings for an input table.
| Long form | table={name="table-name"} |
|---|---|
| Shortcut form | table="table-name" |
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=true | false
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | false |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=true | false
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | false |
|---|
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable={groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
* target="variable-name"
specifies the target or response variable for training. If the variable is numeric, but not specified in the nominal= parameter and nbinstarget= is not specified, then a regression tree is trained.
updateBin=true | false
when set to True, the bin information for all the numerical analysis variables is computed on each node that is used for a split. All the splits make use of the updated bin information. Note that this changes the bin information significantly compared to the bin information that is used in the input model. As a result, this might bias the splits that occur later.
| Default | false |
|---|
userDefinedSplit={userDefinedSplit}
indicates that you will manually define the split of a single node as specified in the nodeId parameter. The userDefinedSplit parameter contains sub-parameters needed for defining this split.
| Long form | userDefinedSplit={splitVar="variable-name"} |
|---|---|
| Shortcut form | userDefinedSplit="variable-name" |
The userDefinedSplit value can be one or more of the following:
intervalSplit={double-1 <, double-2, ...>}
is a list of interval split points. The number of branches is the number of numbers listed plus one. The branches are numbered with the 0 branch being all values less than the least value. Each other branch is numbered in increasing order.
missingBranch=64-bit-integer
indicates to which branch missing values should be assigned. By default, missing values are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
nominalSplit={any-list-or-data-type-1 <, any-list-or-data-type-2, ...>}
is a list of string lists that contain nominal split values. The number of branches is the number of string lists. The branches are ordered in which they are given to the action, with the first branch being numbered as branch 0.
splitVar="variable-name"
indicates which variable the decision tree is to use for splitting.
unseenBranch=64-bit-integer
indicates to which branch unseen levels should be assigned. By default, unseen levels are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
varImp=true | false
specifies whether the variable importance information is generated. The importance value is determined by the total Gini reduction.
| Default | false |
|---|
dtreeSplit Action
Splits decision tree nodes.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametermodelTable |
— |
specifies the table containing the model. |
|
required parametertable |
— |
specifies the settings for an input table. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the table to store the decision tree model in. When not specified, a random name is generated. | |
|
casOut |
requests that the action produce SAS score code. Specify additional parameters. | |
|
— |
specifies the table to store the generated aStore model. |
Parameter Descriptions
alpha=double
specifies the value to use for minimal cost-complexity pruning for regression trees.
| Minimum value | 0 |
|---|
applyRowOrder=True | False
Specifies that you wish the action use a prespecified row ordering. This requires using the orderby and groupby parameters on a preliminary table.partition action call.
| Alias | reproducibleRowOrder |
|---|---|
| Default | False |
attributes=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies temporary attributes, such as a format, to apply to input variables.
For more information about specifying the attributes parameter, see the common casinvardesc parameter.
| Aliases | attribute |
|---|---|
| attrs | |
| attr | |
| varAttrs |
binOrder=True | False
by default, the bin order is preserved for numeric variables. When set to False, the bin order is ignored for numeric variables.
| Default | True |
|---|
bonferroni=True | False
when set to True, specifies to perform a Bonferroni correction when the split criterion uses a chi-square statistic or CHAID.
| Default | False |
|---|
casOut={casouttable}
specifies the table to store the decision tree model in. When not specified, a random name is generated.
For more information about specifying the casOut parameter, see the common casouttable parameter.
cfLev=double
specifies the aggressiveness of tree pruning according to the C4.5 algorithm.
| Default | 0.25 |
|---|
code={codegen}
requests that the action produce SAS score code. Specify additional parameters.
For more information about specifying the code parameter, see the common codegen parameter.
crit="CHAID" | "CHISQUARE" | "FTEST" | "GAIN" | "GAINRATIO" | "GINI" | "VARIANCE"
specifies the split criterion for each tree node.
encodeName=True | False
specifies whether to encode the variable names such as predicted probabilities of a binary or nominal target in the generated casout table. The predicted probabilities are named with the prefix P_ instead of _DT_P_.
| Default | False |
|---|
freq="variable-name"
specifies a numeric variable that contains the frequency of occurrence of each observation.
greedy=True | False
by default, a greedy search or exhaustive search is used to determine the best split for each variable of each tree node. When set to False, a fast and efficient algorithm that is based on clustering is applied. Setting this parameter to False is recommended for variables with high cardinality.
| Default | True |
|---|
includeMissing=True | False
by default, observations with missing values are included. When set to False, observations with missing values for the analysis variables are excluded.
| Default | True |
|---|
* inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the input variables to use in the analysis.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | input |
|---|
leafSize=integer
specifies the minimum number of observations on each node.
| Default | 5 |
|---|---|
| Minimum value | 1 |
maxBranch=integer
specifies the maximum number of children (branches) allowed for each level of the tree.
| Default | 2 |
|---|---|
| Minimum value | 1 |
maxLevel=integer
specifies the maximum number of the tree level.
| Default | 6 |
|---|---|
| Minimum value | 1 |
mergeBin=True | False
by default, when the largest value in one bin matches the lowest value in a neighboring bin, the values are merged into the lower bin. When set to False, the action does not try to merge bins.
| Default | True |
|---|
minGain=double
specifies the minimum value to use to validate a splitting point when the criteria is not chi-square or CHAID.
| Minimum value | 0 |
|---|
minUseInSearch=integer
specifies a threshold for utilizing missing values in the split search when the missing parameter is set to USEINSEARCH. If the number of observations in which the splitting variable has missing values in a node is greater than or equal to the specified value, then the action initiates the USEINSEARCH policy. Otherwise, the missing values are assigned to a popular branch.
| Default | 1 |
|---|
missing="BRANCH" | "MACSMALL" | "POPULAR" | "SIMILAR" | "USEINSEARCH"
specifies the missing policy to handle missing values.
| Default | USEINSEARCH |
|---|
MACSMALL
specifies to treat the missing values for numeric variables as the smallest machine value and to treat missing values for nominal variables as a separate level.
modelId="string"
specifies the model ID variable name to use when generating SAS score code. By default, _DT_ is prefixed to the target variable name and _ is appended.
* modelTable={castable}
specifies the table containing the model.
| Long form | modelTable={"name":"table-name"} |
|---|---|
| Shortcut form | modelTable="table-name" |
| Alias | model |
|---|
The castable value can be one or more of the following:
"caslib":"string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
"computedOnDemand":True | False
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | False |
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"computedVarsProgram":"string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import_ |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* "name":"table-name"
specifies the name of the input table.
"singlePass":True | False
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | False |
|---|
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"where":"where-expression"
specifies an expression for subsetting the input data.
"whereTable":{groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
"casLib":"string"
specifies the caslib for the filter table. By default, the active caslib is used.
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import_ |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* "name":"table-name"
specifies the name of the filter table.
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"where":"where-expression"
specifies an expression for subsetting the data from the filter table.
nBins=integer
specifies the number of bins to use for numeric variables in the calculation of the decision tree.
| Default | 50 |
|---|---|
| Minimum value | 1 |
nodeId=[integer-1 <, integer-2, ...>]
specifies the leaf node IDs to split. When not specified or the list is not valid then all leaf nodes in the tree model are split.
nominals=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the nominal input variables to use in the analysis.
For more information about specifying the nominals parameter, see the common casinvardesc parameter.
| Alias | nominal |
|---|
nominalSearch={tkcasdt_nomSearchOpts}
specifies the method for finding a split on a nominal input.
| Alias | nomSearch |
|---|
The tkcasdt_nomSearchOpts value can be one or more of the following:
"handling":"CLASSIC" | "ENHANCED"
"maxCategories":64-bit-integer
specifies the maximum number of levels for a splitting rule to include.
| Aliases | maxCats |
|---|---|
| maxLevels | |
| maxValues | |
| cluster | |
| minCardCluster | |
| Default | 128 |
| Minimum value | 0 |
"shrinkage":double
specifies how much weight to give the category average in the sort method.
| Default | 10 |
|---|---|
| Minimum value | 0 |
"sort":64-bit-integer
specifies the minimum cardinality of an input to use the sort method.
| Alias | minCardSort |
|---|---|
| Default | 10 |
| Minimum value | 0 |
"sortBy":"COUNT" | "TARGET"
noSplit=True | False
runs the action without splitting nodes or growing the tree. If you specify this parameter, then new columns appear in the output model table that you specify in the casOut parameter. These columns contain statistics about the data specified in the table parameter with respect to the tree model that you specify in the modelTable parameter. You should use this parameter if you want statistics about validation data.
| Default | False |
|---|
prune=True | False
specify true to use a C4.5 pruning method for classification trees or minimal cost-complexity pruning for regression trees.
| Default | False |
|---|
pVal=double
specifies the maximum P-value to use in chi-square and CHAID split criteria to validate a splitting point.
| Default | 1 |
|---|---|
| Range | 0–1 |
quantileBin=True | False
specifies bin boundaries at quantiles of numerical inputs instead of bins of equal width.
| Aliases | qbin |
|---|---|
| qtbin | |
| Default | True |
saveState={casouttable}
specifies the table to store the generated aStore model.
For more information about specifying the saveState parameter, see the common casouttable parameter.
splitOnce=True | False
when set to True, the analysis variable only appears once for each path from the tree root node to the leaf node.
| Default | False |
|---|
stat=True | False
specifies whether to include the variable splitting information for each node in the decision tree.
| Default | False |
|---|
* table={castable}
specifies the settings for an input table.
| Long form | table={"name":"table-name"} |
|---|---|
| Shortcut form | table="table-name" |
The castable value can be one or more of the following:
"caslib":"string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
"computedOnDemand":True | False
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | False |
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"computedVarsProgram":"string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import_ |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* "name":"table-name"
specifies the name of the input table.
"singlePass":True | False
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | False |
|---|
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"where":"where-expression"
specifies an expression for subsetting the input data.
"whereTable":{groupbytable}
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
"casLib":"string"
specifies the caslib for the filter table. By default, the active caslib is used.
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import_ |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* "name":"table-name"
specifies the name of the filter table.
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the length of the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"where":"where-expression"
specifies an expression for subsetting the data from the filter table.
* target="variable-name"
specifies the target or response variable for training. If the variable is numeric, but not specified in the nominal= parameter and nbinstarget= is not specified, then a regression tree is trained.
updateBin=True | False
when set to True, the bin information for all the numerical analysis variables is computed on each node that is used for a split. All the splits make use of the updated bin information. Note that this changes the bin information significantly compared to the bin information that is used in the input model. As a result, this might bias the splits that occur later.
| Default | False |
|---|
userDefinedSplit={userDefinedSplit}
indicates that you will manually define the split of a single node as specified in the nodeId parameter. The userDefinedSplit parameter contains sub-parameters needed for defining this split.
| Long form | userDefinedSplit={"splitVar":"variable-name"} |
|---|---|
| Shortcut form | userDefinedSplit="variable-name" |
The userDefinedSplit value can be one or more of the following:
"intervalSplit":[double-1 <, double-2, ...>]
is a list of interval split points. The number of branches is the number of numbers listed plus one. The branches are numbered with the 0 branch being all values less than the least value. Each other branch is numbered in increasing order.
"missingBranch":64-bit-integer
indicates to which branch missing values should be assigned. By default, missing values are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
"nominalSplit":{any-list-or-data-type-1 <, any-list-or-data-type-2, ...>}
is a list of string lists that contain nominal split values. The number of branches is the number of string lists. The branches are ordered in which they are given to the action, with the first branch being numbered as branch 0.
"splitVar":"variable-name"
indicates which variable the decision tree is to use for splitting.
"unseenBranch":64-bit-integer
indicates to which branch unseen levels should be assigned. By default, unseen levels are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
varImp=True | False
specifies whether the variable importance information is generated. The importance value is determined by the total Gini reduction.
| Default | False |
|---|
dtreeSplit Action
Splits decision tree nodes.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametermodelTable |
— |
specifies the table containing the model. |
|
required parametertable |
— |
specifies the settings for an input table. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the table to store the decision tree model in. When not specified, a random name is generated. | |
|
casOut |
requests that the action produce SAS score code. Specify additional parameters. | |
|
— |
specifies the table to store the generated aStore model. |
Parameter Descriptions
alpha=double
specifies the value to use for minimal cost-complexity pruning for regression trees.
| Minimum value | 0 |
|---|
applyRowOrder=TRUE | FALSE
Specifies that you wish the action use a prespecified row ordering. This requires using the orderby and groupby parameters on a preliminary table.partition action call.
| Alias | reproducibleRowOrder |
|---|---|
| Default | FALSE |
attributes=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies temporary attributes, such as a format, to apply to input variables.
For more information about specifying the attributes parameter, see the common casinvardesc parameter.
| Aliases | attribute |
|---|---|
| attrs | |
| attr | |
| varAttrs |
binOrder=TRUE | FALSE
by default, the bin order is preserved for numeric variables. When set to False, the bin order is ignored for numeric variables.
| Default | TRUE |
|---|
bonferroni=TRUE | FALSE
when set to True, specifies to perform a Bonferroni correction when the split criterion uses a chi-square statistic or CHAID.
| Default | FALSE |
|---|
casOut=list(casouttable)
specifies the table to store the decision tree model in. When not specified, a random name is generated.
For more information about specifying the casOut parameter, see the common casouttable parameter.
cfLev=double
specifies the aggressiveness of tree pruning according to the C4.5 algorithm.
| Default | 0.25 |
|---|
code=list(codegen)
requests that the action produce SAS score code. Specify additional parameters.
For more information about specifying the code parameter, see the common codegen parameter.
crit="CHAID" | "CHISQUARE" | "FTEST" | "GAIN" | "GAINRATIO" | "GINI" | "VARIANCE"
specifies the split criterion for each tree node.
encodeName=TRUE | FALSE
specifies whether to encode the variable names such as predicted probabilities of a binary or nominal target in the generated casout table. The predicted probabilities are named with the prefix P_ instead of _DT_P_.
| Default | FALSE |
|---|
freq="variable-name"
specifies a numeric variable that contains the frequency of occurrence of each observation.
greedy=TRUE | FALSE
by default, a greedy search or exhaustive search is used to determine the best split for each variable of each tree node. When set to False, a fast and efficient algorithm that is based on clustering is applied. Setting this parameter to False is recommended for variables with high cardinality.
| Default | TRUE |
|---|
includeMissing=TRUE | FALSE
by default, observations with missing values are included. When set to False, observations with missing values for the analysis variables are excluded.
| Default | TRUE |
|---|
* inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the input variables to use in the analysis.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | input |
|---|
leafSize=integer
specifies the minimum number of observations on each node.
| Default | 5 |
|---|---|
| Minimum value | 1 |
maxBranch=integer
specifies the maximum number of children (branches) allowed for each level of the tree.
| Default | 2 |
|---|---|
| Minimum value | 1 |
maxLevel=integer
specifies the maximum number of the tree level.
| Default | 6 |
|---|---|
| Minimum value | 1 |
mergeBin=TRUE | FALSE
by default, when the largest value in one bin matches the lowest value in a neighboring bin, the values are merged into the lower bin. When set to False, the action does not try to merge bins.
| Default | TRUE |
|---|
minGain=double
specifies the minimum value to use to validate a splitting point when the criteria is not chi-square or CHAID.
| Minimum value | 0 |
|---|
minUseInSearch=integer
specifies a threshold for utilizing missing values in the split search when the missing parameter is set to USEINSEARCH. If the number of observations in which the splitting variable has missing values in a node is greater than or equal to the specified value, then the action initiates the USEINSEARCH policy. Otherwise, the missing values are assigned to a popular branch.
| Default | 1 |
|---|
missing="BRANCH" | "MACSMALL" | "POPULAR" | "SIMILAR" | "USEINSEARCH"
specifies the missing policy to handle missing values.
| Default | USEINSEARCH |
|---|
MACSMALL
specifies to treat the missing values for numeric variables as the smallest machine value and to treat missing values for nominal variables as a separate level.
modelId="string"
specifies the model ID variable name to use when generating SAS score code. By default, _DT_ is prefixed to the target variable name and _ is appended.
* modelTable=list(castable)
specifies the table containing the model.
| Long form | modelTable=list(name="table-name") |
|---|---|
| Shortcut form | modelTable="table-name" |
| Alias | model |
|---|
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | FALSE |
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
|---|
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable=list(groupbytable)
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
nBins=integer
specifies the number of bins to use for numeric variables in the calculation of the decision tree.
| Default | 50 |
|---|---|
| Minimum value | 1 |
nodeId=list(integer-1 <, integer-2, ...>)
specifies the leaf node IDs to split. When not specified or the list is not valid then all leaf nodes in the tree model are split.
nominals=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the nominal input variables to use in the analysis.
For more information about specifying the nominals parameter, see the common casinvardesc parameter.
| Alias | nominal |
|---|
nominalSearch=list(tkcasdt_nomSearchOpts)
specifies the method for finding a split on a nominal input.
| Alias | nomSearch |
|---|
The tkcasdt_nomSearchOpts value can be one or more of the following:
handling="CLASSIC" | "ENHANCED"
maxCategories=64-bit-integer
specifies the maximum number of levels for a splitting rule to include.
| Aliases | maxCats |
|---|---|
| maxLevels | |
| maxValues | |
| cluster | |
| minCardCluster | |
| Default | 128 |
| Minimum value | 0 |
shrinkage=double
specifies how much weight to give the category average in the sort method.
| Default | 10 |
|---|---|
| Minimum value | 0 |
sort=64-bit-integer
specifies the minimum cardinality of an input to use the sort method.
| Alias | minCardSort |
|---|---|
| Default | 10 |
| Minimum value | 0 |
sortBy="COUNT" | "TARGET"
noSplit=TRUE | FALSE
runs the action without splitting nodes or growing the tree. If you specify this parameter, then new columns appear in the output model table that you specify in the casOut parameter. These columns contain statistics about the data specified in the table parameter with respect to the tree model that you specify in the modelTable parameter. You should use this parameter if you want statistics about validation data.
| Default | FALSE |
|---|
prune=TRUE | FALSE
specify true to use a C4.5 pruning method for classification trees or minimal cost-complexity pruning for regression trees.
| Default | FALSE |
|---|
pVal=double
specifies the maximum P-value to use in chi-square and CHAID split criteria to validate a splitting point.
| Default | 1 |
|---|---|
| Range | 0–1 |
quantileBin=TRUE | FALSE
specifies bin boundaries at quantiles of numerical inputs instead of bins of equal width.
| Aliases | qbin |
|---|---|
| qtbin | |
| Default | TRUE |
saveState=list(casouttable)
specifies the table to store the generated aStore model.
For more information about specifying the saveState parameter, see the common casouttable parameter.
splitOnce=TRUE | FALSE
when set to True, the analysis variable only appears once for each path from the tree root node to the leaf node.
| Default | FALSE |
|---|
stat=TRUE | FALSE
specifies whether to include the variable splitting information for each node in the decision tree.
| Default | FALSE |
|---|
* table=list(castable)
specifies the settings for an input table.
| Long form | table=list(name="table-name") |
|---|---|
| Shortcut form | table="table-name" |
The castable value can be one or more of the following:
caslib="string"
specifies the caslib for the input table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
|---|---|
| Default | FALSE |
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter. If you do not specify this parameter, then all variables from computedVarsProgram are automatically included.
| Alias | compVars |
|---|
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
|---|
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the input table.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
|---|
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variables to use in the action.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
whereTable=list(groupbytable)
specifies an input table that contains rows to use as a WHERE filter. If the vars parameter is not specified, then all the variable names that are common to the input table and the filtering table are used to find matching rows. If the where parameter for the input table and this parameter are specified, then this filtering table is applied first.
The groupbytable value can be one or more of the following:
casLib="string"
specifies the caslib for the filter table. By default, the active caslib is used.
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
specifies data source options.
| Aliases | options |
|---|---|
| dataSource |
For more information about specifying the dataSourceOptions parameter, see the common dataSourceOptions parameter.
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
specifies the settings for reading a table from a data source.
| Alias | import |
|---|
For more information about specifying the importOptions parameter, see the common importOptions parameter.
* name="table-name"
specifies the name of the filter table.
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variable names to use from the filter table.
The casinvardesc value can be one or more of the following:
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the length of the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the data from the filter table.
* target="variable-name"
specifies the target or response variable for training. If the variable is numeric, but not specified in the nominal= parameter and nbinstarget= is not specified, then a regression tree is trained.
updateBin=TRUE | FALSE
when set to True, the bin information for all the numerical analysis variables is computed on each node that is used for a split. All the splits make use of the updated bin information. Note that this changes the bin information significantly compared to the bin information that is used in the input model. As a result, this might bias the splits that occur later.
| Default | FALSE |
|---|
userDefinedSplit=list(userDefinedSplit)
indicates that you will manually define the split of a single node as specified in the nodeId parameter. The userDefinedSplit parameter contains sub-parameters needed for defining this split.
| Long form | userDefinedSplit=list(splitVar="variable-name") |
|---|---|
| Shortcut form | userDefinedSplit="variable-name" |
The userDefinedSplit value can be one or more of the following:
intervalSplit=list(double-1 <, double-2, ...>)
is a list of interval split points. The number of branches is the number of numbers listed plus one. The branches are numbered with the 0 branch being all values less than the least value. Each other branch is numbered in increasing order.
missingBranch=64-bit-integer
indicates to which branch missing values should be assigned. By default, missing values are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
nominalSplit=list(any-list-or-data-type-1 <, any-list-or-data-type-2, ...>)
is a list of string lists that contain nominal split values. The number of branches is the number of string lists. The branches are ordered in which they are given to the action, with the first branch being numbered as branch 0.
splitVar="variable-name"
indicates which variable the decision tree is to use for splitting.
unseenBranch=64-bit-integer
indicates to which branch unseen levels should be assigned. By default, unseen levels are assigned to the left-most branch. For continuous variable splits, this corresponds to the branch with the smallest values.
| Minimum value | 0 |
|---|
varImp=TRUE | FALSE
specifies whether the variable importance information is generated. The importance value is determined by the total Gini reduction.
| Default | FALSE |
|---|