Data Preprocess Action Set: Syntax

Provides actions for data preprocessing and transformation

rustats Action

Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.

dataPreprocess.rustats <result=results> <status=rc> /
casOutStats
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
complementaryStats=TRUE | FALSE,
confidenceInterval=TRUE | FALSE,
excessKurtosis=TRUE | FALSE,
freq="variable-name",
includeMissingGroup=TRUE | FALSE,
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
maxIterations=integer,
nArgumentsForEachVar={integer-1 <, integer-2, ...>},
nNominalVars=integer,
nominalStats
={
bottomK=integer,
fuzzyCompare=double,
iqv=TRUE | FALSE,
topK=integer
},
nominalVarsIndices={integer-1 <, integer-2, ...>},
outputTableOptions
={
forceTableReturn=TRUE | FALSE,
tableNames={"string-1" <, "string-2", ...>}
},
requestPackages
={{
aadLocationUseMean=TRUE | FALSE,
allCentralizedMoments=TRUE | FALSE,
allKurtoses=TRUE | FALSE,
allLocations=TRUE | FALSE,
allNominalStats=TRUE | FALSE,
allScales=TRUE | FALSE,
allSkewnesses=TRUE | FALSE,
alpha=double,
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>},
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>},
maxOrder=integer,
nPercentiles=integer,
percentiles={double-1 <, double-2, ...>},
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>},
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>},
standardizeMoments=TRUE | FALSE
}, {...}},
seed=integer,
standardError=TRUE | FALSE,
required parameter table
={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
tolerance=double,
unbiased=TRUE | FALSE,
varsToArgumentsMap={integer-1 <, integer-2, ...>},
weight="variable-name"
;
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOutStats

—

specifies the settings for an output table.

Parameter Descriptions

casOutStats={casouttable}

specifies the settings for an output table.

For more information about specifying the casOutStats parameter, see the common casouttable parameter.

AliascasOut

complementaryStats=TRUE | FALSE

when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.

DefaultFALSE

confidenceInterval=TRUE | FALSE

when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.

DefaultFALSE

excessKurtosis=TRUE | FALSE

when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.

DefaultTRUE

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

includeMissingGroup=TRUE | FALSE

when set to True, missing values are allowed as group-by keys.

DefaultFALSE

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

Aliasvars

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

nArgumentsForEachVar={integer-1 <, integer-2, ...>}

specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.

AliasnArgsForEachVar

nNominalVars=integer

specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.

Minimum value (exclusive)0

nominalStats={nominalStatistics}

computes nominal statistics. These include indices of qualitative variation and most and least frequent values.

The nominalStatistics value can be one or more of the following:

bottomK=integer

specifies the number of least frequent items to compute.

Default5
distinctCountLimit=integer

specifies the distinct count limit.

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05
iqv=TRUE | FALSE

when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.

DefaultTRUE
topK=integer

specifies the number of most frequent items to compute.

Default5

nominalVarsIndices={integer-1 <, integer-2, ...>}

specifies the indices of the variables to treat as nominal variables.

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliasoutputTableOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=TRUE | FALSE

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFALSE
tableNames={"string-1" <, "string-2", ...>}

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

requestPackages={{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>}

specifies an array of robust univariate statistics request packages to be processed by the action.

AliasreqPacks

The rustatsRequestPackage value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
allCentralizedMoments=TRUE | FALSE

when set to True, computes all centralized moments (up to the sixth order).

DefaultFALSE
allKurtoses=TRUE | FALSE

when set to True, computes all kurtosis estimates.

DefaultFALSE
allLocations=TRUE | FALSE

when set to True, computes all location estimates.

DefaultFALSE
allNominalStats=TRUE | FALSE

computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.

DefaultFALSE
allScales=TRUE | FALSE

when set to True, computes all scale estimates.

DefaultFALSE
allSkewnesses=TRUE | FALSE

when set to True, computes all skewness estimates.

DefaultFALSE
alpha=double

specifies the significance level for the confidence interval of location estimates.

Default0.05
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>}

specifies a list of the kurtosis estimate types to compute.

AVGQUANTILEuses averaged quantile based skewness estimate.
CLASSICALuses the classical kurtosis estimate.
MOORSuses Moor kurtosis estimate.
QUANTILEuses quantile based kurtosis estimate.
kurtosisAlphaQuantile=double

specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtAlphaQuantile
Default2.5
Range(0–50]
kurtosisBetaQuantile=double

specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtBetaQuantile
Default25
Range(0–50]
locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Default6
Minimum value (exclusive)0
locationLowerPercentile=double

specifies the lower percentile for location estimates that need a percentile.

AliaslocLowerPerc
Default5
Range(0, 50)
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>}

specifies a list of the location estimate types to compute.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
locationSymmetricPercentile=double

specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliaslocSymPerc
Default10
Range(0, 100)
locationUpperPercentile=double

specifies the upper percentile for location estimates that need a percentile.

AliaslocUpperPerc
Default95
Range(50, 100)
maxOrder=integer

specifies the maximum order of the centralized moment to compute.

AliasnMoments
Default4
Maximum value6
nPercentiles=integer

computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.

Default0
Minimum value0
percentiles={double-1 <, double-2, ...>}

specifies a list of quantiles or percentiles to compute.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Default9
Minimum value (exclusive)0
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>}

specifies a list of the scale estimate types to compute.

AADuses the absolute deviation about the mean or median as scale.
BIWEIGHTuses Tukey biweight based estimate for scale.
GINIuses the Gini scale for scale.
IQRuses the inter-quartile range for scale.
MADuses the median absolute deviation about the median for scale.
STDuses the standard deviation for scale.
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>}

specifies a list of the skewness estimate types to compute.

AVGQUANTILEuses average quantile based skewness estimate.
CLASSICALuses the classical skewness coefficient for skewness.
PEARMEDuses Pearson median skewness coefficient for skewness.
QUANTILEuses quantile based skewness estimate.
skewnessQuantile=double

specifies the quantile for quantile skewness estimator.

AliasskewQuantile
Default25
Range(0, 50)
standardizeMoments=TRUE | FALSE

specifies that standardized moments be generated, instead of the centralized moments.

AliasstdMoments
DefaultFALSE

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

standardError=TRUE | FALSE

when set to True, standard errors for location are included in the results.

AliasstdErr
DefaultFALSE

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

unbiased=TRUE | FALSE

when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.

DefaultTRUE

varsToArgumentsMap={integer-1 <, integer-2, ...>}

specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.

weight="variable-name"

specifies the weight variable.

rustats Action

Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.

results, info = s:dataPreprocess_rustats{
casOutStats
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
complementaryStats=true | false,
confidenceInterval=true | false,
excessKurtosis=true | false,
freq="variable-name",
includeMissingGroup=true | false,
inputs
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
maxIterations=integer,
nArgumentsForEachVar={integer-1 <, integer-2, ...>},
nNominalVars=integer,
nominalStats
={
bottomK=integer,
fuzzyCompare=double,
iqv=true | false,
topK=integer
},
nominalVarsIndices={integer-1 <, integer-2, ...>},
outputTableOptions
={
forceTableReturn=true | false,
tableNames={"string-1" <, "string-2", ...>}
},
requestPackages
={{
aadLocationUseMean=true | false,
allCentralizedMoments=true | false,
allKurtoses=true | false,
allLocations=true | false,
allNominalStats=true | false,
allScales=true | false,
allSkewnesses=true | false,
alpha=double,
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>},
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>},
maxOrder=integer,
nPercentiles=integer,
percentiles={double-1 <, double-2, ...>},
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>},
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>},
standardizeMoments=true | false
}, {...}},
seed=integer,
standardError=true | false,
required parameter table
={
caslib="string",
computedOnDemand=true | false,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
tolerance=double,
unbiased=true | false,
varsToArgumentsMap={integer-1 <, integer-2, ...>},
weight="variable-name"
}
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOutStats

—

specifies the settings for an output table.

Parameter Descriptions

casOutStats={casouttable}

specifies the settings for an output table.

For more information about specifying the casOutStats parameter, see the common casouttable parameter.

AliascasOut

complementaryStats=true | false

when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.

Defaultfalse

confidenceInterval=true | false

when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.

Defaultfalse

excessKurtosis=true | false

when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.

Defaulttrue

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

includeMissingGroup=true | false

when set to True, missing values are allowed as group-by keys.

Defaultfalse

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

Aliasvars

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

nArgumentsForEachVar={integer-1 <, integer-2, ...>}

specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.

AliasnArgsForEachVar

nNominalVars=integer

specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.

Minimum value (exclusive)0

nominalStats={nominalStatistics}

computes nominal statistics. These include indices of qualitative variation and most and least frequent values.

The nominalStatistics value can be one or more of the following:

bottomK=integer

specifies the number of least frequent items to compute.

Default5
distinctCountLimit=integer

specifies the distinct count limit.

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05
iqv=true | false

when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.

Defaulttrue
topK=integer

specifies the number of most frequent items to compute.

Default5

nominalVarsIndices={integer-1 <, integer-2, ...>}

specifies the indices of the variables to treat as nominal variables.

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliasoutputTableOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=true | false

when set to True, result tables are returned to the client even if the output is also saved as an output table.

Defaultfalse
tableNames={"string-1" <, "string-2", ...>}

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

requestPackages={{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>}

specifies an array of robust univariate statistics request packages to be processed by the action.

AliasreqPacks

The rustatsRequestPackage value can be one or more of the following:

aadLocationUseMean=true | false

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
Defaulttrue
allCentralizedMoments=true | false

when set to True, computes all centralized moments (up to the sixth order).

Defaultfalse
allKurtoses=true | false

when set to True, computes all kurtosis estimates.

Defaultfalse
allLocations=true | false

when set to True, computes all location estimates.

Defaultfalse
allNominalStats=true | false

computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.

Defaultfalse
allScales=true | false

when set to True, computes all scale estimates.

Defaultfalse
allSkewnesses=true | false

when set to True, computes all skewness estimates.

Defaultfalse
alpha=double

specifies the significance level for the confidence interval of location estimates.

Default0.05
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>}

specifies a list of the kurtosis estimate types to compute.

AVGQUANTILEuses averaged quantile based skewness estimate.
CLASSICALuses the classical kurtosis estimate.
MOORSuses Moor kurtosis estimate.
QUANTILEuses quantile based kurtosis estimate.
kurtosisAlphaQuantile=double

specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtAlphaQuantile
Default2.5
Range(0–50]
kurtosisBetaQuantile=double

specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtBetaQuantile
Default25
Range(0–50]
locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Default6
Minimum value (exclusive)0
locationLowerPercentile=double

specifies the lower percentile for location estimates that need a percentile.

AliaslocLowerPerc
Default5
Range(0, 50)
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>}

specifies a list of the location estimate types to compute.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
locationSymmetricPercentile=double

specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliaslocSymPerc
Default10
Range(0, 100)
locationUpperPercentile=double

specifies the upper percentile for location estimates that need a percentile.

AliaslocUpperPerc
Default95
Range(50, 100)
maxOrder=integer

specifies the maximum order of the centralized moment to compute.

AliasnMoments
Default4
Maximum value6
nPercentiles=integer

computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.

Default0
Minimum value0
percentiles={double-1 <, double-2, ...>}

specifies a list of quantiles or percentiles to compute.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Default9
Minimum value (exclusive)0
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>}

specifies a list of the scale estimate types to compute.

AADuses the absolute deviation about the mean or median as scale.
BIWEIGHTuses Tukey biweight based estimate for scale.
GINIuses the Gini scale for scale.
IQRuses the inter-quartile range for scale.
MADuses the median absolute deviation about the median for scale.
STDuses the standard deviation for scale.
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>}

specifies a list of the skewness estimate types to compute.

AVGQUANTILEuses average quantile based skewness estimate.
CLASSICALuses the classical skewness coefficient for skewness.
PEARMEDuses Pearson median skewness coefficient for skewness.
QUANTILEuses quantile based skewness estimate.
skewnessQuantile=double

specifies the quantile for quantile skewness estimator.

AliasskewQuantile
Default25
Range(0, 50)
standardizeMoments=true | false

specifies that standardized moments be generated, instead of the centralized moments.

AliasstdMoments
Defaultfalse

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

standardError=true | false

when set to True, standard errors for location are included in the results.

AliasstdErr
Defaultfalse

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

unbiased=true | false

when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.

Defaulttrue

varsToArgumentsMap={integer-1 <, integer-2, ...>}

specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.

weight="variable-name"

specifies the weight variable.

rustats Action

Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.

results=s.dataPreprocess.rustats(
casOutStats
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
complementaryStats=True | False,
confidenceInterval=True | False,
excessKurtosis=True | False,
freq="variable-name",
includeMissingGroup=True | False,
inputs
=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
maxIterations=integer,
nArgumentsForEachVar=[integer-1 <, integer-2, ...>],
nNominalVars=integer,
nominalStats
={
"bottomK":integer,
"fuzzyCompare":double,
"iqv":True | False,
"topK":integer
},
nominalVarsIndices=[integer-1 <, integer-2, ...>],
outputTableOptions
={
"forceTableReturn":True | False,
"tableNames":["string-1" <, "string-2", ...>]
},
quantileSketch
={
"epsilon":double
},
requestPackages
=[{
"aadLocationUseMean":True | False,
"allCentralizedMoments":True | False,
"allKurtoses":True | False,
"allLocations":True | False,
"allNominalStats":True | False,
"allScales":True | False,
"allSkewnesses":True | False,
"alpha":double,
"kurtoses":["AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"] | ["string-1" <, "string-2", ...>],
"locations":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | ["string-1" <, "string-2", ...>],
"maxOrder":integer,
"nPercentiles":integer,
"percentiles":[double-1 <, double-2, ...>],
"scales":["AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"] | ["string-1" <, "string-2", ...>],
"skewnesses":["AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"] | ["string-1" <, "string-2", ...>],
"standardizeMoments":True | False
}<, {...}>],
seed=integer,
standardError=True | False,
required parameter table
={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"groupByMode":"NOSORT" | "REDISTRIBUTE",
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression",
"whereTable"
:{
"casLib":"string"
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter "name":"table-name"
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>]
"where":"where-expression"
}
},
tolerance=double,
unbiased=True | False,
varsToArgumentsMap=[integer-1 <, integer-2, ...>],
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOutStats

—

specifies the settings for an output table.

Parameter Descriptions

casOutStats={casouttable}

specifies the settings for an output table.

For more information about specifying the casOutStats parameter, see the common casouttable parameter.

AliascasOut

complementaryStats=True | False

when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.

DefaultFalse

confidenceInterval=True | False

when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.

DefaultFalse

excessKurtosis=True | False

when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.

DefaultTrue

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

includeMissingGroup=True | False

when set to True, missing values are allowed as group-by keys.

DefaultFalse

inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

Aliasvars

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

nArgumentsForEachVar=[integer-1 <, integer-2, ...>]

specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.

AliasnArgsForEachVar

nNominalVars=integer

specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.

Minimum value (exclusive)0

nominalStats={nominalStatistics}

computes nominal statistics. These include indices of qualitative variation and most and least frequent values.

The nominalStatistics value can be one or more of the following:

"bottomK":integer

specifies the number of least frequent items to compute.

Default5
"distinctCountLimit":integer

specifies the distinct count limit.

"fuzzyCompare":double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05
"iqv":True | False

when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.

DefaultTrue
"topK":integer

specifies the number of most frequent items to compute.

Default5

nominalVarsIndices=[integer-1 <, integer-2, ...>]

specifies the indices of the variables to treat as nominal variables.

outputTableOptions={outputTableOptions}

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliasoutputTableOpts

The outputTableOptions value can be one or more of the following:

"forceTableReturn":True | False

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFalse
"tableNames":["string-1" <, "string-2", ...>]

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch={quantileSketchOptions}

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

"compressionFactor":double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
"epsilon":double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

requestPackages=[{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>]

specifies an array of robust univariate statistics request packages to be processed by the action.

AliasreqPacks

The rustatsRequestPackage value can be one or more of the following:

"aadLocationUseMean":True | False

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTrue
"allCentralizedMoments":True | False

when set to True, computes all centralized moments (up to the sixth order).

DefaultFalse
"allKurtoses":True | False

when set to True, computes all kurtosis estimates.

DefaultFalse
"allLocations":True | False

when set to True, computes all location estimates.

DefaultFalse
"allNominalStats":True | False

computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.

DefaultFalse
"allScales":True | False

when set to True, computes all scale estimates.

DefaultFalse
"allSkewnesses":True | False

when set to True, computes all skewness estimates.

DefaultFalse
"alpha":double

specifies the significance level for the confidence interval of location estimates.

Default0.05
"kurtoses":["AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"] | ["string-1" <, "string-2", ...>]

specifies a list of the kurtosis estimate types to compute.

AVGQUANTILEuses averaged quantile based skewness estimate.
CLASSICALuses the classical kurtosis estimate.
MOORSuses Moor kurtosis estimate.
QUANTILEuses quantile based kurtosis estimate.
"kurtosisAlphaQuantile":double

specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtAlphaQuantile
Default2.5
Range(0–50]
"kurtosisBetaQuantile":double

specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtBetaQuantile
Default25
Range(0–50]
"locationBiweightTuning":double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Default6
Minimum value (exclusive)0
"locationLowerPercentile":double

specifies the lower percentile for location estimates that need a percentile.

AliaslocLowerPerc
Default5
Range(0, 50)
"locations":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | ["string-1" <, "string-2", ...>]

specifies a list of the location estimate types to compute.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
"locationSymmetricPercentile":double

specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliaslocSymPerc
Default10
Range(0, 100)
"locationUpperPercentile":double

specifies the upper percentile for location estimates that need a percentile.

AliaslocUpperPerc
Default95
Range(50, 100)
"maxOrder":integer

specifies the maximum order of the centralized moment to compute.

AliasnMoments
Default4
Maximum value6
"nPercentiles":integer

computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.

Default0
Minimum value0
"percentiles":[double-1 <, double-2, ...>]

specifies a list of quantiles or percentiles to compute.

"scaleBiweightTuning":double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Default9
Minimum value (exclusive)0
"scales":["AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"] | ["string-1" <, "string-2", ...>]

specifies a list of the scale estimate types to compute.

AADuses the absolute deviation about the mean or median as scale.
BIWEIGHTuses Tukey biweight based estimate for scale.
GINIuses the Gini scale for scale.
IQRuses the inter-quartile range for scale.
MADuses the median absolute deviation about the median for scale.
STDuses the standard deviation for scale.
"skewnesses":["AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"] | ["string-1" <, "string-2", ...>]

specifies a list of the skewness estimate types to compute.

AVGQUANTILEuses average quantile based skewness estimate.
CLASSICALuses the classical skewness coefficient for skewness.
PEARMEDuses Pearson median skewness coefficient for skewness.
QUANTILEuses quantile based skewness estimate.
"skewnessQuantile":double

specifies the quantile for quantile skewness estimator.

AliasskewQuantile
Default25
Range(0, 50)
"standardizeMoments":True | False

specifies that standardized moments be generated, instead of the centralized moments.

AliasstdMoments
DefaultFalse

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

standardError=True | False

when set to True, standard errors for location are included in the results.

AliasstdErr
DefaultFalse

* table={castable}

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

unbiased=True | False

when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.

DefaultTrue

varsToArgumentsMap=[integer-1 <, integer-2, ...>]

specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.

weight="variable-name"

specifies the weight variable.

rustats Action

Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.

results <– cas.dataPreprocess.rustats(s,
casOutStats
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
complementaryStats=TRUE | FALSE,
confidenceInterval=TRUE | FALSE,
excessKurtosis=TRUE | FALSE,
freq="variable-name",
includeMissingGroup=TRUE | FALSE,
inputs
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
maxIterations=integer,
nArgumentsForEachVar=list(integer-1 <, integer-2, ...>),
nNominalVars=integer,
nominalStats
=list(
bottomK=integer,
fuzzyCompare=double,
iqv=TRUE | FALSE,
topK=integer
),
nominalVarsIndices=list(integer-1 <, integer-2, ...>),
outputTableOptions
=list(
forceTableReturn=TRUE | FALSE,
tableNames=list("string-1" <, "string-2", ...>)
),
quantileSketch
=list(
epsilon=double
),
requestPackages
=list( list(
aadLocationUseMean=TRUE | FALSE,
allCentralizedMoments=TRUE | FALSE,
allKurtoses=TRUE | FALSE,
allLocations=TRUE | FALSE,
allNominalStats=TRUE | FALSE,
allScales=TRUE | FALSE,
allSkewnesses=TRUE | FALSE,
alpha=double,
kurtoses=list("AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE") | list("string-1" <, "string-2", ...>),
locations=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list("string-1" <, "string-2", ...>),
maxOrder=integer,
nPercentiles=integer,
percentiles=list(double-1 <, double-2, ...>),
scales=list("AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD") | list("string-1" <, "string-2", ...>),
skewnesses=list("AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE") | list("string-1" <, "string-2", ...>),
standardizeMoments=TRUE | FALSE
) <, list(...)>),
seed=integer,
standardError=TRUE | FALSE,
required parameter table
=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
computedVarsProgram="string",
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression",
whereTable
=list(
casLib="string"
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | hadoop-parameters | hana-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
required parameter name="table-name"
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>)
where="where-expression"
)
),
tolerance=double,
unbiased=TRUE | FALSE,
varsToArgumentsMap=list(integer-1 <, integer-2, ...>),
weight="variable-name"
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

required parametertable

—

specifies the table name, caslib, and other common parameters.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 casOutStats

—

specifies the settings for an output table.

Parameter Descriptions

casOutStats=list(casouttable)

specifies the settings for an output table.

For more information about specifying the casOutStats parameter, see the common casouttable parameter.

AliascasOut

complementaryStats=TRUE | FALSE

when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.

DefaultFALSE

confidenceInterval=TRUE | FALSE

when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.

DefaultFALSE

excessKurtosis=TRUE | FALSE

when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.

DefaultTRUE

freq="variable-name"

specifies the frequency variable.

Aliasfrequency

includeMissingGroup=TRUE | FALSE

when set to True, missing values are allowed as group-by keys.

DefaultFALSE

inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.

For more information about specifying the inputs parameter, see the common casinvardesc parameter.

Aliasvars

maxIterations=integer

specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.

AliasesmaxIters
rustatsMaxNiters

nArgumentsForEachVar=list(integer-1 <, integer-2, ...>)

specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.

AliasnArgsForEachVar

nNominalVars=integer

specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.

Minimum value (exclusive)0

nominalStats=list(nominalStatistics)

computes nominal statistics. These include indices of qualitative variation and most and least frequent values.

The nominalStatistics value can be one or more of the following:

bottomK=integer

specifies the number of least frequent items to compute.

Default5
distinctCountLimit=integer

specifies the distinct count limit.

fuzzyCompare=double

specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.

Aliasprecision
Range0–1E-05
iqv=TRUE | FALSE

when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.

DefaultTRUE
topK=integer

specifies the number of most frequent items to compute.

Default5

nominalVarsIndices=list(integer-1 <, integer-2, ...>)

specifies the indices of the variables to treat as nominal variables.

outputTableOptions=list(outputTableOptions)

specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.

AliasoutputTableOpts

The outputTableOptions value can be one or more of the following:

forceTableReturn=TRUE | FALSE

when set to True, result tables are returned to the client even if the output is also saved as an output table.

DefaultFALSE
tableNames=list("string-1" <, "string-2", ...>)

specifies the names of result tables to generate. By default, all result tables are returned.

AliasoutputTables

percentileDefinition=integer

specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.

AliaspctlDef
Default6
Range1–6

percentileMaxIterations=integer

specifies the maximum number of iterations for percentile computation.

AliaspctlMaxIters

percentileTolerance=double

specifies the tolerance for percentile computation.

AliaspctlEpsilon
Default1E-05

quantileSketch=list(quantileSketchOptions)

specifies the options for quantile sketch.

AliasquantileSketchOptions

The quantileSketchOptions value can be one or more of the following:

compressionFactor=double

specifies the compression factor to use for quantile sketch.

Default10
Minimum value1
epsilon=double

specifies the tolerance to use for quantile sketch.

Default0.001
Range1E-06–0.1

requestPackages=list( list(rustatsRequestPackage-1) <, list(rustatsRequestPackage-2), ...>)

specifies an array of robust univariate statistics request packages to be processed by the action.

AliasreqPacks

The rustatsRequestPackage value can be one or more of the following:

aadLocationUseMean=TRUE | FALSE

when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.

AliasaadLocUseMean
DefaultTRUE
allCentralizedMoments=TRUE | FALSE

when set to True, computes all centralized moments (up to the sixth order).

DefaultFALSE
allKurtoses=TRUE | FALSE

when set to True, computes all kurtosis estimates.

DefaultFALSE
allLocations=TRUE | FALSE

when set to True, computes all location estimates.

DefaultFALSE
allNominalStats=TRUE | FALSE

computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.

DefaultFALSE
allScales=TRUE | FALSE

when set to True, computes all scale estimates.

DefaultFALSE
allSkewnesses=TRUE | FALSE

when set to True, computes all skewness estimates.

DefaultFALSE
alpha=double

specifies the significance level for the confidence interval of location estimates.

Default0.05
kurtoses=list("AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE") | list("string-1" <, "string-2", ...>)

specifies a list of the kurtosis estimate types to compute.

AVGQUANTILEuses averaged quantile based skewness estimate.
CLASSICALuses the classical kurtosis estimate.
MOORSuses Moor kurtosis estimate.
QUANTILEuses quantile based kurtosis estimate.
kurtosisAlphaQuantile=double

specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtAlphaQuantile
Default2.5
Range(0–50]
kurtosisBetaQuantile=double

specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.

AliaskurtBetaQuantile
Default25
Range(0–50]
locationBiweightTuning=double

specifies the tuning factor for the Tukey biweight location estimator.

AliaslocBiweightTuning
Default6
Minimum value (exclusive)0
locationLowerPercentile=double

specifies the lower percentile for location estimates that need a percentile.

AliaslocLowerPerc
Default5
Range(0, 50)
locations=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list("string-1" <, "string-2", ...>)

specifies a list of the location estimate types to compute.

BIWEIGHTuses Tukey biweight based estimate for location.
GEOMETRICMEANuses the geometric mean for location.
HARMONICMEANuses the harmonic mean for location.
MEANuses the arithmetic mean for location.
MEDIANuses the median value for location.
TRIMMEDMEANuses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
WINSORIZEDMEANuses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters.
locationSymmetricPercentile=double

specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.

AliaslocSymPerc
Default10
Range(0, 100)
locationUpperPercentile=double

specifies the upper percentile for location estimates that need a percentile.

AliaslocUpperPerc
Default95
Range(50, 100)
maxOrder=integer

specifies the maximum order of the centralized moment to compute.

AliasnMoments
Default4
Maximum value6
nPercentiles=integer

computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.

Default0
Minimum value0
percentiles=list(double-1 <, double-2, ...>)

specifies a list of quantiles or percentiles to compute.

scaleBiweightTuning=double

specifies the tuning factor for the Tukey biweight scale estimator.

AliassclBiweightTuning
Default9
Minimum value (exclusive)0
scales=list("AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD") | list("string-1" <, "string-2", ...>)

specifies a list of the scale estimate types to compute.

AADuses the absolute deviation about the mean or median as scale.
BIWEIGHTuses Tukey biweight based estimate for scale.
GINIuses the Gini scale for scale.
IQRuses the inter-quartile range for scale.
MADuses the median absolute deviation about the median for scale.
STDuses the standard deviation for scale.
skewnesses=list("AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE") | list("string-1" <, "string-2", ...>)

specifies a list of the skewness estimate types to compute.

AVGQUANTILEuses average quantile based skewness estimate.
CLASSICALuses the classical skewness coefficient for skewness.
PEARMEDuses Pearson median skewness coefficient for skewness.
QUANTILEuses quantile based skewness estimate.
skewnessQuantile=double

specifies the quantile for quantile skewness estimator.

AliasskewQuantile
Default25
Range(0, 50)
standardizeMoments=TRUE | FALSE

specifies that standardized moments be generated, instead of the centralized moments.

AliasstdMoments
DefaultFALSE

seed=integer

specifies a seed value. The seed is used to generate random values.

Default0

standardError=TRUE | FALSE

when set to True, standard errors for location are included in the results.

AliasstdErr
DefaultFALSE

* table=list(castable)

specifies the table name, caslib, and other common parameters.

For more information about specifying the table parameter, see the common castable parameter.

tolerance=double

specifies the tolerance for the iterative robust univariate statistics.

Default1E-05

unbiased=TRUE | FALSE

when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.

DefaultTRUE

varsToArgumentsMap=list(integer-1 <, integer-2, ...>)

specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.

weight="variable-name"

specifies the weight variable.

Last updated: February 15, 2023