Data Preprocess Action Set: Syntax
Provides actions for data preprocessing and transformation
rustats Action
Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. |
Parameter Descriptions
casOutStats={casouttable}
specifies the settings for an output table.
For more information about specifying the casOutStats parameter, see the common casouttable parameter.
| Alias | casOut |
|---|
complementaryStats=TRUE | FALSE
when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.
| Default | FALSE |
|---|
confidenceInterval=TRUE | FALSE
when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.
| Default | FALSE |
|---|
excessKurtosis=TRUE | FALSE
when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.
| Default | TRUE |
|---|
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
includeMissingGroup=TRUE | FALSE
when set to True, missing values are allowed as group-by keys.
| Default | FALSE |
|---|
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | vars |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
nArgumentsForEachVar={integer-1 <, integer-2, ...>}
specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.
| Alias | nArgsForEachVar |
|---|
nNominalVars=integer
specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.
| Minimum value (exclusive) | 0 |
|---|
nominalStats={nominalStatistics}
computes nominal statistics. These include indices of qualitative variation and most and least frequent values.
The nominalStatistics value can be one or more of the following:
bottomK=integer
specifies the number of least frequent items to compute.
| Default | 5 |
|---|
distinctCountLimit=integer
specifies the distinct count limit.
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
iqv=TRUE | FALSE
when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.
| Default | TRUE |
|---|
topK=integer
specifies the number of most frequent items to compute.
| Default | 5 |
|---|
nominalVarsIndices={integer-1 <, integer-2, ...>}
specifies the indices of the variables to treat as nominal variables.
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | outputTableOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=TRUE | FALSE
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | FALSE |
|---|
tableNames={"string-1" <, "string-2", ...>}
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
requestPackages={{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>}
specifies an array of robust univariate statistics request packages to be processed by the action.
| Alias | reqPacks |
|---|
The rustatsRequestPackage value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
allCentralizedMoments=TRUE | FALSE
when set to True, computes all centralized moments (up to the sixth order).
| Default | FALSE |
|---|
allKurtoses=TRUE | FALSE
when set to True, computes all kurtosis estimates.
| Default | FALSE |
|---|
allLocations=TRUE | FALSE
when set to True, computes all location estimates.
| Default | FALSE |
|---|
allNominalStats=TRUE | FALSE
computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.
| Default | FALSE |
|---|
allScales=TRUE | FALSE
when set to True, computes all scale estimates.
| Default | FALSE |
|---|
allSkewnesses=TRUE | FALSE
when set to True, computes all skewness estimates.
| Default | FALSE |
|---|
alpha=double
specifies the significance level for the confidence interval of location estimates.
| Default | 0.05 |
|---|
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>}
specifies a list of the kurtosis estimate types to compute.
| AVGQUANTILE | uses averaged quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical kurtosis estimate. |
| MOORS | uses Moor kurtosis estimate. |
| QUANTILE | uses quantile based kurtosis estimate. |
kurtosisAlphaQuantile=double
specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtAlphaQuantile |
|---|---|
| Default | 2.5 |
| Range | (0–50] |
kurtosisBetaQuantile=double
specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtBetaQuantile |
|---|---|
| Default | 25 |
| Range | (0–50] |
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Default | 6 |
| Minimum value (exclusive) | 0 |
locationLowerPercentile=double
specifies the lower percentile for location estimates that need a percentile.
| Alias | locLowerPerc |
|---|---|
| Default | 5 |
| Range | (0, 50) |
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>}
specifies a list of the location estimate types to compute.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
locationSymmetricPercentile=double
specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | locSymPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
locationUpperPercentile=double
specifies the upper percentile for location estimates that need a percentile.
| Alias | locUpperPerc |
|---|---|
| Default | 95 |
| Range | (50, 100) |
maxOrder=integer
specifies the maximum order of the centralized moment to compute.
| Alias | nMoments |
|---|---|
| Default | 4 |
| Maximum value | 6 |
nPercentiles=integer
computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.
| Default | 0 |
|---|---|
| Minimum value | 0 |
percentiles={double-1 <, double-2, ...>}
specifies a list of quantiles or percentiles to compute.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Default | 9 |
| Minimum value (exclusive) | 0 |
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>}
specifies a list of the scale estimate types to compute.
| AAD | uses the absolute deviation about the mean or median as scale. |
|---|---|
| BIWEIGHT | uses Tukey biweight based estimate for scale. |
| GINI | uses the Gini scale for scale. |
| IQR | uses the inter-quartile range for scale. |
| MAD | uses the median absolute deviation about the median for scale. |
| STD | uses the standard deviation for scale. |
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>}
specifies a list of the skewness estimate types to compute.
| AVGQUANTILE | uses average quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical skewness coefficient for skewness. |
| PEARMED | uses Pearson median skewness coefficient for skewness. |
| QUANTILE | uses quantile based skewness estimate. |
skewnessQuantile=double
specifies the quantile for quantile skewness estimator.
| Alias | skewQuantile |
|---|---|
| Default | 25 |
| Range | (0, 50) |
standardizeMoments=TRUE | FALSE
specifies that standardized moments be generated, instead of the centralized moments.
| Alias | stdMoments |
|---|---|
| Default | FALSE |
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
standardError=TRUE | FALSE
when set to True, standard errors for location are included in the results.
| Alias | stdErr |
|---|---|
| Default | FALSE |
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
unbiased=TRUE | FALSE
when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.
| Default | TRUE |
|---|
varsToArgumentsMap={integer-1 <, integer-2, ...>}
specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.
weight="variable-name"
specifies the weight variable.
rustats Action
Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. |
Parameter Descriptions
casOutStats={casouttable}
specifies the settings for an output table.
For more information about specifying the casOutStats parameter, see the common casouttable parameter.
| Alias | casOut |
|---|
complementaryStats=true | false
when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.
| Default | false |
|---|
confidenceInterval=true | false
when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.
| Default | false |
|---|
excessKurtosis=true | false
when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.
| Default | true |
|---|
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
includeMissingGroup=true | false
when set to True, missing values are allowed as group-by keys.
| Default | false |
|---|
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | vars |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
nArgumentsForEachVar={integer-1 <, integer-2, ...>}
specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.
| Alias | nArgsForEachVar |
|---|
nNominalVars=integer
specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.
| Minimum value (exclusive) | 0 |
|---|
nominalStats={nominalStatistics}
computes nominal statistics. These include indices of qualitative variation and most and least frequent values.
The nominalStatistics value can be one or more of the following:
bottomK=integer
specifies the number of least frequent items to compute.
| Default | 5 |
|---|
distinctCountLimit=integer
specifies the distinct count limit.
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
iqv=true | false
when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.
| Default | true |
|---|
topK=integer
specifies the number of most frequent items to compute.
| Default | 5 |
|---|
nominalVarsIndices={integer-1 <, integer-2, ...>}
specifies the indices of the variables to treat as nominal variables.
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | outputTableOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=true | false
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | false |
|---|
tableNames={"string-1" <, "string-2", ...>}
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
requestPackages={{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>}
specifies an array of robust univariate statistics request packages to be processed by the action.
| Alias | reqPacks |
|---|
The rustatsRequestPackage value can be one or more of the following:
aadLocationUseMean=true | false
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | true |
allCentralizedMoments=true | false
when set to True, computes all centralized moments (up to the sixth order).
| Default | false |
|---|
allKurtoses=true | false
when set to True, computes all kurtosis estimates.
| Default | false |
|---|
allLocations=true | false
when set to True, computes all location estimates.
| Default | false |
|---|
allNominalStats=true | false
computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.
| Default | false |
|---|
allScales=true | false
when set to True, computes all scale estimates.
| Default | false |
|---|
allSkewnesses=true | false
when set to True, computes all skewness estimates.
| Default | false |
|---|
alpha=double
specifies the significance level for the confidence interval of location estimates.
| Default | 0.05 |
|---|
kurtoses={"AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"} | {"string-1" <, "string-2", ...>}
specifies a list of the kurtosis estimate types to compute.
| AVGQUANTILE | uses averaged quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical kurtosis estimate. |
| MOORS | uses Moor kurtosis estimate. |
| QUANTILE | uses quantile based kurtosis estimate. |
kurtosisAlphaQuantile=double
specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtAlphaQuantile |
|---|---|
| Default | 2.5 |
| Range | (0–50] |
kurtosisBetaQuantile=double
specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtBetaQuantile |
|---|---|
| Default | 25 |
| Range | (0–50] |
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Default | 6 |
| Minimum value (exclusive) | 0 |
locationLowerPercentile=double
specifies the lower percentile for location estimates that need a percentile.
| Alias | locLowerPerc |
|---|---|
| Default | 5 |
| Range | (0, 50) |
locations={"BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"} | {"string-1" <, "string-2", ...>}
specifies a list of the location estimate types to compute.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
locationSymmetricPercentile=double
specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | locSymPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
locationUpperPercentile=double
specifies the upper percentile for location estimates that need a percentile.
| Alias | locUpperPerc |
|---|---|
| Default | 95 |
| Range | (50, 100) |
maxOrder=integer
specifies the maximum order of the centralized moment to compute.
| Alias | nMoments |
|---|---|
| Default | 4 |
| Maximum value | 6 |
nPercentiles=integer
computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.
| Default | 0 |
|---|---|
| Minimum value | 0 |
percentiles={double-1 <, double-2, ...>}
specifies a list of quantiles or percentiles to compute.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Default | 9 |
| Minimum value (exclusive) | 0 |
scales={"AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"} | {"string-1" <, "string-2", ...>}
specifies a list of the scale estimate types to compute.
| AAD | uses the absolute deviation about the mean or median as scale. |
|---|---|
| BIWEIGHT | uses Tukey biweight based estimate for scale. |
| GINI | uses the Gini scale for scale. |
| IQR | uses the inter-quartile range for scale. |
| MAD | uses the median absolute deviation about the median for scale. |
| STD | uses the standard deviation for scale. |
skewnesses={"AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"} | {"string-1" <, "string-2", ...>}
specifies a list of the skewness estimate types to compute.
| AVGQUANTILE | uses average quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical skewness coefficient for skewness. |
| PEARMED | uses Pearson median skewness coefficient for skewness. |
| QUANTILE | uses quantile based skewness estimate. |
skewnessQuantile=double
specifies the quantile for quantile skewness estimator.
| Alias | skewQuantile |
|---|---|
| Default | 25 |
| Range | (0, 50) |
standardizeMoments=true | false
specifies that standardized moments be generated, instead of the centralized moments.
| Alias | stdMoments |
|---|---|
| Default | false |
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
standardError=true | false
when set to True, standard errors for location are included in the results.
| Alias | stdErr |
|---|---|
| Default | false |
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
unbiased=true | false
when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.
| Default | true |
|---|
varsToArgumentsMap={integer-1 <, integer-2, ...>}
specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.
weight="variable-name"
specifies the weight variable.
rustats Action
Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. |
Parameter Descriptions
casOutStats={casouttable}
specifies the settings for an output table.
For more information about specifying the casOutStats parameter, see the common casouttable parameter.
| Alias | casOut |
|---|
complementaryStats=True | False
when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.
| Default | False |
|---|
confidenceInterval=True | False
when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.
| Default | False |
|---|
excessKurtosis=True | False
when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.
| Default | True |
|---|
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
includeMissingGroup=True | False
when set to True, missing values are allowed as group-by keys.
| Default | False |
|---|
inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | vars |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
nArgumentsForEachVar=[integer-1 <, integer-2, ...>]
specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.
| Alias | nArgsForEachVar |
|---|
nNominalVars=integer
specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.
| Minimum value (exclusive) | 0 |
|---|
nominalStats={nominalStatistics}
computes nominal statistics. These include indices of qualitative variation and most and least frequent values.
The nominalStatistics value can be one or more of the following:
"bottomK":integer
specifies the number of least frequent items to compute.
| Default | 5 |
|---|
"distinctCountLimit":integer
specifies the distinct count limit.
"fuzzyCompare":double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
"iqv":True | False
when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.
| Default | True |
|---|
"topK":integer
specifies the number of most frequent items to compute.
| Default | 5 |
|---|
nominalVarsIndices=[integer-1 <, integer-2, ...>]
specifies the indices of the variables to treat as nominal variables.
outputTableOptions={outputTableOptions}
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | outputTableOpts |
|---|
The outputTableOptions value can be one or more of the following:
"forceTableReturn":True | False
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | False |
|---|
"tableNames":["string-1" <, "string-2", ...>]
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch={quantileSketchOptions}
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
"compressionFactor":double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
"epsilon":double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
requestPackages=[{rustatsRequestPackage-1} <, {rustatsRequestPackage-2}, ...>]
specifies an array of robust univariate statistics request packages to be processed by the action.
| Alias | reqPacks |
|---|
The rustatsRequestPackage value can be one or more of the following:
"aadLocationUseMean":True | False
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | True |
"allCentralizedMoments":True | False
when set to True, computes all centralized moments (up to the sixth order).
| Default | False |
|---|
"allKurtoses":True | False
when set to True, computes all kurtosis estimates.
| Default | False |
|---|
"allLocations":True | False
when set to True, computes all location estimates.
| Default | False |
|---|
"allNominalStats":True | False
computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.
| Default | False |
|---|
"allScales":True | False
when set to True, computes all scale estimates.
| Default | False |
|---|
"allSkewnesses":True | False
when set to True, computes all skewness estimates.
| Default | False |
|---|
"alpha":double
specifies the significance level for the confidence interval of location estimates.
| Default | 0.05 |
|---|
"kurtoses":["AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE"] | ["string-1" <, "string-2", ...>]
specifies a list of the kurtosis estimate types to compute.
| AVGQUANTILE | uses averaged quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical kurtosis estimate. |
| MOORS | uses Moor kurtosis estimate. |
| QUANTILE | uses quantile based kurtosis estimate. |
"kurtosisAlphaQuantile":double
specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtAlphaQuantile |
|---|---|
| Default | 2.5 |
| Range | (0–50] |
"kurtosisBetaQuantile":double
specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtBetaQuantile |
|---|---|
| Default | 25 |
| Range | (0–50] |
"locationBiweightTuning":double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Default | 6 |
| Minimum value (exclusive) | 0 |
"locationLowerPercentile":double
specifies the lower percentile for location estimates that need a percentile.
| Alias | locLowerPerc |
|---|---|
| Default | 5 |
| Range | (0, 50) |
"locations":["BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN"] | ["string-1" <, "string-2", ...>]
specifies a list of the location estimate types to compute.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
"locationSymmetricPercentile":double
specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | locSymPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
"locationUpperPercentile":double
specifies the upper percentile for location estimates that need a percentile.
| Alias | locUpperPerc |
|---|---|
| Default | 95 |
| Range | (50, 100) |
"maxOrder":integer
specifies the maximum order of the centralized moment to compute.
| Alias | nMoments |
|---|---|
| Default | 4 |
| Maximum value | 6 |
"nPercentiles":integer
computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.
| Default | 0 |
|---|---|
| Minimum value | 0 |
"percentiles":[double-1 <, double-2, ...>]
specifies a list of quantiles or percentiles to compute.
"scaleBiweightTuning":double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Default | 9 |
| Minimum value (exclusive) | 0 |
"scales":["AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD"] | ["string-1" <, "string-2", ...>]
specifies a list of the scale estimate types to compute.
| AAD | uses the absolute deviation about the mean or median as scale. |
|---|---|
| BIWEIGHT | uses Tukey biweight based estimate for scale. |
| GINI | uses the Gini scale for scale. |
| IQR | uses the inter-quartile range for scale. |
| MAD | uses the median absolute deviation about the median for scale. |
| STD | uses the standard deviation for scale. |
"skewnesses":["AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE"] | ["string-1" <, "string-2", ...>]
specifies a list of the skewness estimate types to compute.
| AVGQUANTILE | uses average quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical skewness coefficient for skewness. |
| PEARMED | uses Pearson median skewness coefficient for skewness. |
| QUANTILE | uses quantile based skewness estimate. |
"skewnessQuantile":double
specifies the quantile for quantile skewness estimator.
| Alias | skewQuantile |
|---|---|
| Default | 25 |
| Range | (0, 50) |
"standardizeMoments":True | False
specifies that standardized moments be generated, instead of the centralized moments.
| Alias | stdMoments |
|---|---|
| Default | False |
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
standardError=True | False
when set to True, standard errors for location are included in the results.
| Alias | stdErr |
|---|---|
| Default | False |
* table={castable}
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
unbiased=True | False
when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.
| Default | True |
|---|
varsToArgumentsMap=[integer-1 <, integer-2, ...>]
specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.
weight="variable-name"
specifies the weight variable.
rustats Action
Computes robust univariate statistics, centralized moments, quantiles, and frequency distribution statistics.
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
required parametertable |
— |
specifies the table name, caslib, and other common parameters. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the settings for an output table. |
Parameter Descriptions
casOutStats=list(casouttable)
specifies the settings for an output table.
For more information about specifying the casOutStats parameter, see the common casouttable parameter.
| Alias | casOut |
|---|
complementaryStats=TRUE | FALSE
when set to True, complementary univariate statistics such as uncorrected and corrected sum of squared, and coefficient of variation are computed.
| Default | FALSE |
|---|
confidenceInterval=TRUE | FALSE
when set to True, confidence intervals are included in the results. This parameter enables the stdErr parameter.
| Default | FALSE |
|---|
excessKurtosis=TRUE | FALSE
when set to True, excess kurtosis values, instead of kurtosis values, are computed. The excess kurtosis is defined with respect to the normal distribution.
| Default | TRUE |
|---|
freq="variable-name"
specifies the frequency variable.
| Alias | frequency |
|---|
includeMissingGroup=TRUE | FALSE
when set to True, missing values are allowed as group-by keys.
| Default | FALSE |
|---|
inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variables to use for the analysis. You can specify a subset of the variables from the input table.
For more information about specifying the inputs parameter, see the common casinvardesc parameter.
| Alias | vars |
|---|
maxIterations=integer
specifies the maximum number of iterations for the iterative robust univariate statistics such as MAD scale, GINI scale, and Medcouple skewness estimates. This parameter can be used if the ZSCORE outlier definition is used.
| Aliases | maxIters |
|---|---|
| rustatsMaxNiters |
nArgumentsForEachVar=list(integer-1 <, integer-2, ...>)
specifies the number of arguments (request packages) for each variable. If not set, then all request packages are included for all variables.
| Alias | nArgsForEachVar |
|---|
nNominalVars=integer
specifies to treat the last nNomVars variables as nominal if you do not provide a value for the nomVarsIndices parameter.
| Minimum value (exclusive) | 0 |
|---|
nominalStats=list(nominalStatistics)
computes nominal statistics. These include indices of qualitative variation and most and least frequent values.
The nominalStatistics value can be one or more of the following:
bottomK=integer
specifies the number of least frequent items to compute.
| Default | 5 |
|---|
distinctCountLimit=integer
specifies the distinct count limit.
fuzzyCompare=double
specifies the fuzzy comparison threshold that is used to determine distinctness of numeric values.
| Alias | precision |
|---|---|
| Range | 0–1E-05 |
iqv=TRUE | FALSE
when set to True, computes indices of qualitative variation. These include Shannon entropy, Gini index and a few other frequency distribution dispersion measures.
| Default | TRUE |
|---|
topK=integer
specifies the number of most frequent items to compute.
| Default | 5 |
|---|
nominalVarsIndices=list(integer-1 <, integer-2, ...>)
specifies the indices of the variables to treat as nominal variables.
outputTableOptions=list(outputTableOptions)
specifies options for result tables. You can specify which result tables the server returns and how group-by results are handled.
| Alias | outputTableOpts |
|---|
The outputTableOptions value can be one or more of the following:
forceTableReturn=TRUE | FALSE
when set to True, result tables are returned to the client even if the output is also saved as an output table.
| Default | FALSE |
|---|
tableNames=list("string-1" <, "string-2", ...>)
specifies the names of result tables to generate. By default, all result tables are returned.
| Alias | outputTables |
|---|
percentileDefinition=integer
specifies the percentile definition to use. The definitions are numbered 1 to 6. The default value is 6.
| Alias | pctlDef |
|---|---|
| Default | 6 |
| Range | 1–6 |
percentileMaxIterations=integer
specifies the maximum number of iterations for percentile computation.
| Alias | pctlMaxIters |
|---|
percentileTolerance=double
specifies the tolerance for percentile computation.
| Alias | pctlEpsilon |
|---|---|
| Default | 1E-05 |
quantileSketch=list(quantileSketchOptions)
specifies the options for quantile sketch.
| Alias | quantileSketchOptions |
|---|
The quantileSketchOptions value can be one or more of the following:
compressionFactor=double
specifies the compression factor to use for quantile sketch.
| Default | 10 |
|---|---|
| Minimum value | 1 |
epsilon=double
specifies the tolerance to use for quantile sketch.
| Default | 0.001 |
|---|---|
| Range | 1E-06–0.1 |
requestPackages=list( list(rustatsRequestPackage-1) <, list(rustatsRequestPackage-2), ...>)
specifies an array of robust univariate statistics request packages to be processed by the action.
| Alias | reqPacks |
|---|
The rustatsRequestPackage value can be one or more of the following:
aadLocationUseMean=TRUE | FALSE
when set to True, the mean is used, instead of the median, as the center for the absolute average deviation (AAD) scale estimator.
| Alias | aadLocUseMean |
|---|---|
| Default | TRUE |
allCentralizedMoments=TRUE | FALSE
when set to True, computes all centralized moments (up to the sixth order).
| Default | FALSE |
|---|
allKurtoses=TRUE | FALSE
when set to True, computes all kurtosis estimates.
| Default | FALSE |
|---|
allLocations=TRUE | FALSE
when set to True, computes all location estimates.
| Default | FALSE |
|---|
allNominalStats=TRUE | FALSE
computes all nominal statistics. These include indices of qualitative variation and most and least frequent values.
| Default | FALSE |
|---|
allScales=TRUE | FALSE
when set to True, computes all scale estimates.
| Default | FALSE |
|---|
allSkewnesses=TRUE | FALSE
when set to True, computes all skewness estimates.
| Default | FALSE |
|---|
alpha=double
specifies the significance level for the confidence interval of location estimates.
| Default | 0.05 |
|---|
kurtoses=list("AVGQUANTILE", "CLASSICAL", "MOORS", "QUANTILE") | list("string-1" <, "string-2", ...>)
specifies a list of the kurtosis estimate types to compute.
| AVGQUANTILE | uses averaged quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical kurtosis estimate. |
| MOORS | uses Moor kurtosis estimate. |
| QUANTILE | uses quantile based kurtosis estimate. |
kurtosisAlphaQuantile=double
specifies the lower (alpha) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtAlphaQuantile |
|---|---|
| Default | 2.5 |
| Range | (0–50] |
kurtosisBetaQuantile=double
specifies the upper (beta) quantile for quantile or the average quantile kurtosis estimator.
| Alias | kurtBetaQuantile |
|---|---|
| Default | 25 |
| Range | (0–50] |
locationBiweightTuning=double
specifies the tuning factor for the Tukey biweight location estimator.
| Alias | locBiweightTuning |
|---|---|
| Default | 6 |
| Minimum value (exclusive) | 0 |
locationLowerPercentile=double
specifies the lower percentile for location estimates that need a percentile.
| Alias | locLowerPerc |
|---|---|
| Default | 5 |
| Range | (0, 50) |
locations=list("BIWEIGHT", "GEOMETRICMEAN", "HARMONICMEAN", "MEAN", "MEDIAN", "TRIMMEDMEAN", "WINSORIZEDMEAN") | list("string-1" <, "string-2", ...>)
specifies a list of the location estimate types to compute.
| BIWEIGHT | uses Tukey biweight based estimate for location. |
|---|---|
| GEOMETRICMEAN | uses the geometric mean for location. |
| HARMONICMEAN | uses the harmonic mean for location. |
| MEAN | uses the arithmetic mean for location. |
| MEDIAN | uses the median value for location. |
| TRIMMEDMEAN | uses the trimmed mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
| WINSORIZEDMEAN | uses the Winsorized mean for location. You can specify bounds with the lowerPerc, upperPerc, or symPerc parameters. |
locationSymmetricPercentile=double
specifies the symmetric percentile for the location estimates that need percentiles. For example, a value of 20 indicates to set a lower percentile to 10 and the upper percentile to 90.
| Alias | locSymPerc |
|---|---|
| Default | 10 |
| Range | (0, 100) |
locationUpperPercentile=double
specifies the upper percentile for location estimates that need a percentile.
| Alias | locUpperPerc |
|---|---|
| Default | 95 |
| Range | (50, 100) |
maxOrder=integer
specifies the maximum order of the centralized moment to compute.
| Alias | nMoments |
|---|---|
| Default | 4 |
| Maximum value | 6 |
nPercentiles=integer
computes the equal frequency percentiles. For example, setting the value to 3 asks for the computation of 25, 50, and 75 percentiles.
| Default | 0 |
|---|---|
| Minimum value | 0 |
percentiles=list(double-1 <, double-2, ...>)
specifies a list of quantiles or percentiles to compute.
scaleBiweightTuning=double
specifies the tuning factor for the Tukey biweight scale estimator.
| Alias | sclBiweightTuning |
|---|---|
| Default | 9 |
| Minimum value (exclusive) | 0 |
scales=list("AAD", "BIWEIGHT", "GINI", "IQR", "MAD", "STD") | list("string-1" <, "string-2", ...>)
specifies a list of the scale estimate types to compute.
| AAD | uses the absolute deviation about the mean or median as scale. |
|---|---|
| BIWEIGHT | uses Tukey biweight based estimate for scale. |
| GINI | uses the Gini scale for scale. |
| IQR | uses the inter-quartile range for scale. |
| MAD | uses the median absolute deviation about the median for scale. |
| STD | uses the standard deviation for scale. |
skewnesses=list("AVGQUANTILE", "CLASSICAL", "PEARMED", "QUANTILE") | list("string-1" <, "string-2", ...>)
specifies a list of the skewness estimate types to compute.
| AVGQUANTILE | uses average quantile based skewness estimate. |
|---|---|
| CLASSICAL | uses the classical skewness coefficient for skewness. |
| PEARMED | uses Pearson median skewness coefficient for skewness. |
| QUANTILE | uses quantile based skewness estimate. |
skewnessQuantile=double
specifies the quantile for quantile skewness estimator.
| Alias | skewQuantile |
|---|---|
| Default | 25 |
| Range | (0, 50) |
standardizeMoments=TRUE | FALSE
specifies that standardized moments be generated, instead of the centralized moments.
| Alias | stdMoments |
|---|---|
| Default | FALSE |
seed=integer
specifies a seed value. The seed is used to generate random values.
| Default | 0 |
|---|
standardError=TRUE | FALSE
when set to True, standard errors for location are included in the results.
| Alias | stdErr |
|---|---|
| Default | FALSE |
* table=list(castable)
specifies the table name, caslib, and other common parameters.
For more information about specifying the table parameter, see the common castable parameter.
tolerance=double
specifies the tolerance for the iterative robust univariate statistics.
| Default | 1E-05 |
|---|
unbiased=TRUE | FALSE
when set to True, the standard deviation, variance, classical skewness and classical kurtosis are corrected for bias.
| Default | TRUE |
|---|
varsToArgumentsMap=list(integer-1 <, integer-2, ...>)
specifies which request packages to compute for each variable. If a value is specified for the nArgsForEachVar parameter, then you must set this. Otherwise, both parameters are ignored and all request packages are computed for all variables.
weight="variable-name"
specifies the weight variable.