Robust PCA Action Set: Syntax
Provides actions for robust principal component analysis (RPCA) and moving windows principal component analysis (MWPCA)
robustpca Action
Performs robust principal component analysis.
CASL Syntax
Parameter Descriptions
attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}
changes the attributes of variables that are used in this action.
For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
center=TRUE | FALSE
when set to True, centers the numeric variables by the mean of each column.
| Alias | centering |
| Default | FALSE |
code={rpcaCodegen}
produces SAS score code.
casOut={casouttable}
specifies the settings for an output table.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
comment=TRUE | FALSE
when set to True, adds comments to the DATA step code.
| Default | FALSE |
fmtWdth=integer
specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.
| Alias | fmtWidth |
| Default | 20 |
| Range | 0–32 |
indentSize=integer
specifies the number of spaces to indent the DATA step code for each indent level.
| Default | 3 |
| Range | 0–10 |
labelId=integer
specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.
lineSize=integer
specifies the line size for the generated code.
| Default | 120 |
| Range | 64–254 |
noTrim=TRUE | FALSE
requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.
| Default | FALSE |
projectionType="LRS" | "PCA"
tabForm=TRUE | FALSE
when set to True, generates the code in way that is appropriate for storing in a table.
| Alias | tableForm |
| Default | FALSE |
colStatistics={casouttable}
specifies the name of the output table to contain simple statistics for the variables of the input data set.
For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).
cumEigPctTol=double
specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.
| Default | 1 |
| Range | (0–1] |
decomp="NONE" | "PCA" | "SVD"
specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.
| Default | NONE |
display={displayTables}
specifies a list of results tables to send to the client for display.
For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).
fixedMu=TRUE | FALSE
when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.
| Default | FALSE |
id={"variable-name-1" <, "variable-name-2", ...>}
specifies the variables to use as record identifiers.
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.
For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
lambda=double
specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.
| Default | -1 |
| Range | (0–10000000000] |
lambdaWeight=double
specifies the weight of lambda.
| Default | 1 |
| Range | (0–10000000000] |
maxIter=integer
specifies the maximum number of iterations for robust principal component analysis algorithms.
| Default | 1000 |
| Minimum value | 0 |
method="ALM" | "APG"
mu=double
specifies an initial value of mu in the objective function for the accelerated proximal gradient method.
| Default | 0.001 |
| Range | 0–10000000000 |
nThreads=integer
specifies the maximum number of threads to use on each computation node.
| Default | 16 |
| Minimum value | 1 |
outMat={outRpcaTabs}
specifies a list of parameters for the output tables of the robust principal component analysis method.
errMat={casouttable}
specifies the name of the output table for the error matrix.
| Alias | outError |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
lowRankMat={casouttable}
specifies the name of the output table for the low-rank matrix.
| Alias | outLowRank |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
sparseMat={casouttable}
specifies the name of the output table for the sparse matrix.
| Alias | outSparse |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outPca={outPcaTabs}
specifies a list of parameters for the output tables of the principal component analysis.
pcLoadings={casouttable}
specifies the name of the output table for the principal component loadings.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
pcScores={casouttable}
specifies the name of the output table for the principal component scores.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outSvd={outSvdTabs}
specifies a list of parameters for the output tables of the singular value decomposition.
svdDiag={casouttable}
specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
svdLeft={casouttable}
specifies the name of the output table for the left-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
svdRight={casouttable}
specifies the name of the output table for the right-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outputTables={outputTables}
lists the names of results tables to save as CAS tables on the server.
For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).
pcPrefix="string"
specifies a prefix for naming the principal components.
| Default | "Prin" |
saveState={casouttable}
specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.
For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).
scale=TRUE | FALSE
when set to True, scales the numeric variables by the standard deviation of each column.
| Alias | scaling |
| Default | FALSE |
svdMaxRank=integer
specifies the maximum value for rank to be considered in the singular value decomposition solver.
| Default | 0 |
| Minimum value | 1 |
svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"
svdRand={randomizedSvd}
specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.
power=integer
specifies the parameter power.
| Default | 0 |
| Minimum value | 0 |
randSeed=integer
specifies the seed value.
| Default | 0 |
| Minimum value | 1 |
* table={castable}
specifies the settings for an input table.
| Long form | table={name="table-name"} |
| Shortcut form | table="table-name" |
| The castable value can be one or more of the following: |
caslib="string"
specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
| Default | FALSE |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.
| Alias | compVars |
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Alias | options, dataSource |
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).
* name="table-name"
specifies the name of the table to use.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
tolerance=double
specifies the convergence criterion for the robust principal component analysis algorithms.
| Alias | stopcriterion |
| Default | 1e-07 |
| Minimum value | 1e-10 |
robustpca Action
Performs robust principal component analysis.
Lua Syntax
Parameter Descriptions
attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}
changes the attributes of variables that are used in this action.
For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
center=true | false
when set to True, centers the numeric variables by the mean of each column.
| Alias | centering |
| Default | false |
code={rpcaCodegen}
produces SAS score code.
casOut={casouttable}
specifies the settings for an output table.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
comment=true | false
when set to True, adds comments to the DATA step code.
| Default | false |
fmtWdth=integer
specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.
| Alias | fmtWidth |
| Default | 20 |
| Range | 0–32 |
indentSize=integer
specifies the number of spaces to indent the DATA step code for each indent level.
| Default | 3 |
| Range | 0–10 |
labelId=integer
specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.
lineSize=integer
specifies the line size for the generated code.
| Default | 120 |
| Range | 64–254 |
noTrim=true | false
requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.
| Default | false |
projectionType="LRS" | "PCA"
tabForm=true | false
when set to True, generates the code in way that is appropriate for storing in a table.
| Alias | tableForm |
| Default | false |
colStatistics={casouttable}
specifies the name of the output table to contain simple statistics for the variables of the input data set.
For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).
cumEigPctTol=double
specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.
| Default | 1 |
| Range | (0–1] |
decomp="NONE" | "PCA" | "SVD"
specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.
| Default | NONE |
display={displayTables}
specifies a list of results tables to send to the client for display.
For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).
fixedMu=true | false
when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.
| Default | false |
id={"variable-name-1" <, "variable-name-2", ...>}
specifies the variables to use as record identifiers.
inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.
For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
lambda=double
specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.
| Default | -1 |
| Range | (0–10000000000] |
lambdaWeight=double
specifies the weight of lambda.
| Default | 1 |
| Range | (0–10000000000] |
maxIter=integer
specifies the maximum number of iterations for robust principal component analysis algorithms.
| Default | 1000 |
| Minimum value | 0 |
method="ALM" | "APG"
mu=double
specifies an initial value of mu in the objective function for the accelerated proximal gradient method.
| Default | 0.001 |
| Range | 0–10000000000 |
nThreads=integer
specifies the maximum number of threads to use on each computation node.
| Default | 16 |
| Minimum value | 1 |
outMat={outRpcaTabs}
specifies a list of parameters for the output tables of the robust principal component analysis method.
errMat={casouttable}
specifies the name of the output table for the error matrix.
| Alias | outError |
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
lowRankMat={casouttable}
specifies the name of the output table for the low-rank matrix.
| Alias | outLowRank |
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
sparseMat={casouttable}
specifies the name of the output table for the sparse matrix.
| Alias | outSparse |
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outPca={outPcaTabs}
specifies a list of parameters for the output tables of the principal component analysis.
pcLoadings={casouttable}
specifies the name of the output table for the principal component loadings.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
pcScores={casouttable}
specifies the name of the output table for the principal component scores.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outSvd={outSvdTabs}
specifies a list of parameters for the output tables of the singular value decomposition.
svdDiag={casouttable}
specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
svdLeft={casouttable}
specifies the name of the output table for the left-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
svdRight={casouttable}
specifies the name of the output table for the right-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=true | false
when set to True, data compression is applied to the table.
| Default | false |
indexVars={"variable-name-1" <, "variable-name-2", ...>}
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=true | false
This parameter is deprecated.
| Default | true |
promote=true | false
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | false |
replace=true | false
when set to True, overwrites an existing table with the same name.
| Default | false |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where={"string-1" <, "string-2", ...>}
specifies an expression for subsetting the output data.
outputTables={outputTables}
lists the names of results tables to save as CAS tables on the server.
For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).
pcPrefix="string"
specifies a prefix for naming the principal components.
| Default | "Prin" |
saveState={casouttable}
specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.
For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).
scale=true | false
when set to True, scales the numeric variables by the standard deviation of each column.
| Alias | scaling |
| Default | false |
svdMaxRank=integer
specifies the maximum value for rank to be considered in the singular value decomposition solver.
| Default | 0 |
| Minimum value | 1 |
svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"
svdRand={randomizedSvd}
specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.
power=integer
specifies the parameter power.
| Default | 0 |
| Minimum value | 0 |
randSeed=integer
specifies the seed value.
| Default | 0 |
| Minimum value | 1 |
* table={castable}
specifies the settings for an input table.
| Long form | table={name="table-name"} |
| Shortcut form | table="table-name" |
| The castable value can be one or more of the following: |
caslib="string"
specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=true | false
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
| Default | false |
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.
| Alias | compVars |
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}
specifies data source options.
| Alias | options, dataSource |
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import |
The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).
* name="table-name"
specifies the name of the table to use.
onDemand=true | false
This parameter is deprecated.
| Default | true |
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
singlePass=true | false
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | false |
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}
specifies the variables to use in the action.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
tolerance=double
specifies the convergence criterion for the robust principal component analysis algorithms.
| Alias | stopcriterion |
| Default | 1e-07 |
| Minimum value | 1e-10 |
robustpca Action
Performs robust principal component analysis.
Python Syntax
Parameter Descriptions
attributes=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
changes the attributes of variables that are used in this action.
For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
center=True | False
when set to True, centers the numeric variables by the mean of each column.
| Alias | centering |
| Default | False |
code={rpcaCodegen}
produces SAS score code.
"casOut":{casouttable}
specifies the settings for an output table.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"comment":True | False
when set to True, adds comments to the DATA step code.
| Default | False |
"fmtWdth":integer
specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.
| Alias | fmtWidth |
| Default | 20 |
| Range | 0–32 |
"indentSize":integer
specifies the number of spaces to indent the DATA step code for each indent level.
| Default | 3 |
| Range | 0–10 |
"labelId":integer
specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.
"lineSize":integer
specifies the line size for the generated code.
| Default | 120 |
| Range | 64–254 |
"noTrim":True | False
requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.
| Default | False |
"projectionType":"LRS" | "PCA"
"tabForm":True | False
when set to True, generates the code in way that is appropriate for storing in a table.
| Alias | tableForm |
| Default | False |
colStatistics={casouttable}
specifies the name of the output table to contain simple statistics for the variables of the input data set.
For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).
cumEigPctTol=double
specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.
| Default | 1 |
| Range | (0–1] |
decomp="NONE" | "PCA" | "SVD"
specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.
| Default | NONE |
display={displayTables}
specifies a list of results tables to send to the client for display.
For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).
fixedMu=True | False
when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.
| Default | False |
id=["variable-name-1" <, "variable-name-2", ...>]
specifies the variables to use as record identifiers.
inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.
For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
lambda_=double
specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.
| Default | -1 |
| Range | (0–10000000000] |
lambdaWeight=double
specifies the weight of lambda.
| Default | 1 |
| Range | (0–10000000000] |
maxIter=integer
specifies the maximum number of iterations for robust principal component analysis algorithms.
| Default | 1000 |
| Minimum value | 0 |
method="ALM" | "APG"
mu=double
specifies an initial value of mu in the objective function for the accelerated proximal gradient method.
| Default | 0.001 |
| Range | 0–10000000000 |
nThreads=integer
specifies the maximum number of threads to use on each computation node.
| Default | 16 |
| Minimum value | 1 |
outMat={outRpcaTabs}
specifies a list of parameters for the output tables of the robust principal component analysis method.
"errMat":{casouttable}
specifies the name of the output table for the error matrix.
| Alias | outError |
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"lowRankMat":{casouttable}
specifies the name of the output table for the low-rank matrix.
| Alias | outLowRank |
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"sparseMat":{casouttable}
specifies the name of the output table for the sparse matrix.
| Alias | outSparse |
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
outPca={outPcaTabs}
specifies a list of parameters for the output tables of the principal component analysis.
"pcLoadings":{casouttable}
specifies the name of the output table for the principal component loadings.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"pcScores":{casouttable}
specifies the name of the output table for the principal component scores.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
outSvd={outSvdTabs}
specifies a list of parameters for the output tables of the singular value decomposition.
"svdDiag":{casouttable}
specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"svdLeft":{casouttable}
specifies the name of the output table for the left-singular vectors.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
"svdRight":{casouttable}
specifies the name of the output table for the right-singular vectors.
"caslib":"string"
specifies the name of the caslib to use.
"compress":True | False
when set to True, data compression is applied to the table.
| Default | False |
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
specifies the list of variables to create indexes for in the output data
"label":"string"
specifies the descriptive label to associate with the table.
"maxMemSize":64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
"name":"table-name"
specifies the name to associate with the table.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"promote":True | False
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | False |
"replace":True | False
when set to True, overwrites an existing table with the same name.
| Default | False |
"replication":integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
"timeStamp":"string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
"where":["string-1" <, "string-2", ...>]
specifies an expression for subsetting the output data.
outputTables={outputTables}
lists the names of results tables to save as CAS tables on the server.
For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).
pcPrefix="string"
specifies a prefix for naming the principal components.
| Default | "Prin" |
saveState={casouttable}
specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.
For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).
scale=True | False
when set to True, scales the numeric variables by the standard deviation of each column.
| Alias | scaling |
| Default | False |
svdMaxRank=integer
specifies the maximum value for rank to be considered in the singular value decomposition solver.
| Default | 0 |
| Minimum value | 1 |
svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"
svdRand={randomizedSvd}
specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.
"power":integer
specifies the parameter power.
| Default | 0 |
| Minimum value | 0 |
"randSeed":integer
specifies the seed value.
| Default | 0 |
| Minimum value | 1 |
* table={castable}
specifies the settings for an input table.
| Long form | table={"name":"table-name"} |
| Shortcut form | table="table-name" |
| The castable value can be one or more of the following: |
"caslib":"string"
specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
"computedOnDemand":True | False
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
| Default | False |
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.
| Alias | compVars |
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"computedVarsProgram":"string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}
specifies data source options.
| Alias | options, dataSource |
"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}
specifies the settings for reading a table from a data source.
| Alias | import_ |
The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).
* "name":"table-name"
specifies the name of the table to use.
"onDemand":True | False
This parameter is deprecated.
| Default | True |
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"singlePass":True | False
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | False |
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]
specifies the variables to use in the action.
"format":"string"
specifies the format to apply to the variable.
"formattedLength":integer
specifies the length of format field plus the format precision.
"label":"string"
specifies the descriptive label for the variable.
* "name":"variable-name"
specifies the name for the variable.
"nfd":integer
specifies the length of the format precision.
"nfl":integer
specifies the length of the format field.
"where":"where-expression"
specifies an expression for subsetting the input data.
tolerance=double
specifies the convergence criterion for the robust principal component analysis algorithms.
| Alias | stopcriterion |
| Default | 1e-07 |
| Minimum value | 1e-10 |
robustpca Action
Performs robust principal component analysis.
R Syntax
Parameter Descriptions
attributes=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
changes the attributes of variables that are used in this action.
For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
center=TRUE | FALSE
when set to True, centers the numeric variables by the mean of each column.
| Alias | centering |
| Default | FALSE |
code=list(rpcaCodegen)
produces SAS score code.
casOut=list(casouttable)
specifies the settings for an output table.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
comment=TRUE | FALSE
when set to True, adds comments to the DATA step code.
| Default | FALSE |
fmtWdth=integer
specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.
| Alias | fmtWidth |
| Default | 20 |
| Range | 0–32 |
indentSize=integer
specifies the number of spaces to indent the DATA step code for each indent level.
| Default | 3 |
| Range | 0–10 |
labelId=integer
specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.
lineSize=integer
specifies the line size for the generated code.
| Default | 120 |
| Range | 64–254 |
noTrim=TRUE | FALSE
requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.
| Default | FALSE |
projectionType="LRS" | "PCA"
tabForm=TRUE | FALSE
when set to True, generates the code in way that is appropriate for storing in a table.
| Alias | tableForm |
| Default | FALSE |
colStatistics=list(casouttable)
specifies the name of the output table to contain simple statistics for the variables of the input data set.
For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).
cumEigPctTol=double
specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.
| Default | 1 |
| Range | (0–1] |
decomp="NONE" | "PCA" | "SVD"
specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.
| Default | NONE |
display=list(displayTables)
specifies a list of results tables to send to the client for display.
For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).
fixedMu=TRUE | FALSE
when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.
| Default | FALSE |
id=list("variable-name-1" <, "variable-name-2", ...>)
specifies the variables to use as record identifiers.
inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.
For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).
lambda=double
specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.
| Default | -1 |
| Range | (0–10000000000] |
lambdaWeight=double
specifies the weight of lambda.
| Default | 1 |
| Range | (0–10000000000] |
maxIter=integer
specifies the maximum number of iterations for robust principal component analysis algorithms.
| Default | 1000 |
| Minimum value | 0 |
method="ALM" | "APG"
mu=double
specifies an initial value of mu in the objective function for the accelerated proximal gradient method.
| Default | 0.001 |
| Range | 0–10000000000 |
nThreads=integer
specifies the maximum number of threads to use on each computation node.
| Default | 16 |
| Minimum value | 1 |
outMat=list(outRpcaTabs)
specifies a list of parameters for the output tables of the robust principal component analysis method.
errMat=list(casouttable)
specifies the name of the output table for the error matrix.
| Alias | outError |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
lowRankMat=list(casouttable)
specifies the name of the output table for the low-rank matrix.
| Alias | outLowRank |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
sparseMat=list(casouttable)
specifies the name of the output table for the sparse matrix.
| Alias | outSparse |
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
outPca=list(outPcaTabs)
specifies a list of parameters for the output tables of the principal component analysis.
pcLoadings=list(casouttable)
specifies the name of the output table for the principal component loadings.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
pcScores=list(casouttable)
specifies the name of the output table for the principal component scores.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
outSvd=list(outSvdTabs)
specifies a list of parameters for the output tables of the singular value decomposition.
svdDiag=list(casouttable)
specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
svdLeft=list(casouttable)
specifies the name of the output table for the left-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
svdRight=list(casouttable)
specifies the name of the output table for the right-singular vectors.
caslib="string"
specifies the name of the caslib to use.
compress=TRUE | FALSE
when set to True, data compression is applied to the table.
| Default | FALSE |
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
specifies the list of variables to create indexes for in the output data
label="string"
specifies the descriptive label to associate with the table.
maxMemSize=64-bit-integer
specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.
| TIP | You can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes. |
name="table-name"
specifies the name to associate with the table.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
promote=TRUE | FALSE
when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.
| Default | FALSE |
replace=TRUE | FALSE
when set to True, overwrites an existing table with the same name.
| Default | FALSE |
replication=integer
specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.
| Default | 1 |
| Minimum value | 0 |
timeStamp="string"
specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.
where=list("string-1" <, "string-2", ...>)
specifies an expression for subsetting the output data.
outputTables=list(outputTables)
lists the names of results tables to save as CAS tables on the server.
For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).
pcPrefix="string"
specifies a prefix for naming the principal components.
| Default | "Prin" |
saveState=list(casouttable)
specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.
For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).
scale=TRUE | FALSE
when set to True, scales the numeric variables by the standard deviation of each column.
| Alias | scaling |
| Default | FALSE |
svdMaxRank=integer
specifies the maximum value for rank to be considered in the singular value decomposition solver.
| Default | 0 |
| Minimum value | 1 |
svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"
svdRand=list(randomizedSvd)
specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.
power=integer
specifies the parameter power.
| Default | 0 |
| Minimum value | 0 |
randSeed=integer
specifies the seed value.
| Default | 0 |
| Minimum value | 1 |
* table=list(castable)
specifies the settings for an input table.
| Long form | table=list(name="table-name") |
| Shortcut form | table="table-name" |
| The castable value can be one or more of the following: |
caslib="string"
specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.
computedOnDemand=TRUE | FALSE
when set to True, creates the computed variables when the table is loaded instead of when the action begins.
| Alias | compOnDemand |
| Default | FALSE |
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.
| Alias | compVars |
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
computedVarsProgram="string"
specifies an expression for each computed variable that you include in the computedVars parameter.
| Alias | compPgm |
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)
specifies data source options.
| Alias | options, dataSource |
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)
specifies the settings for reading a table from a data source.
| Alias | import |
The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).
* name="table-name"
specifies the name of the table to use.
onDemand=TRUE | FALSE
This parameter is deprecated.
| Default | TRUE |
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
singlePass=TRUE | FALSE
when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.
| Default | FALSE |
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)
specifies the variables to use in the action.
format="string"
specifies the format to apply to the variable.
formattedLength=integer
specifies the length of format field plus the format precision.
label="string"
specifies the descriptive label for the variable.
* name="variable-name"
specifies the name for the variable.
nfd=integer
specifies the length of the format precision.
nfl=integer
specifies the length of the format field.
where="where-expression"
specifies an expression for subsetting the input data.
tolerance=double
specifies the convergence criterion for the robust principal component analysis algorithms.
| Alias | stopcriterion |
| Default | 1e-07 |
| Minimum value | 1e-10 |