Robust PCA Action Set: Syntax

Provides actions for robust principal component analysis (RPCA) and moving windows principal component analysis (MWPCA)

robustpca Action

Performs robust principal component analysis.

CASL Syntax

robustPca.robustpca <result=results> <status=rc> /
attributes={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
center=TRUE | FALSE
code={
casOut={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=TRUE | FALSE,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
tabForm=TRUE | FALSE
}
colStatistics={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
display={
caseSensitive=TRUE | FALSE,
exclude=TRUE | FALSE,
excludeAll=TRUE | FALSE,
keyIsPath=TRUE | FALSE,
names={"string-1" <, "string-2", ...>},
pathType="LABEL" | "NAME",
traceNames=TRUE | FALSE
}
fixedMu=TRUE | FALSE
id={"variable-name-1" <, "variable-name-2", ...>}
inputs={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
lambda=double
maxIter=integer
mu=double
nThreads=integer
outMat={
errMat={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
lowRankMat={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
sparseMat={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
}
outPca={
pcLoadings={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
pcScores={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
}
outSvd={
svdDiag={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
svdLeft={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
svdRight={
caslib="string"
compress=TRUE | FALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
}
outputTables={
groupByVarsRaw=TRUE | FALSE,
includeAll=TRUE | FALSE,
names={"string-1" <, "string-2", ...>} | {key-1={casouttable-1} <, key-2={casouttable-2}, ...>},
repeated=TRUE | FALSE,
replace=TRUE | FALSE
}
pcPrefix="string"
saveState={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
scale=TRUE | FALSE
svdMaxRank=integer
svdRand={
power=integer,
randSeed=integer
}
required parameter table={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
}
tolerance=double
;

Parameter Descriptions

attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

center=TRUE | FALSE

when set to True, centers the numeric variables by the mean of each column.

Aliascentering
DefaultFALSE

code={rpcaCodegen}

produces SAS score code.

casOut={casouttable}

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

comment=TRUE | FALSE

when set to True, adds comments to the DATA step code.

DefaultFALSE
fmtWdth=integer

specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.

AliasfmtWidth
Default20
Range0–32
indentSize=integer

specifies the number of spaces to indent the DATA step code for each indent level.

Default3
Range0–10
labelId=integer

specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.

lineSize=integer

specifies the line size for the generated code.

Default120
Range64–254
noTrim=TRUE | FALSE

requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.

DefaultFALSE
projectionType="LRS" | "PCA"

specifies the type of scoring.

DefaultPCA
LRS

projects the scoring observations in the low-rank space.

PCA

projects the scoring observations onto the principal components.

tabForm=TRUE | FALSE

when set to True, generates the code in way that is appropriate for storing in a table.

AliastableForm
DefaultFALSE

colStatistics={casouttable}

specifies the name of the output table to contain simple statistics for the variables of the input data set.

For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

cumEigPctTol=double

specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.

Default1
Range(0–1]

decomp="NONE" | "PCA" | "SVD"

specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.

DefaultNONE
NONE

performs neither analysis.

PCA

performs principal component analysis.

SVD

performs singular value decomposition.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

fixedMu=TRUE | FALSE

when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.

DefaultFALSE

id={"variable-name-1" <, "variable-name-2", ...>}

specifies the variables to use as record identifiers.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

lambda=double

specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.

Default-1
Range(0–10000000000]

lambdaWeight=double

specifies the weight of lambda.

Default1
Range(0–10000000000]

maxIter=integer

specifies the maximum number of iterations for robust principal component analysis algorithms.

Default1000
Minimum value0

method="ALM" | "APG"

specifies the method to use to perform the robust principal component analysis.

DefaultALM
ALM

uses the augmented Lagrange multiplier method.

APG

uses the accelerated proximal gradient method.

mu=double

specifies an initial value of mu in the objective function for the accelerated proximal gradient method.

Default0.001
Range0–10000000000

nThreads=integer

specifies the maximum number of threads to use on each computation node.

Default16
Minimum value1

outMat={outRpcaTabs}

specifies a list of parameters for the output tables of the robust principal component analysis method.

errMat={casouttable}

specifies the name of the output table for the error matrix.

AliasoutError
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

lowRankMat={casouttable}

specifies the name of the output table for the low-rank matrix.

AliasoutLowRank
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

sparseMat={casouttable}

specifies the name of the output table for the sparse matrix.

AliasoutSparse
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outPca={outPcaTabs}

specifies a list of parameters for the output tables of the principal component analysis.

pcLoadings={casouttable}

specifies the name of the output table for the principal component loadings.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

pcScores={casouttable}

specifies the name of the output table for the principal component scores.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outSvd={outSvdTabs}

specifies a list of parameters for the output tables of the singular value decomposition.

svdDiag={casouttable}

specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

svdLeft={casouttable}

specifies the name of the output table for the left-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

svdRight={casouttable}

specifies the name of the output table for the right-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

pcPrefix="string"

specifies a prefix for naming the principal components.

Default"Prin"

saveState={casouttable}

specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.

For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scale=TRUE | FALSE

when set to True, scales the numeric variables by the standard deviation of each column.

Aliasscaling
DefaultFALSE

svdMaxRank=integer

specifies the maximum value for rank to be considered in the singular value decomposition solver.

Default0
Minimum value1

svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"

specifies the type of the singular value decomposition solver.

DefaultEIGEN
EIGEN

uses the eigenvalue decomposition method.

ITERATIVE

uses the iterative singular value decomposition method.

RANDOM

uses the randomized singular value decomposition method.

svdRand={randomizedSvd}

specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.

power=integer

specifies the parameter power.

Default0
Minimum value0
randSeed=integer

specifies the seed value.

Default0
Minimum value1

* table={castable}

specifies the settings for an input table.

Long formtable={name="table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the convergence criterion for the robust principal component analysis algorithms.

Aliasstopcriterion
Default1e-07
Minimum value1e-10

robustpca Action

Performs robust principal component analysis.

Lua Syntax

results, info = s:robustPca_robustpca{
attributes={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
center=true | false,
code={
casOut={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
comment=true | false,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=true | false,
tabForm=true | false
},
colStatistics={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
cumEigPctTol=double,
display={
caseSensitive=true | false,
exclude=true | false,
excludeAll=true | false,
keyIsPath=true | false,
names={"string-1" <, "string-2", ...>},
pathType="LABEL" | "NAME",
traceNames=true | false
},
fixedMu=true | false,
id={"variable-name-1" <, "variable-name-2", ...>},
inputs={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
lambda=double,
lambdaWeight=double,
maxIter=integer,
mu=double,
nThreads=integer,
outMat={
errMat={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
lowRankMat={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
sparseMat={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
},
outPca={
pcLoadings={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
pcScores={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
},
outSvd={
svdDiag={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
svdLeft={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
},
svdRight={
caslib="string"
compress=true | false
indexVars={"variable-name-1" <, "variable-name-2", ...>}
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=true | false
promote=true | false
replace=true | false
replication=integer
timeStamp="string"
where={"string-1" <, "string-2", ...>}
}
},
outputTables={
groupByVarsRaw=true | false,
includeAll=true | false,
names={"string-1" <, "string-2", ...>} | {key-1={casouttable-1} <, key-2={casouttable-2}, ...>},
repeated=true | false,
replace=true | false
},
pcPrefix="string",
saveState={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
scale=true | false,
svdMaxRank=integer,
svdRand={
power=integer,
randSeed=integer
},
required parameter table={
caslib="string",
computedOnDemand=true | false,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
},
tolerance=double
}

Parameter Descriptions

attributes={{casinvardesc-1} <, {casinvardesc-2}, ...>}

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

center=true | false

when set to True, centers the numeric variables by the mean of each column.

Aliascentering
Defaultfalse

code={rpcaCodegen}

produces SAS score code.

casOut={casouttable}

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

comment=true | false

when set to True, adds comments to the DATA step code.

Defaultfalse
fmtWdth=integer

specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.

AliasfmtWidth
Default20
Range0–32
indentSize=integer

specifies the number of spaces to indent the DATA step code for each indent level.

Default3
Range0–10
labelId=integer

specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.

lineSize=integer

specifies the line size for the generated code.

Default120
Range64–254
noTrim=true | false

requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.

Defaultfalse
projectionType="LRS" | "PCA"

specifies the type of scoring.

DefaultPCA
LRS

projects the scoring observations in the low-rank space.

PCA

projects the scoring observations onto the principal components.

tabForm=true | false

when set to True, generates the code in way that is appropriate for storing in a table.

AliastableForm
Defaultfalse

colStatistics={casouttable}

specifies the name of the output table to contain simple statistics for the variables of the input data set.

For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

cumEigPctTol=double

specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.

Default1
Range(0–1]

decomp="NONE" | "PCA" | "SVD"

specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.

DefaultNONE
NONE

performs neither analysis.

PCA

performs principal component analysis.

SVD

performs singular value decomposition.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

fixedMu=true | false

when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.

Defaultfalse

id={"variable-name-1" <, "variable-name-2", ...>}

specifies the variables to use as record identifiers.

inputs={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

lambda=double

specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.

Default-1
Range(0–10000000000]

lambdaWeight=double

specifies the weight of lambda.

Default1
Range(0–10000000000]

maxIter=integer

specifies the maximum number of iterations for robust principal component analysis algorithms.

Default1000
Minimum value0

method="ALM" | "APG"

specifies the method to use to perform the robust principal component analysis.

DefaultALM
ALM

uses the augmented Lagrange multiplier method.

APG

uses the accelerated proximal gradient method.

mu=double

specifies an initial value of mu in the objective function for the accelerated proximal gradient method.

Default0.001
Range0–10000000000

nThreads=integer

specifies the maximum number of threads to use on each computation node.

Default16
Minimum value1

outMat={outRpcaTabs}

specifies a list of parameters for the output tables of the robust principal component analysis method.

errMat={casouttable}

specifies the name of the output table for the error matrix.

AliasoutError
caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

lowRankMat={casouttable}

specifies the name of the output table for the low-rank matrix.

AliasoutLowRank
caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

sparseMat={casouttable}

specifies the name of the output table for the sparse matrix.

AliasoutSparse
caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outPca={outPcaTabs}

specifies a list of parameters for the output tables of the principal component analysis.

pcLoadings={casouttable}

specifies the name of the output table for the principal component loadings.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

pcScores={casouttable}

specifies the name of the output table for the principal component scores.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outSvd={outSvdTabs}

specifies a list of parameters for the output tables of the singular value decomposition.

svdDiag={casouttable}

specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

svdLeft={casouttable}

specifies the name of the output table for the left-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

svdRight={casouttable}

specifies the name of the output table for the right-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=true | false

when set to True, data compression is applied to the table.

Defaultfalse
indexVars={"variable-name-1" <, "variable-name-2", ...>}

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
promote=true | false

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

Defaultfalse
replace=true | false

when set to True, overwrites an existing table with the same name.

Defaultfalse
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where={"string-1" <, "string-2", ...>}

specifies an expression for subsetting the output data.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

pcPrefix="string"

specifies a prefix for naming the principal components.

Default"Prin"

saveState={casouttable}

specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.

For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scale=true | false

when set to True, scales the numeric variables by the standard deviation of each column.

Aliasscaling
Defaultfalse

svdMaxRank=integer

specifies the maximum value for rank to be considered in the singular value decomposition solver.

Default0
Minimum value1

svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"

specifies the type of the singular value decomposition solver.

DefaultEIGEN
EIGEN

uses the eigenvalue decomposition method.

ITERATIVE

uses the iterative singular value decomposition method.

RANDOM

uses the randomized singular value decomposition method.

svdRand={randomizedSvd}

specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.

power=integer

specifies the parameter power.

Default0
Minimum value0
randSeed=integer

specifies the seed value.

Default0
Minimum value1

* table={castable}

specifies the settings for an input table.

Long formtable={name="table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the convergence criterion for the robust principal component analysis algorithms.

Aliasstopcriterion
Default1e-07
Minimum value1e-10

robustpca Action

Performs robust principal component analysis.

Python Syntax

results= s.robustPca.robustpca(
attributes=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
center=True | False,
code={
"casOut":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"comment":True | False,
"fmtWdth":integer,
"indentSize":integer,
"labelId":integer,
"lineSize":integer,
"noTrim":True | False,
"tabForm":True | False
},
colStatistics={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
cumEigPctTol=double,
display={
"caseSensitive":True | False,
"exclude":True | False,
"excludeAll":True | False,
"keyIsPath":True | False,
"names":["string-1" <, "string-2", ...>],
"pathType":"LABEL" | "NAME",
"traceNames":True | False
},
fixedMu=True | False,
id=["variable-name-1" <, "variable-name-2", ...>],
inputs=[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
lambda_=double,
lambdaWeight=double,
maxIter=integer,
mu=double,
nThreads=integer,
outMat={
"errMat":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"lowRankMat":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"sparseMat":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
}
},
outPca={
"pcLoadings":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"pcScores":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
}
},
outSvd={
"svdDiag":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"svdLeft":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
},
"svdRight":{
"caslib":"string"
"compress":True | False
"indexVars":["variable-name-1" <, "variable-name-2", ...>]
"label":"string"
"maxMemSize":64-bit-integer
"name":"table-name"
"onDemand":True | False
"promote":True | False
"replace":True | False
"replication":integer
"timeStamp":"string"
"where":["string-1" <, "string-2", ...>]
}
},
outputTables={
"groupByVarsRaw":True | False,
"includeAll":True | False,
"names":["string-1" <, "string-2", ...>] | {"key-1":{casouttable-1} <, "key-2":{casouttable-2}, ...>},
"repeated":True | False,
"replace":True | False
},
pcPrefix="string",
saveState={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
scale=True | False,
svdMaxRank=integer,
svdRand={
"power":integer,
"randSeed":integer
},
required parameter table={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"importOptions":{"fileType":"AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression"
},
tolerance=double
)

Parameter Descriptions

attributes=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

center=True | False

when set to True, centers the numeric variables by the mean of each column.

Aliascentering
DefaultFalse

code={rpcaCodegen}

produces SAS score code.

"casOut":{casouttable}

specifies the settings for an output table.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"comment":True | False

when set to True, adds comments to the DATA step code.

DefaultFalse
"fmtWdth":integer

specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.

AliasfmtWidth
Default20
Range0–32
"indentSize":integer

specifies the number of spaces to indent the DATA step code for each indent level.

Default3
Range0–10
"labelId":integer

specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.

"lineSize":integer

specifies the line size for the generated code.

Default120
Range64–254
"noTrim":True | False

requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.

DefaultFalse
"projectionType":"LRS" | "PCA"

specifies the type of scoring.

DefaultPCA
LRS

projects the scoring observations in the low-rank space.

PCA

projects the scoring observations onto the principal components.

"tabForm":True | False

when set to True, generates the code in way that is appropriate for storing in a table.

AliastableForm
DefaultFalse

colStatistics={casouttable}

specifies the name of the output table to contain simple statistics for the variables of the input data set.

For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

cumEigPctTol=double

specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.

Default1
Range(0–1]

decomp="NONE" | "PCA" | "SVD"

specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.

DefaultNONE
NONE

performs neither analysis.

PCA

performs principal component analysis.

SVD

performs singular value decomposition.

display={displayTables}

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

fixedMu=True | False

when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.

DefaultFalse

id=["variable-name-1" <, "variable-name-2", ...>]

specifies the variables to use as record identifiers.

inputs=[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

lambda_=double

specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.

Default-1
Range(0–10000000000]

lambdaWeight=double

specifies the weight of lambda.

Default1
Range(0–10000000000]

maxIter=integer

specifies the maximum number of iterations for robust principal component analysis algorithms.

Default1000
Minimum value0

method="ALM" | "APG"

specifies the method to use to perform the robust principal component analysis.

DefaultALM
ALM

uses the augmented Lagrange multiplier method.

APG

uses the accelerated proximal gradient method.

mu=double

specifies an initial value of mu in the objective function for the accelerated proximal gradient method.

Default0.001
Range0–10000000000

nThreads=integer

specifies the maximum number of threads to use on each computation node.

Default16
Minimum value1

outMat={outRpcaTabs}

specifies a list of parameters for the output tables of the robust principal component analysis method.

"errMat":{casouttable}

specifies the name of the output table for the error matrix.

AliasoutError
"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"lowRankMat":{casouttable}

specifies the name of the output table for the low-rank matrix.

AliasoutLowRank
"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"sparseMat":{casouttable}

specifies the name of the output table for the sparse matrix.

AliasoutSparse
"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

outPca={outPcaTabs}

specifies a list of parameters for the output tables of the principal component analysis.

"pcLoadings":{casouttable}

specifies the name of the output table for the principal component loadings.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"pcScores":{casouttable}

specifies the name of the output table for the principal component scores.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

outSvd={outSvdTabs}

specifies a list of parameters for the output tables of the singular value decomposition.

"svdDiag":{casouttable}

specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"svdLeft":{casouttable}

specifies the name of the output table for the left-singular vectors.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

"svdRight":{casouttable}

specifies the name of the output table for the right-singular vectors.

"caslib":"string"

specifies the name of the caslib to use.

"compress":True | False

when set to True, data compression is applied to the table.

DefaultFalse
"indexVars":["variable-name-1" <, "variable-name-2", ...>]

specifies the list of variables to create indexes for in the output data

"label":"string"

specifies the descriptive label to associate with the table.

"maxMemSize":64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
"name":"table-name"

specifies the name to associate with the table.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"promote":True | False

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFalse
"replace":True | False

when set to True, overwrites an existing table with the same name.

DefaultFalse
"replication":integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
"timeStamp":"string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

"where":["string-1" <, "string-2", ...>]

specifies an expression for subsetting the output data.

outputTables={outputTables}

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

pcPrefix="string"

specifies a prefix for naming the principal components.

Default"Prin"

saveState={casouttable}

specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.

For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scale=True | False

when set to True, scales the numeric variables by the standard deviation of each column.

Aliasscaling
DefaultFalse

svdMaxRank=integer

specifies the maximum value for rank to be considered in the singular value decomposition solver.

Default0
Minimum value1

svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"

specifies the type of the singular value decomposition solver.

DefaultEIGEN
EIGEN

uses the eigenvalue decomposition method.

ITERATIVE

uses the iterative singular value decomposition method.

RANDOM

uses the randomized singular value decomposition method.

svdRand={randomizedSvd}

specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.

"power":integer

specifies the parameter power.

Default0
Minimum value0
"randSeed":integer

specifies the seed value.

Default0
Minimum value1

* table={castable}

specifies the settings for an input table.

Long formtable={"name":"table-name"}
Shortcut formtable="table-name"
The castable value can be one or more of the following:
"caslib":"string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasoptions, dataSource
"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the table to use.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"where":"where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the convergence criterion for the robust principal component analysis algorithms.

Aliasstopcriterion
Default1e-07
Minimum value1e-10

robustpca Action

Performs robust principal component analysis.

R Syntax

results <– cas.robustPca.robustpca(s,
attributes=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
center=TRUE | FALSE,
code=list(
casOut=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
comment=TRUE | FALSE,
fmtWdth=integer,
indentSize=integer,
labelId=integer,
lineSize=integer,
noTrim=TRUE | FALSE,
tabForm=TRUE | FALSE
),
colStatistics=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
cumEigPctTol=double,
display=list(
caseSensitive=TRUE | FALSE,
exclude=TRUE | FALSE,
excludeAll=TRUE | FALSE,
keyIsPath=TRUE | FALSE,
names=list("string-1" <, "string-2", ...>),
pathType="LABEL" | "NAME",
traceNames=TRUE | FALSE
),
fixedMu=TRUE | FALSE,
id=list("variable-name-1" <, "variable-name-2", ...>),
inputs=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
lambda=double,
lambdaWeight=double,
maxIter=integer,
mu=double,
nThreads=integer,
outMat=list(
errMat=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
lowRankMat=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
sparseMat=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
)
),
outPca=list(
pcLoadings=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
pcScores=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
)
),
outSvd=list(
svdDiag=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
svdLeft=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
),
svdRight=list(
caslib="string"
compress=TRUE | FALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)
label="string"
maxMemSize=64-bit-integer
name="table-name"
onDemand=TRUE | FALSE
promote=TRUE | FALSE
replace=TRUE | FALSE
replication=integer
timeStamp="string"
where=list("string-1" <, "string-2", ...>)
)
),
outputTables=list(
groupByVarsRaw=TRUE | FALSE,
includeAll=TRUE | FALSE,
names=list("string-1" <, "string-2", ...>) | list(key-1=list(casouttable-1) <, key-2=list(casouttable-2), ...>),
repeated=TRUE | FALSE,
replace=TRUE | FALSE
),
pcPrefix="string",
saveState=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
scale=TRUE | FALSE,
svdMaxRank=integer,
svdRand=list(
power=integer,
randSeed=integer
),
required parameter table=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression"
),
tolerance=double
)

Parameter Descriptions

attributes=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

changes the attributes of variables that are used in this action.

For more information about specifying the attributes parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

center=TRUE | FALSE

when set to True, centers the numeric variables by the mean of each column.

Aliascentering
DefaultFALSE

code=list(rpcaCodegen)

produces SAS score code.

casOut=list(casouttable)

specifies the settings for an output table.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

comment=TRUE | FALSE

when set to True, adds comments to the DATA step code.

DefaultFALSE
fmtWdth=integer

specifies the width to use for formatting derived numbers such as parameter estimates in the DATA step code.

AliasfmtWidth
Default20
Range0–32
indentSize=integer

specifies the number of spaces to indent the DATA step code for each indent level.

Default3
Range0–10
labelId=integer

specifies the label ID to use in array names and statement labels in the DATA step code. By default, a random positive integer is used.

lineSize=integer

specifies the line size for the generated code.

Default120
Range64–254
noTrim=TRUE | FALSE

requests that the comparison of variables with formatted values be based on the full format width, with padding. By default, leading and trailing blanks are removed from the formatted values.

DefaultFALSE
projectionType="LRS" | "PCA"

specifies the type of scoring.

DefaultPCA
LRS

projects the scoring observations in the low-rank space.

PCA

projects the scoring observations onto the principal components.

tabForm=TRUE | FALSE

when set to True, generates the code in way that is appropriate for storing in a table.

AliastableForm
DefaultFALSE

colStatistics=list(casouttable)

specifies the name of the output table to contain simple statistics for the variables of the input data set.

For more information about specifying the colStatistics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

cumEigPctTol=double

specifies the significance level of the eigenvalues that determine the rank of the low-rank matrix.

Default1
Range(0–1]

decomp="NONE" | "PCA" | "SVD"

specifies the decomposition method for the low-rank matrix. If the value of the maxiter parameter is 0, decomposition is applied to the original input data instead of to the low-rank matrix.

DefaultNONE
NONE

performs neither analysis.

PCA

performs principal component analysis.

SVD

performs singular value decomposition.

display=list(displayTables)

specifies a list of results tables to send to the client for display.

For more information about specifying the display parameter, see the common displayTables parameter (Appendix A: Common Parameters).

fixedMu=TRUE | FALSE

when set to True, fixes mu in each iteration of the accelerated proximal gradient method. Otherwise, mu is dynamically updated in each iteration.

DefaultFALSE

id=list("variable-name-1" <, "variable-name-2", ...>)

specifies the variables to use as record identifiers.

inputs=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the numeric variables to be analyzed. If this parameter is omitted, all numeric variables that are not specified in other parameters are analyzed.

For more information about specifying the inputs parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

lambda=double

specifies the value of lambda, which is multiplied by the L1 norm of the sparse matrix in the objective function.

Default-1
Range(0–10000000000]

lambdaWeight=double

specifies the weight of lambda.

Default1
Range(0–10000000000]

maxIter=integer

specifies the maximum number of iterations for robust principal component analysis algorithms.

Default1000
Minimum value0

method="ALM" | "APG"

specifies the method to use to perform the robust principal component analysis.

DefaultALM
ALM

uses the augmented Lagrange multiplier method.

APG

uses the accelerated proximal gradient method.

mu=double

specifies an initial value of mu in the objective function for the accelerated proximal gradient method.

Default0.001
Range0–10000000000

nThreads=integer

specifies the maximum number of threads to use on each computation node.

Default16
Minimum value1

outMat=list(outRpcaTabs)

specifies a list of parameters for the output tables of the robust principal component analysis method.

errMat=list(casouttable)

specifies the name of the output table for the error matrix.

AliasoutError
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

lowRankMat=list(casouttable)

specifies the name of the output table for the low-rank matrix.

AliasoutLowRank
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

sparseMat=list(casouttable)

specifies the name of the output table for the sparse matrix.

AliasoutSparse
caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

outPca=list(outPcaTabs)

specifies a list of parameters for the output tables of the principal component analysis.

pcLoadings=list(casouttable)

specifies the name of the output table for the principal component loadings.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

pcScores=list(casouttable)

specifies the name of the output table for the principal component scores.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

outSvd=list(outSvdTabs)

specifies a list of parameters for the output tables of the singular value decomposition.

svdDiag=list(casouttable)

specifies the name of the output table for the diagonal vector of the rectangular diagonal matrix.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

svdLeft=list(casouttable)

specifies the name of the output table for the left-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

svdRight=list(casouttable)

specifies the name of the output table for the right-singular vectors.

caslib="string"

specifies the name of the caslib to use.

compress=TRUE | FALSE

when set to True, data compression is applied to the table.

DefaultFALSE
indexVars=list("variable-name-1" <, "variable-name-2", ...>)

specifies the list of variables to create indexes for in the output data

label="string"

specifies the descriptive label to associate with the table.

maxMemSize=64-bit-integer

specifies the maximum amount of memory, in bytes, that each thread should allocate for in-memory blocks before converting to a memory-mapped file. Files are written in the directories that are specified in the CAS_DISK_CACHE environment variable.

TIPYou can enclose the value in quotation marks and specify B, K, M, G, or T as a suffix to indicate the units. For example, "8M" specifies eight megabytes.
name="table-name"

specifies the name to associate with the table.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
promote=TRUE | FALSE

when set to True, adds the output table with a global scope. This enables other sessions to access the table, subject to access controls. The target caslib must also have a global scope.

DefaultFALSE
replace=TRUE | FALSE

when set to True, overwrites an existing table with the same name.

DefaultFALSE
replication=integer

specifies the number of copies of the table to make for fault tolerance. Larger values result in slower performance and use more memory, but provide high availability for data in the event of a node failure.

Default1
Minimum value0
timeStamp="string"

specifies the timestamp to apply to the table. Specify the value in the form that is appropriate for your session locale.

where=list("string-1" <, "string-2", ...>)

specifies an expression for subsetting the output data.

outputTables=list(outputTables)

lists the names of results tables to save as CAS tables on the server.

For more information about specifying the outputTables parameter, see the common outputTables parameter (Appendix A: Common Parameters).

pcPrefix="string"

specifies a prefix for naming the principal components.

Default"Prin"

saveState=list(casouttable)

specifies the output data table in which to save the scoring results to be used in the score action of the aStore action set. You can specify the RPCA_PROJECTION_TYPE subparameter in the options parameter in the score action: the value 0 projects the scoring observations onto the principal component space, and the value 1 projects the scoring observations onto the low-rank subspace.

For more information about specifying the saveState parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scale=TRUE | FALSE

when set to True, scales the numeric variables by the standard deviation of each column.

Aliasscaling
DefaultFALSE

svdMaxRank=integer

specifies the maximum value for rank to be considered in the singular value decomposition solver.

Default0
Minimum value1

svdMethod="EIGEN" | "ITERATIVE" | "RANDOM"

specifies the type of the singular value decomposition solver.

DefaultEIGEN
EIGEN

uses the eigenvalue decomposition method.

ITERATIVE

uses the iterative singular value decomposition method.

RANDOM

uses the randomized singular value decomposition method.

svdRand=list(randomizedSvd)

specifies a list of parameters to use when the value of the svdMethod parameter is RANDOM.

power=integer

specifies the parameter power.

Default0
Minimum value0
randSeed=integer

specifies the seed value.

Default0
Minimum value1

* table=list(castable)

specifies the settings for an input table.

Long formtable=list(name="table-name")
Shortcut formtable="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasoptions, dataSource
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the convergence criterion for the robust principal component analysis algorithms.

Aliasstopcriterion
Default1e-07
Minimum value1e-10
Last updated: June 07, 2018