The PLSMOD Procedure
The PROC PLSMOD statement invokes the PLSMOD procedure. Table 3 summarizes the options available in the PROC PLSMOD statement.
Table 3: PROC PLSMOD Statement Options
| Option | Description |
|---|
| Basic Options |
|---|
|
DATA= | Specifies the CAS input data table |
| Model Fitting Options |
|---|
|
CVTEST | Requests that van der Voet’s (1994) randomization-based model comparison test be performed |
|
METHOD= | Specifies the general factor extraction method to be used |
|
NFAC= | Specifies the number of factors to extract |
|
NOCENTER | Suppresses centering of the responses and predictors before fitting |
|
NOCVSTDIZE | Suppresses re-centering and rescaling of the responses and predictors when cross validating |
|
NOSCALE | Suppresses scaling of the responses and predictors before fitting |
| Output Options |
|---|
|
CENSCALE | Displays the centering and scaling information |
|
DETAILS | Displays the details of the fitted model |
| NOCLPRINT | Limits or suppresses the display of class levels |
|
VARSS | Displays the amount of variation accounted for in each response and predictor |
The following list provides details about these options.
-
CENSCALE
lists the centering and scaling information for each response and predictor.
-
CVTEST <(cvtest-options)>
-
requests that van der Voet’s (1994) randomization-based model comparison test be performed to test models that have different numbers of extracted factors against the model that minimizes the predicted residual sum of squares. For more information, see the section Test Set Validation. You can also specify the following cvtest-options in parentheses:
-
NSAMP=number
specifies the number of randomizations to perform. By default, NSAMP=1000.
-
PVAL=number
specifies the cutoff probability for declaring an insignificant difference. By default, PVAL=0.10.
-
SEED=number
-
specifies the seed value for the random number stream. If you do not specify this option or if number is less than or equal to 0, the seed is generated by reading the time of day from the computer’s clock.
Analyses that use the same (nonzero) seed are not completely reproducible if they are executed on a different number of compute nodes, because the random number streams in separate compute nodes are independent.
-
STAT=PRESS | T2
-
specifies the test statistic for the model comparison. You can specify the following values:
- PRESS
uses the predicted residual sum of squares.
- T2
uses Hotelling’s
statistic.
By default, STAT=T2.
-
DATA=CAS-libref.data-table
-
names the input data table for PROC PLSMOD to use. The default is the most recently created data table. CAS-libref.data-table is a two-level name, where
- CAS-libref
refers to a collection of information that is defined in the LIBNAME statement and includes the caslib, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about CAS-libref, see the section Using CAS Sessions and CAS Engine Librefs.
- data-table
specifies the name of the input data table.
-
DETAILS
lists the details of the fitted model for each successive factor. The listed details are different for different extraction methods. For more information, see the section Displayed Output.
-
METHOD=PLS<(PLS-options)> | SIMPLS | PCR | RRR
-
specifies the general factor extraction method to be used. You can specify the following values:
-
PCR
uses principal component regression.
-
PLS<(PLS-options)>
-
uses partial least squares. You can also specify the following optional PLS-options in parentheses:
-
ALGORITHM=NIPALS | SVD | EIG
-
names the specific algorithm to use to compute extracted PLS factors. You can specify the following values:
- NIPALS
requests the usual iterative NIPALS algorithm.
- SVD
bases the extraction on the singular value decomposition of
. This algorithm is the most accurate but least efficient approach.
- EIG
bases the extraction on the eigenvalue decomposition of
.
By default, ALGORITHM=NIPALS.
-
EPSILON=number
specifies the convergence criterion for the NIPALS algorithm. By default, EPSILON=
.
-
MAXITER=number
specifies the maximum number of iterations for the NIPALS algorithm. By default, MAXITER=200.
-
RRR
uses reduced rank regression.
-
SIMPLS
uses the straightforward implementation of a statistically inspired modification of the partial least squares (SIMPLS) method of De Jong (1993).
By default, METHOD=PLS(NIPALS).
-
NFAC=number
specifies the number of factors to extract. The default is
, where p is the number of predictors (or the number of response variables when METHOD=RRR) and N is the number of runs (observations). You probably do not need to extract this many factors for most applications. Extracting too many factors can lead to an overfitted model (one that matches the training data too well), sacrificing predictive ability. Thus, if you use the default, you should also either specify the PARTITION statement to select the appropriate number of factors for the final model or consider the analysis to be preliminary and examine the results to determine the appropriate number of factors for a subsequent analysis.
-
NOCENTER
suppresses centering of the responses and predictors before fitting. This option is useful if the analysis variables are already centered and scaled. For more information, see the section Centering and Scaling.
-
NOCLPRINT<=number>
suppresses the display of the "Class Level Information" table if you do not specify number. If you specify number, the values of the classification variables are displayed only for variables whose number of levels is less than number. Specifying a number helps to reduce the size of the "Class Level Information" table if some classification variables have a large number of levels.
-
NOCVSTDIZE
suppresses re-centering and rescaling of the responses and predictors before each model is fit in the cross validation. For more information, see the section Centering and Scaling.
-
NOSCALE
suppresses scaling of the responses and predictors before fitting. This option is useful if the analysis variables are already centered and scaled. For more information, see the section Centering and Scaling.
lists, in addition to the average response and predictor sum of squares accounted for by each successive factor, the amount of variation accounted for in each response and predictor.
Last updated: September 17, 2021