EFA Procedure
Getting Started: EFA Procedure
(View the complete code for this example.)
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This example demonstrates how you can use the EFA procedure to perform common factor analysis and factor rotation. The example uses the jobratings data set that is described in the examples for the FACTOR procedure in SAS/STAT software. In the original data set, 103 police officers were rated by their supervisors on 14 scales (variables). The current example excludes one of these scales—the overall rating variable.
The following DATA step creates the table jobratings:
options;
data mylib.jobratings;
input ('Communication Skills'n
'Problem Solving'n
'Learning Ability'n
'Judgment under Pressure'n
'Observational Skills'n
'Willingness to Confront Problems'n
'Interest in People'n
'Interpersonal Sensitivity'n
'Desire for Self-Improvement'n
'Appearance'n
'Dependability'n
'Physical Ability'n
'Integrity'n) (1.);
datalines;
2683885387986
7475887685766
5675786377587
6786977798899
9999779887888
8989789988879
8999988989979
8779479846888
3565233514311
... more lines ...
9989989989989
7665639956748
;
These statements assume that your libref is named mylib, but you can substitute any appropriately defined libref.
You conduct a common factor analysis on these variables to see what latent factors are operating behind these ratings. As part of this analysis, you want to determine the number of latent factors that are likely to explain the data.
The following statements invoke the EFA procedure:
proc efa data=mylib.jobratings rotate=varimax;
nfactors type=parallel alpha=0.05 seed=1031 nsimulations=20000;
nfactors type=parallel alpha=0.05 seed=1128 nsimulations=20000;
nfactors type=parallel alpha=0.10 seed=1031 nsimulations=20000;
nfactors type=parallel alpha=0.10 seed=1128 nsimulations=20000;
nfactors type=eigenvalue threshold=0.5 status=inactive;
nfactors type=eigenvalue threshold=1.0 status=inactive;
nfactors type=proportion threshold=0.9 status=inactive;
nfactors type=proportion threshold=0.95 status=inactive;
nfactors type=map2 status=inactive;
nfactors type=map4 status=inactive;
run;
In the common factor model, you always assume that observed variables are functions of underlying latent factors. (For more information about the common factor model, see the section The Common Factor Model.) The goal of factor analysis is to extract and interpret a factor pattern that describes these relationships. By default, PROC EFA uses the principal factor extraction method, with prior communalities set to the squared multiple correlations (SMC) of each variable with all the other variables. You can use the METHOD= option to specify a different factor extraction method. You can use the PRIORS= option to specify different prior communality estimates. To facilitate interpretations, the ROTATE= option specifies the varimax orthogonal factor rotation.
PROC EFA supports various criteria that you can use to suggest the number of factors to extract. You use the NFACTORS statement to specify a criterion. You can use multiple NFACTORS statements to specify multiple criteria. In this example, you determine the number of factors to extract by using multiple parallel analyses, each with 20,000 simulations. You include two different thresholds for the parallel analyses and two different random number seeds. By default, the actual number of factors that are extracted is the minimum of the numbers that are suggested by the four parallel analyses. You can use the NFACTORS=method option to change this default behavior, where method is MIN, MAX, MEDIAN, or MEAN.
In this example, you specify multiple additional NFACTORS statements, all of which include the STATUS=INACTIVE option. When a criterion has an inactive status, it is computed and the results are displayed, but it is not used to determine the final number of factors that are extracted. This can be useful for investigating the robustness of the number of extracted factors to alternative criteria. The minimum eigenvalue criterion (TYPE=EIGENVALUE) computes the number of factors to extract by comparing the eigenvalues of the reduced correlation matrix to a threshold value. The proportion of common variance criterion (TYPE=PROPORTION) computes the number of factors to extract by determining the smallest number of factors such that the cumulative proportion of common variance explained exceeds a threshold value. The minimum average partial correlation analyses (TYPE=MAP2 and MAP4) compute the number of factors by determining the number of principal components that, when partialed out, results in the smallest average squared partial (or residual) correlations.
If you know the exact number of factors that you want to extract in an analysis, you can omit the NFACTORS statement and specify the NFACTORS=n option in the PROC EFA statement, where n is the number of factors that you want to extract.
The output from the factor analysis is displayed in Figure 1 through Figure 11.
The first outputs that PROC EFA produces are the "Model Information" table and a table that summarizes the number of observations in the input data table. These are shown in Figure 1. Together, these tables summarize important information about the input data that are analyzed as well as the primary analysis options that you specified. For this analysis, the numbers of observations that are used is the same as the number that are read. These numbers would be different if some observations were discarded from the analysis. PROC EFA uses only complete cases for analysis, so a record is discarded if any field contains a missing value.
Figure 1: Model Information and Number of Observations
| Model Information | |
|---|---|
| Data Source | JOBRATINGS |
| Data Type | Raw |
| Prior Communalities | Squared Multiple Correlations |
| Factor Extraction | Principal Factor Method |
| Factor Rotation | Varimax |
| Number of Observations Read | 103 |
|---|---|
| Number of Observations Used | 103 |
Next, the "Simple Statistics" table in Figure 2 summarizes the means and standard deviations of the analysis variables.
Figure 2: Simple Statistics
| Simple Statistics | ||
|---|---|---|
| Variable | Mean | Standard Deviation |
| Communication Skills | 6.65049 | 1.76407 |
| Problem Solving | 6.63107 | 1.59035 |
| Learning Ability | 6.99029 | 1.33941 |
| Judgment under Pressure | 6.73786 | 1.73183 |
| Observational Skills | 6.93204 | 1.76158 |
| Willingness to Confront Problems | 7.29126 | 1.52516 |
| Interest in People | 6.70874 | 1.89235 |
| Interpersonal Sensitivity | 6.62136 | 1.76077 |
| Desire for Self-Improvement | 6.57282 | 1.72980 |
| Appearance | 7.00000 | 1.79869 |
| Dependability | 6.82524 | 1.91704 |
| Physical Ability | 7.20388 | 1.55525 |
| Integrity | 7.21359 | 1.84524 |
The summary output for the determination of the number of factors is shown in Figure 3. For most of the specified criteria, the results suggest that two latent factors should be retained to explain the data. The minimum eigenvalue criterion, with a threshold of 0.50, and the cumulative proportion criterion, with a threshold of 0.95, suggest the existence of a third latent factor. All four parallel analysis criteria suggest two latent factors, so the final number of latent factors that the analysis suggests is two.
Figure 3: Number of Factors
| Determination of Number of Factors | ||
|---|---|---|
| Criterion | Description | Factors |
| 1 | Parallel Analysis (Alpha = 0.05, Seed = 1031, Sims = 20000) | 2 |
| 2 | Parallel Analysis (Alpha = 0.05, Seed = 1128, Sims = 20000) | 2 |
| 3 | Parallel Analysis (Alpha = 0.10, Seed = 1031, Sims = 20000) | 2 |
| 4 | Parallel Analysis (Alpha = 0.10, Seed = 1128, Sims = 20000) | 2 |
| 5 | Minimum Eigenvalue (Eigenvalue > 0.50) | 3 |
| 6 | Minimum Eigenvalue (Eigenvalue > 1.00) | 2 |
| 7 | Total Proportion (Threshold = 0.90) | 2 |
| 8 | Total Proportion (Threshold = 0.95) | 3 |
| 9 | Min Ave Squared Partial Corr (MAP2) | 2 |
| 10 | Min Ave 4th-powered Partial Corr (MAP4) | 2 |
| Min = 2, Max = 2, Mean = 2, Median = 2 | ||
Additional information about the minimum eigenvalue and cumulative proportion criteria is presented in the "Eigenvalues of the Reduced Correlation Matrix" table in Figure 4. After the data correlation matrix is reduced by replacing its diagonal elements with the SMC prior communality estimates, the largest eigenvalue of the resulting matrix is approximately 6.18, and the corresponding latent factor explains approximately 77% of the common variance of the input data. The second eigenvalue is approximately 1.46, and the latent factors that correspond to these first two eigenvalues explain almost 95% of the common variance. A third eigenvalue has a value of approximately 0.56, and the latent factors that correspond to these three eigenvalues explain more than 100% of the common variance. Thus, two latent factors are sufficient to explain up to 90% of the common variance in the input data, and a third factor is needed to explain more than 95% of the common variance. Similarly, a minimum eigenvalue threshold of 1.00 produces an estimate of two latent factors, whereas a third latent factor is estimated if the minimum eigenvalue threshold is 0.50. (The fourth-largest eigenvalue for the reduced correlation matrix has a value of approximately 0.28.)
Figure 4: Eigenvalues and Common Variance Explained
| Eigenvalues of the Reduced Correlation Matrix | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | 6.177605 | 4.715319 | 0.7660 | 0.7660 |
| 2 | 1.462286 | 0.901833 | 0.1813 | 0.9473 |
| 3 | 0.560453 | 0.280939 | 0.0695 | 1.0168 |
| 4 | 0.279513 | 0.047660 | 0.0347 | 1.0515 |
| 5 | 0.231853 | 0.161134 | 0.0287 | 1.0802 |
| 6 | 0.070719 | 0.074896 | 0.0088 | 1.0890 |
| 7 | -0.004177 | 0.033875 | -0.0005 | 1.0885 |
| 8 | -0.038053 | 0.047765 | -0.0047 | 1.0838 |
| 9 | -0.085818 | 0.024381 | -0.0106 | 1.0731 |
| 10 | -0.110199 | 0.014527 | -0.0137 | 1.0595 |
| 11 | -0.124726 | 0.023565 | -0.0155 | 1.0440 |
| 12 | -0.148291 | 0.058236 | -0.0184 | 1.0256 |
| 13 | -0.206527 | -0.0256 | 1.0000 | |
The parallel analysis results are summarized in the tables in Figure 5. For each of the four specified parallel analysis criteria, correlation matrices are simulated from a Wishart distribution. These simulations represent correlation matrices that might be realized for a data-generating process in which all variables are independently and identically distributed according to the standard normal distribution. The eigenvalues of the simulated correlation matrices are computed and rank-ordered, and the eigenvalues of the input data correlation matrix are compared to the distributions of the eigenvalues at each position (largest eigenvalue, second-largest eigenvalue, and so on). When an eigenvalue of the input data correlation matrix exceeds the eigenvalue at some critical point (for example, the 95th percentile), this is understood to indicate the presence of a latent factor in the data. In all four parallel analyses in this example, the largest and second-largest eigenvalues of the input data correlation matrix exceed the simulated critical values. Thus all four parallel analyses suggest that two latent factors should be retained.
Figure 5: Parallel Analysis
| Parallel Analysis (Criterion 1) | ||
|---|---|---|
| Component | Simulated Crit Val | Observed Eigenvalue |
| 1 | 1.7988 | 6.5474* |
| 2 | 1.5809 | 1.7727* |
| 3 | 1.4326 | 1.0052 |
| 4 | 1.3112 | 0.7431 |
| 5 | 1.2084 | 0.6783 |
| 6 | 1.1153 | 0.4514 |
| 7 | 1.0308 | 0.3822 |
| 8 | 0.9510 | 0.2978 |
| 9 | 0.8732 | 0.2744 |
| 10 | 0.7988 | 0.2623 |
| 11 | 0.7234 | 0.2446 |
| 12 | 0.6482 | 0.1978 |
| 13 | 0.5652 | 0.1427 |
| * Retained Dimension (Obs > Crit) | ||
| Parallel Analysis (Criterion 2) | ||
|---|---|---|
| Component | Simulated Crit Val | Observed Eigenvalue |
| 1 | 1.7981 | 6.5474* |
| 2 | 1.5812 | 1.7727* |
| 3 | 1.4303 | 1.0052 |
| 4 | 1.3124 | 0.7431 |
| 5 | 1.2081 | 0.6783 |
| 6 | 1.1152 | 0.4514 |
| 7 | 1.0288 | 0.3822 |
| 8 | 0.9503 | 0.2978 |
| 9 | 0.8728 | 0.2744 |
| 10 | 0.7986 | 0.2623 |
| 11 | 0.7235 | 0.2446 |
| 12 | 0.6479 | 0.1978 |
| 13 | 0.5651 | 0.1427 |
| * Retained Dimension (Obs > Crit) | ||
| Parallel Analysis (Criterion 3) | ||
|---|---|---|
| Component | Simulated Crit Val | Observed Eigenvalue |
| 1 | 1.7599 | 6.5474* |
| 2 | 1.5528 | 1.7727* |
| 3 | 1.4106 | 1.0052 |
| 4 | 1.2929 | 0.7431 |
| 5 | 1.1922 | 0.6783 |
| 6 | 1.1016 | 0.4514 |
| 7 | 1.0168 | 0.3822 |
| 8 | 0.9372 | 0.2978 |
| 9 | 0.8601 | 0.2744 |
| 10 | 0.7842 | 0.2623 |
| 11 | 0.7095 | 0.2446 |
| 12 | 0.6328 | 0.1978 |
| 13 | 0.5499 | 0.1427 |
| * Retained Dimension (Obs > Crit) | ||
| Parallel Analysis (Criterion 4) | ||
|---|---|---|
| Component | Simulated Crit Val | Observed Eigenvalue |
| 1 | 1.7600 | 6.5474* |
| 2 | 1.5533 | 1.7727* |
| 3 | 1.4101 | 1.0052 |
| 4 | 1.2943 | 0.7431 |
| 5 | 1.1921 | 0.6783 |
| 6 | 1.1013 | 0.4514 |
| 7 | 1.0156 | 0.3822 |
| 8 | 0.9366 | 0.2978 |
| 9 | 0.8594 | 0.2744 |
| 10 | 0.7845 | 0.2623 |
| 11 | 0.7097 | 0.2446 |
| 12 | 0.6337 | 0.1978 |
| 13 | 0.5495 | 0.1427 |
| * Retained Dimension (Obs > Crit) | ||
The minimum average partial correlation analysis (MAP) results are summarized in the "Average Partial Correlations Controlling Principal Components" table in Figure 6. In this example, analyses are performed that use both squared (MAP2) and fourth-powered (MAP4) partial correlations. These two analyses are summarized in separate columns in the table. The number of latent factors that are estimated by a minimum average partial correlation analysis is the number of principal components that, when partialed out of the correlation matrix, produce the smallest average squared or fourth-powered partial correlations. As shown in Figure 6, the minimum value is obtained for both the MAP2 and the MAP4 analysis when two principal components are partialed out.
Figure 6: Minimum Average Partial Correlation Analysis
| Average Partial Correlations Controlling Principal Components | ||
|---|---|---|
| Number of Components | Squared | Fourth- Powered |
| 0 | 0.229076 | 0.069057 |
| 1 | 0.063655 | 0.011875 |
| 2 | 0.036328* | 0.003287* |
| 3 | 0.042970 | 0.004975 |
| 4 | 0.057102 | 0.009267 |
| 5 | 0.065326 | 0.014253 |
| 6 | 0.083345 | 0.022845 |
| 7 | 0.115706 | 0.037743 |
| 8 | 0.153423 | 0.050831 |
| 9 | 0.233063 | 0.120901 |
| 10 | 0.309145 | 0.167278 |
| 11 | 0.544163 | 0.416722 |
| 12 | 1.000000 | 1.000000 |
| * MAP = Minimum Values in Columns | ||
Now that the number of factors is determined, it is time to extract and rotate a factor pattern that corresponds to a two-factor solution. For this analysis, the prior communality estimates are set to the squared multiple correlations. These are shown in Figure 7.
Figure 7: Prior Communalities
| Prior Communality Estimates | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Communication Skills | Problem Solving | Learning Ability | Judgment under Pressure | Observational Skills | Willingness to Confront Problems | Interest in People | Interpersonal Sensitivity | Desire for Self-Improvement | Appearance | Dependability | Physical Ability | Integrity |
| 0.62981394 | 0.58657431 | 0.61009871 | 0.63766021 | 0.67187583 | 0.64779805 | 0.75641519 | 0.75584891 | 0.57460176 | 0.45505304 | 0.63449045 | 0.42245324 | 0.68195454 |
Figure 8 displays the initial factor pattern matrix. The factor pattern matrix represents standardized regression coefficients for predicting the variables by using the extracted factors. Because the initial factors are uncorrelated, the pattern matrix is also equal to the correlations between variables and the common factors.
The pattern matrix suggests that Factor1 represents general ability. All loadings for Factor1 in the Factor Pattern are at least 0.5. Factor2 consists of high positive loadings on certain task-related skills (Willingness to Confront Problems, Observational Skills, and Learning Ability) and high negative loadings on some interpersonal skills (Interpersonal Sensitivity, Interest in People, and Integrity). This factor measures individuals’ relative strength in these skills. Theoretically, individuals with high positive scores on this factor would exhibit better task-related skills than interpersonal skills. Individuals with high negative scores would exhibit better interpersonal skills than task-related skills. Individuals with scores near zero would have those skills balanced.
Figure 8: Initial Factor Pattern
| Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| Communication Skills | 0.75441 | 0.07707 |
| Problem Solving | 0.68590 | 0.08026 |
| Learning Ability | 0.65904 | 0.34808 |
| Judgment under Pressure | 0.73391 | -0.21405 |
| Observational Skills | 0.69039 | 0.45292 |
| Willingness to Confront Problems | 0.66458 | 0.47460 |
| Interest in People | 0.70770 | -0.53427 |
| Interpersonal Sensitivity | 0.64668 | -0.61284 |
| Desire for Self-Improvement | 0.73820 | 0.12506 |
| Appearance | 0.57188 | 0.20052 |
| Dependability | 0.79475 | -0.04516 |
| Physical Ability | 0.51285 | 0.10251 |
| Integrity | 0.74906 | -0.35091 |
Figure 9 displays the proportion of variance explained by each factor and the final communality estimates, including the total communality. The final communality estimates are the proportion of variance of the variables accounted for by the common factors. When the factors are orthogonal, the final communalities are calculated by taking the sum of squares of each row of the factor pattern matrix.
Figure 9: Initial Solution: Variance Explained and Final Communality Estimates
| Variance Explained by Each Factor | |
|---|---|
| Factor1 | Factor2 |
| 6.1776055 | 1.4622860 |
| Final Communality Estimates: Total = 7.639892 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Communication Skills | Problem Solving | Learning Ability | Judgment under Pressure | Observational Skills | Willingness to Confront Problems | Interest in People | Interpersonal Sensitivity | Desire for Self-Improvement | Appearance | Dependability | Physical Ability | Integrity |
| 0.57507775 | 0.47690128 | 0.55548854 | 0.58444301 | 0.68176986 | 0.66690409 | 0.78628588 | 0.79376521 | 0.56057957 | 0.36724994 | 0.63367286 | 0.27352712 | 0.68422640 |
Figure 10 displays the results of the varimax rotation of the two extracted factors and the corresponding orthogonal transformation matrix. The rotated factor pattern matrix is calculated by postmultiplying the original factor pattern matrix by the transformation matrix.
The rotated factor pattern matrix is somewhat easier to interpret. If a magnitude of at least 0.5 is required to indicate a salient variable-factor relationship, Factor1 now measures job enthusiasm and cognitive skills (Observational Skills, Willingness to Confront Problems, Learning Ability, Desire for Self-Improvement, Communication Skills, Problem Solving, Dependability, and Appearance) and Factor2 represents interpersonal skills (Interpersonal Sensitivity, Interest in People, Integrity, Judgment under Pressure, and Dependability).
Figure 10: Transformation Matrix and Rotated Factor Pattern
| Orthogonal Transformation Matrix | ||
|---|---|---|
| 1 | 2 | |
| 1 | 0.75128 | 0.65999 |
| 2 | 0.65999 | -0.75128 |
| Rotated Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| Communication Skills | 0.61764 | 0.44000 |
| Problem Solving | 0.56827 | 0.39239 |
| Learning Ability | 0.72485 | 0.17346 |
| Judgment under Pressure | 0.41009 | 0.64519 |
| Observational Skills | 0.81759 | 0.11538 |
| Willingness to Confront Problems | 0.81251 | 0.08206 |
| Interest in People | 0.17906 | 0.86846 |
| Interpersonal Sensitivity | 0.08137 | 0.88721 |
| Desire for Self-Improvement | 0.63713 | 0.39325 |
| Appearance | 0.56198 | 0.22679 |
| Dependability | 0.56727 | 0.55846 |
| Physical Ability | 0.45295 | 0.26147 |
| Integrity | 0.33116 | 0.75800 |
To make the rotated factor pattern matrix even easier to interpret, you can use the REORDER option to rearrange the rows of the factor pattern. When you specify the REORDER option, variables whose highest absolute loading (reference structure loading for oblique rotations) is on the first factor are displayed first, from largest to smallest loading, followed by variables whose highest absolute loading is on the second factor, and so on. For an example of how to use the REORDER option, see Example 10.2: Principal Factor Analysis. You can also specify the FUZZ=0.5 option, which causes all the entries of the factor pattern smaller than 0.5 in absolute value to be printed as missing. For an example of how to use the FUZZ= option, see Example 10.4: Extracting and Interpreting Solutions with Maximum Likelihood Factor Analysis.
Figure 11 displays the variance explained by each factor and the final communality estimates after the orthogonal rotation. Even though the variances explained by the rotated factors are different from those of the unrotated factors (compare to Figure 9), the cumulative variance explained by the common factors remains the same. Note also that the final communalities for variables, as well as the total communality, remain unchanged after rotation. Although rotating a factor solution does not increase or decrease the statistical quality of the factor model, it can simplify the interpretations of the factors and redistribute the variance explained by the factors.
Figure 11: Rotated Solution: Variance Explained and Final Communality Estimates
| Variance Explained by Each Factor | |
|---|---|
| Factor1 | Factor2 |
| 4.1236793 | 3.5162122 |
| Final Communality Estimates: Total = 7.639892 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Communication Skills | Problem Solving | Learning Ability | Judgment under Pressure | Observational Skills | Willingness to Confront Problems | Interest in People | Interpersonal Sensitivity | Desire for Self-Improvement | Appearance | Dependability | Physical Ability | Integrity |
| 0.57507775 | 0.47690128 | 0.55548854 | 0.58444301 | 0.68176986 | 0.66690409 | 0.78628588 | 0.79376521 | 0.56057957 | 0.36724994 | 0.63367286 | 0.27352712 | 0.68422640 |