EFA Procedure
Example 10.2 Principal Factor Analysis
(View the complete code for this example.)
This example shows how to use the EFA procedure to perform a principal factor analysis. It uses the socioeconomics data set that is described in the examples for the FACTOR procedure in SAS/STAT software. This data set contains five variables, which represent total population (Population), median years of school (School), total employment (Employment), miscellaneous professional services (Services), and median house value (HouseValue). Each observation represents one of twelve census tracts in the Los Angeles Standard Metropolitan Statistical Area.
For this example, it is assumed that your libref is named mylib, but you can substitute any appropriately defined libref.
Principal Factor Analysis (SAS/STAT User's Guide) and Maximum Likelihood Factor Analysis (SAS/STAT User's Guide) show that the socioeconomics data set is well described by a two-factor model. The following code uses PROC EFA to perform a principal factor analysis to extract two factors:
proc efa data=mylib.socioeconomics nfactors=2
method=principal rotate=promax reorder;
run;
When you use PROC EFA to extract a known number of factors, you use the NFACTORS= option to specify that number of factors. In this example, you specify the METHOD=PRINCIPAL option to perform principal factor extraction. PROC EFA uses squared multiple correlations by default to estimate the prior communalities for an analysis. You can use the PRIORS= option to specify alternative prior communality estimates.
The factors that are initially extracted by this method are orthogonal. To facilitate the interpretation of these factors, you specify the ROTATE=PROMAX option to perform factor rotation by the promax method. By default, the promax rotation first computes an orthogonal varimax prerotation, which is followed by an oblique Procrustes rotation. You can use the PREROTATE= option to specify a different prerotation method, which can be either orthogonal or oblique.
Finally, you specify the REORDER option so that the rows of the various factor matrices are ordered in such a way that variables whose highest absolute loading is on the first factor are displayed first, followed by variables whose highest absolute loading is on the second factor, and so on.
The first outputs that PROC EFA prints are the "Model Information" table and the "Number of Observations" table. These are shown in Output 10.2.1. Together, these tables summarize important information about the input data that are analyzed, as well as the primary analysis options that you specified. Next, the "Simple Statistics" table in Output 10.2.2 summarizes the means and standard deviations of the analysis variables.
Output 10.2.1: Model Information and Number of Observations
| Model Information | |
|---|---|
| Data Source | SOCIOECONOMICS |
| Data Type | Raw |
| Prior Communalities | Squared Multiple Correlations |
| Factor Extraction | Principal Factor Method |
| Factor Rotation | Promax |
| Number of Observations Read | 12 |
|---|---|
| Number of Observations Used | 12 |
Output 10.2.2: Simple Statistics
| Simple Statistics | ||
|---|---|---|
| Variable | Mean | Standard Deviation |
| Population | 6241.66667 | 3439.99427 |
| School | 11.44167 | 1.78654 |
| Employment | 2333.33333 | 1241.21153 |
| Services | 120.83333 | 114.92751 |
| HouseValue | 17000 | 6367.53128 |
The prior communality estimates are shown in the table in Output 10.2.3. The prior communalities represent initial estimates of the common variance for each analysis variable. (For more information about communalities and common variance, see the section The Common Factor Model.) The prior communalities are used to compute the reduced correlation matrix. The reduced correlation matrix is the observed correlation matrix for the analysis variables in which the diagonal entries of that matrix are replaced by the prior communalities.
Output 10.2.3: Prior Communalities
| Prior Communality Estimates | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 0.96859160 | 0.82228514 | 0.96918082 | 0.78572440 | 0.84701921 |
Next, the eigenvalues of the reduced correlation matrix are computed. These eigenvalues are summarized in the "Preliminary Eigenvalues" table in Output 10.2.4. In this example, the first two largest positive eigenvalues of the reduced correlation matrix account for 101.31% of the common variance. This is possible because the reduced correlation matrix, in general, is not necessarily positive definite, and negative eigenvalues for the matrix are possible. A pattern like this adds further support to the assumption that you might not need more than two common factors in this analysis.
Output 10.2.4: Preliminary Eigenvalues
| Eigenvalues of the Reduced Correlation Matrix: Total = 4.39280116 Average = 0.87856023 | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | 2.734301 | 1.018232 | 0.6225 | 0.6225 |
| 2 | 1.716069 | 1.676506 | 0.3907 | 1.0131 |
| 3 | 0.039563 | 0.064086 | 0.0090 | 1.0221 |
| 4 | -0.024523 | 0.048084 | -0.0056 | 1.0165 |
| 5 | -0.072608 | -0.0165 | 1.0000 | |
The principal factor pattern with the two extracted factors is displayed in Output 10.2.5. The Services variable has the largest loading on the first factor, and the Population variable has the smallest. The Population and Employment variables have large positive loadings on the second factor, and the HouseValue and School variables have large negative loadings.
Output 10.2.5: Factor Pattern
| Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| Services | 0.87899 | -0.15847 |
| HouseValue | 0.74215 | -0.57806 |
| Employment | 0.71447 | 0.67936 |
| School | 0.71370 | -0.55515 |
| Population | 0.62533 | 0.76621 |
The variance explained by the extracted factors is summarized in Output 10.2.6, and the final communality estimates are shown in Output 10.2.7.
Output 10.2.6: Variance Explained (Initial Solution)
| Variance Explained by Each Factor | |
|---|---|
| Factor1 | Factor2 |
| 2.7343008 | 1.7160687 |
Output 10.2.7: Final Communalities (Initial Solution)
| Final Communality Estimates: Total = 4.450370 | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 0.97811334 | 0.81756387 | 0.97199928 | 0.79774304 | 0.88494998 |
Next, you want to rotate the factor pattern such that most variables would have zero loadings on most factors. In this example, you specified a promax rotation with a varimax prerotation. To yield the prerotated factor pattern, the initial factor loading matrix is postmultiplied by an orthogonal transformation matrix that satisfies the varimax criterion. This orthogonal transformation matrix, as well as the varimax-rotated factor pattern, is shown in Output 10.2.8. The rotation leads to small loadings of Population and Employment on the first factor and small loadings of HouseValue and School on the second factor. Services appears to have a larger loading on the first factor than it has on the second factor, although both loadings are substantial. Hence, Services appears to be factorially complex.
The REORDER option helps you see the variable clusters clearly in the factor pattern. The first factor is associated more with the first three variables, HouseValue, School, and Services. The second factor is associated more with the last two variables, Population and Employment. For orthogonal factor solutions such as the current varimax-rotated solution, you can also interpret the values in the factor loading matrix as correlations. For example, HouseValue and Factor1 have a high correlation of 0.94, whereas Population and Factor1 have a low correlation of 0.03.
Output 10.2.8: Orthogonal Transformation and Factor Pattern (Prerotated Solution)
| Orthogonal Transformation Matrix | ||
|---|---|---|
| 1 | 2 | |
| 1 | 0.78895 | 0.61446 |
| 2 | -0.61446 | 0.78895 |
| Rotated Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| HouseValue | 0.94072 | -0.00004 |
| School | 0.90419 | 0.00055 |
| Services | 0.79085 | 0.41509 |
| Population | 0.02255 | 0.98874 |
| Employment | 0.14625 | 0.97499 |
The variance explained by the prerotated factors, as well as the final communality estimates, is shown in Output 10.2.9. The variance explained by the factors is more evenly distributed in the varimax-rotated solution than in the unrotated solution. Indeed, this is a typical fact for any kind of factor rotation. In the current example, before the varimax rotation, the two factors explain 2.73 and 1.72, respectively, of the common variance (see Output 10.2.6). After the varimax rotation, the two rotated factors explain 2.35 and 2.10, respectively, of the common variance. However, the total variance explained by the factors remains unchanged after the varimax rotation. This invariance property is also observed for the communalities of the variables after the rotation, as you can see by comparing the current communality estimates in Output 10.2.6 to those in Output 10.2.7.
Output 10.2.9: Variance Explained and Final Communalities (Prerotated Solution)
| Variance Explained by Each Factor | |
|---|---|
| Factor1 | Factor2 |
| 2.3498567 | 2.1005128 |
| Final Communality Estimates: Total = 4.450370 | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 0.97811334 | 0.81756387 | 0.97199928 | 0.79774304 | 0.88494998 |
Some factor analysts believe that common factors are seldom orthogonal. Thus an obliquely rotated factor solution might be more desirable, or at least it should be attempted. In the current example, you specified the ROTATE=PROMAX option to perform a promax rotation, which allows factors to be correlated. Output 10.2.10 shows the Procrustes target, followed by the Procrustes transformation matrix. This is the matrix that transforms the varimax factor pattern so that the rotated pattern is as close as possible to the Procrustes target, which is created by exponentiating the elements of the prerotated factor pattern (while keeping the signs of the values). However, because the variances of factors must be fixed at 1 during the oblique transformation, a normalized version of the Procrustes transformation matrix is the one that is actually used in the transformation. This normalized transformation matrix is shown at the bottom of Output 10.2.10.
Output 10.2.10: Promax Rotation: Procrustes Target and Transformation Matrices
| Target Matrix for Procrustes Transformation | ||
|---|---|---|
| Factor1 | Factor2 | |
| HouseValue | 1.00000 | -0.00000 |
| School | 1.00000 | 0.00000 |
| Services | 0.69421 | 0.10045 |
| Population | 0.00001 | 1.00000 |
| Employment | 0.00326 | 0.96793 |
| Procrustes Transformation Matrix | ||
|---|---|---|
| 1 | 2 | |
| 1 | 1.04117 | -0.09865 |
| 2 | -0.10572 | 0.96303 |
| Normalized Oblique Transformation Matrix | ||
|---|---|---|
| 1 | 2 | |
| 1 | 0.73803 | 0.54202 |
| 2 | -0.70555 | 0.86528 |
Using this transformation matrix leads to the promax-rotated factor solution, as shown in Output 10.2.11. After the promax rotation, the factors are no longer uncorrelated. As shown in the "Interfactor Correlations" table, the correlation of the two factors is now 0.20. In the unrotated and the varimax solutions, the two factors are not correlated.
Output 10.2.11: Promax Rotation: Factor Correlations and Rotated Factor Pattern
| Interfactor Correlations | ||
|---|---|---|
| Factor1 | Factor2 | |
| Factor1 | 1.00000 | 0.20188 |
| Factor2 | 0.20188 | 1.00000 |
| Rotated Factor Pattern (Standardized Regression Coefficients) | ||
|---|---|---|
| Factor1 | Factor2 | |
| HouseValue | 0.95558 | -0.09792 |
| School | 0.91842 | -0.09352 |
| Services | 0.76053 | 0.33932 |
| Population | -0.07908 | 1.00192 |
| Employment | 0.04799 | 0.97509 |
Unlike orthogonal factor solutions, where you can interpret the factor loadings as correlations between variables and factors, in oblique factor solutions such as the promax solution, you have to turn to the factor structure matrix to examine the correlations between variables and factors. Output 10.2.12 shows the factor structures for the promax-rotated solution. In this example, the factor structure matrix reflects a pattern similar to that of the factor pattern matrix shown in Output 10.2.11. The critical difference is that you can have the correlation interpretation only by using the factor structure matrix. For example, in the factor structure matrix, the correlation between Population and Factor 2 is 0.986. The corresponding value shown in the factor pattern matrix is 1.002, which you certainly cannot interpret as a correlation coefficient.
Output 10.2.12: Factor Structure
| Factor Structure (Correlations) | ||
|---|---|---|
| Factor1 | Factor2 | |
| HouseValue | 0.93582 | 0.09500 |
| School | 0.89954 | 0.09189 |
| Services | 0.82903 | 0.49286 |
| Population | 0.12319 | 0.98596 |
| Employment | 0.24484 | 0.98478 |
When factors are correlated, the common variance explained by each factor is not as straightforward as it is in the case when the factors are uncorrelated. The contributions from the other factors must be either eliminated or ignored. Both computations are shown in Output 10.2.13.
When you ignore the other factor, the common variance explained by the promax-rotated factors is 2.45 and 2.20, respectively. In contrast with orthogonal factor solutions (such as the prerotated varimax solution), the variance explained by these promax-rotated factors (ignoring other factors) does not sum to the total communality estimate of 4.45 (see Output 10.2.14).
Also shown in Output 10.2.13 is the "Variance Explained by Each Factor, Eliminating Other Factors" table. The values in this table represent unique variances explained by the factors, partialing out the variances explained by all other factors (or eliminating all other factors, as the table title suggests). In the current example, Factor1 explains 2.25 of the variable variances, partialing out all variable variances explained by Factor2.
Output 10.2.13: Variance Explained
| Variance Explained by Each Factor, Ignoring Other Factors | |
|---|---|
| Factor1 | Factor2 |
| 2.4473495 | 2.2022803 |
| Variance Explained by Each Factor, Eliminating Other Factors | |
|---|---|
| Factor1 | Factor2 |
| 2.2480892 | 2.0030200 |
The communalities of the variables, as shown in Output 10.2.14, do not change from rotation to rotation. They are still the same set of communalities in the initial, varimax-rotated, and promax-rotated solutions. This is a basic fact about factor rotations: they redistribute the variance explained by the factors; the total variance explained by the factors for any variable (that is, the communality of the variable) remains unchanged.
Output 10.2.14: Final Communalities
| Final Communality Estimates: Total = 4.450370 | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 0.97811334 | 0.81756387 | 0.97199928 | 0.79774304 | 0.88494998 |