EFA Procedure
Example 10.3 Iterative Factor Extraction
(View the complete code for this example.)
This example shows how to use the EFA procedure to perform an iterative factor analysis. It uses the socioeconomics data set that is described in the examples for the FACTOR procedure in SAS/STAT software. These are the same data that are analyzed in Example 10.2: Principal Factor Analysis.
For this example, it is assumed that your libref is named mylib, but you can substitute any appropriately defined libref.
In Example 10.2, the socioeconomics data set is used in a principal factor analysis. In that analysis, squared multiple correlations are used for prior communality estimates. After two factors are extracted, the factor pattern is used to compute updated estimates of the communalities. You might wonder whether this process can be repeated. That is, can you use the final communalities as prior communalities in a subsequent analysis? This is a motivation behind iterative methods of factor extraction.
The following code uses PROC EFA to perform an iterative principal factor analysis with squared multiple correlations for the prior communality estimates:
proc efa data=mylib.socioeconomics nfactors=2 method=prinit heywood;
run;
This code is very similar to the code in Example 10.2. The major difference is the use of the METHOD=PRINIT option here to specify an iterative principal factor extraction. This means that the prior communalities are used to compute an initial factor pattern, which is then used to compute new communalities, which is then used to compute a new factor pattern, and so on. This process continues until the change in the estimated communalities is sufficiently small. You can use the CONVERGE= option to set a specific value for the convergence threshold. The default value is 0.001.
Practically speaking, communalities are squared correlations and so must be bounded above by 1.0. However, during the iterative process of factor extraction, it is possible for a communality to meet or exceed this upper bound. When a communality reaches this bound, it is called a Heywood case. When a communality exceeds this bound, it is called an ultra-Heywood case. (For more information about communalities and Heywood cases, see the section Heywood Cases and Other Anomalies of Communality Estimates.) In this analysis, you use the HEYWOOD option to request that the communalities be bounded above by 1.0 during the iterative process. This means that if a communality exceeds 1.0 during the iterations, it is replaced by a value of 1.0. If you omit the HEYWOOD option, then by default PROC EFA terminates with an error when it encounters a Heywood or ultra-Heywood case.
Note that no rotation is requested in the current analysis. This is simply for clarity of exposition, so that this example can focus on iterative extraction methods. You can use the ROTATE= option to request a rotation for a factor pattern that is extracted using an iterative procedure in the same manner that you would for a noniterative extraction method.
The initial outputs that are produced by PROC EFA are shown in Output 10.3.1 through Output 10.3.4. These are identical to the outputs that are produced when you perform a noniterative principal factor analysis. For a discussion of these outputs, see Example 10.2.
Output 10.3.1: Model Information and Number of Observations
| Model Information | |
|---|---|
| Data Source | SOCIOECONOMICS |
| Data Type | Raw |
| Prior Communalities | Squared Multiple Correlations |
| Factor Extraction | Iterative Principal Factor Method |
| Factor Rotation | None |
| Number of Observations Read | 12 |
|---|---|
| Number of Observations Used | 12 |
Output 10.3.2: Simple Statistics
| Simple Statistics | ||
|---|---|---|
| Variable | Mean | Standard Deviation |
| Population | 6241.66667 | 3439.99427 |
| School | 11.44167 | 1.78654 |
| Employment | 2333.33333 | 1241.21153 |
| Services | 120.83333 | 114.92751 |
| HouseValue | 17000 | 6367.53128 |
Output 10.3.3: Prior Communalities
| Prior Communality Estimates | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 0.96859160 | 0.82228514 | 0.96918082 | 0.78572440 | 0.84701921 |
Output 10.3.4: Preliminary Eigenvalues
| Preliminary Eigenvalues: Total = 4.39280116 Average = 0.87856023 | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | 2.734301 | 1.018232 | 0.6225 | 0.6225 |
| 2 | 1.716069 | 1.676506 | 0.3907 | 1.0131 |
| 3 | 0.039563 | 0.064086 | 0.0090 | 1.0221 |
| 4 | -0.024523 | 0.048084 | -0.0056 | 1.0165 |
| 5 | -0.072608 | -0.0165 | 1.0000 | |
The iterative factor extraction process is summarized in Output 10.3.5. You can see that a Heywood case is encountered in the eighth iteration of the algorithm and remains in place until the algorithm converges.
Output 10.3.5: Iteration History
| Iteration | Change | Communalities | ||||
|---|---|---|---|---|---|---|
| 1 | 0.0379 | 0.97811 | 0.81756 | 0.97200 | 0.79774 | 0.88495 |
| 2 | 0.0235 | 0.98398 | 0.80831 | 0.97162 | 0.80096 | 0.90845 |
| 3 | 0.0161 | 0.98814 | 0.79921 | 0.97003 | 0.80132 | 0.92453 |
| 4 | 0.0117 | 0.99143 | 0.79155 | 0.96804 | 0.80085 | 0.93622 |
| 5 | 0.0088 | 0.99421 | 0.78548 | 0.96597 | 0.80021 | 0.94501 |
| 6 | 0.0067 | 0.99665 | 0.78078 | 0.96397 | 0.79961 | 0.95171 |
| 7 | 0.0052 | 0.99886 | 0.77717 | 0.96207 | 0.79909 | 0.95687 |
| 8 | 0.0040 | 1.00000 | 0.77442 | 0.96030 | 0.79867 | 0.96086 |
| 9 | 0.0031 | 1.00000 | 0.77232 | 0.95884 | 0.79834 | 0.96395 |
| 10 | 0.0024 | 1.00000 | 0.77072 | 0.95785 | 0.79810 | 0.96635 |
| 11 | 0.0019 | 1.00000 | 0.76949 | 0.95717 | 0.79791 | 0.96822 |
| 12 | 0.0015 | 1.00000 | 0.76854 | 0.95671 | 0.79776 | 0.96967 |
| 13 | 0.0011 | 1.00000 | 0.76782 | 0.95640 | 0.79764 | 0.97080 |
| 14 | 0.0009 | 1.00000 | 0.76726 | 0.95618 | 0.79754 | 0.97168 |
The eigenvalues of the reduced correlation matrix, the factor pattern, and the variance explained by the extracted factors are summarized in Output 10.3.6, Output 10.3.7, and Output 10.3.8, respectively. The final communality estimates are displayed in Output 10.3.9. In that table, you can see that the final communality estimate for Population is an ultra-Heywood case. Although the HEYWOOD option causes PROC EFA to replace out-of-bound communality estimates with ones during the iterative steps, this is not sufficient to effectively constrain the final community estimates after factoring the reduced correlation matrix. Consequently, ultra-Heywood cases might still be observed for the iterative principal factor method even if you specify this option. The ultra-Heywood case implies that some unique factor has negative variance—a clear indication that the factor solution in this case is problematic, if not invalid.
Output 10.3.6: Final Eigenvalues
| Eigenvalues of the Reduced Correlation Matrix: Total = 4.49264739 Average = 0.89852948 | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | 2.755916 | 1.016328 | 0.6134 | 0.6134 |
| 2 | 1.739588 | 1.708479 | 0.3872 | 1.0006 |
| 3 | 0.031109 | 0.033620 | 0.0069 | 1.0076 |
| 4 | -0.002511 | 0.028943 | -0.0006 | 1.0070 |
| 5 | -0.031455 | -0.0070 | 1.0000 | |
Output 10.3.7: Factor Pattern
| Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| Population | 0.62254 | 0.78441 |
| School | 0.70213 | -0.52371 |
| Employment | 0.70144 | 0.68130 |
| Services | 0.88116 | -0.14526 |
| HouseValue | 0.77905 | -0.60395 |
Output 10.3.8: Variance Explained
| Variance Explained by Each Factor | |
|---|---|
| Factor1 | Factor2 |
| 2.7559158 | 1.7395881 |
Output 10.3.9: Final Communalities
| Final Communality Estimates: Total = 4.495504 | ||||
|---|---|---|---|---|
| Population | School | Employment | Services | HouseValue |
| 1.00285106 | 0.76725612 | 0.95618481 | 0.79753541 | 0.97167647 |
You might instead consider an alternative factor extraction method to see if you can obtain a different factor solution. The following code uses maximum likelihood (ML) estimation for factor extraction:
proc efa data=mylib.socioeconomics nfactors=2 method=ml heywood;
run;
Like iterative principal factor extraction, the ML method is an iterative method. The iteration history is summarized in Output 10.3.10. The ML method converges in only 6 iterations, compared to the 14 iterations that are required for the iterative principal factor extraction method. Both methods encounter a Heywood case; the ML method encounters the Heywood case in the very first iteration.
Output 10.3.10: Maximum Likelihood Iteration History
| Iteration | Criterion | Ridge | Change | Communalities | ||||
|---|---|---|---|---|---|---|---|---|
| 1 | 0.4092306 | 0.0000 | 0.0908 | 1.00000 | 0.79229 | 0.93333 | 0.80068 | 0.93780 |
| 2 | 0.3088104 | 0.0000 | 0.0255 | 1.00000 | 0.79690 | 0.95885 | 0.81497 | 0.93545 |
| 3 | 0.3069206 | 0.0000 | 0.0130 | 1.00000 | 0.80872 | 0.95904 | 0.81890 | 0.92249 |
| 4 | 0.3067496 | 0.0000 | 0.0035 | 1.00000 | 0.80887 | 0.95952 | 0.81536 | 0.92334 |
| 5 | 0.3067333 | 0.0000 | 0.0013 | 1.00000 | 0.80995 | 0.95952 | 0.81589 | 0.92207 |
| 6 | 0.3067316 | 0.0000 | 0.0004 | 1.00000 | 0.80992 | 0.95956 | 0.81552 | 0.92219 |
In ML factor analysis, the eigenvalues of the weighted reduced correlation matrix are computed. The weights that are used are the reciprocals of the unique error variances (that is, one minus the communalities). When a Heywood case is encountered, the weight for that variable is not defined. For this reason, one of the eigenvalues in the "Eigenvalues of the Weighted Reduced Correlation Matrix" table in Output 10.3.11 is missing.
Output 10.3.11: Maximum Likelihood Final Eigenvalues
| Eigenvalues of the Weighted Reduced Correlation Matrix: Total = 19.8268757 Average = 4.95671892 | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | . | . | . | . |
| 2 | 19.826875 | 19.283728 | 1.0000 | 1.0000 |
| 3 | 0.543148 | 0.582869 | 0.0274 | 1.0274 |
| 4 | -0.039721 | 0.463706 | -0.0020 | 1.0254 |
| 5 | -0.503427 | -0.0254 | 1.0000 | |
As with iterative principal factor extraction, the ML method encounters a Heywood case. However, unlike the iterative principal factor results, the converged solution for the ML method has only a Heywood case rather than an ultra-Heywood case. You can see this in the "Final Communality Estimates" table in Output 10.3.13. Although a Heywood case does not necessarily invalidate a factor solution, it should raise suspicion and requires further investigation.
Output 10.3.12: Maximum Likelihood Solution
| Eigenvalues of the Weighted Reduced Correlation Matrix: Total = 19.8268757 Average = 4.95671892 | ||||
|---|---|---|---|---|
| Eigenvalue | Difference | Proportion | Cumulative | |
| 1 | . | . | . | . |
| 2 | 19.826875 | 19.283728 | 1.0000 | 1.0000 |
| 3 | 0.543148 | 0.582869 | 0.0274 | 1.0274 |
| 4 | -0.039721 | 0.463706 | -0.0020 | 1.0254 |
| 5 | -0.503427 | -0.0254 | 1.0000 | |
| Factor Pattern | ||
|---|---|---|
| Factor1 | Factor2 | |
| Population | 1.00000 | 0.00000 |
| School | 0.00975 | 0.89992 |
| Employment | 0.97245 | 0.11791 |
| Services | 0.43887 | 0.78926 |
| HouseValue | 0.02241 | 0.96004 |
| Variance Explained by Each Factor | ||
|---|---|---|
| Factor | Weighted | Unweighted |
| Factor1 | 24.436381 | 2.138861 |
| Factor2 | 19.826875 | 2.368369 |
Output 10.3.13: Maximum Likelihood Final Communalities
| Final Communality Estimates and Variable Weights | ||
|---|---|---|
| Total Communality: Weighted = . Unweighted = 4.507229 | ||
| Variable | Communality | Weight |
| Population | 1.00000000 | . |
| School | 0.80995115 | 5.26102907 |
| Employment | 0.95955909 | 24.7292489 |
| Services | 0.81553930 | 5.42072311 |
| HouseValue | 0.92217989 | 12.8522557 |