The FACTOR Procedure

Example 38.2 Principal Factor Analysis

(View the complete code for this example.)

This example uses the data presented in Example 38.1 and performs a principal factor analysis with squared multiple correlations for the prior communality estimates. Unlike Example 38.1, which analyzes the principal components (with default PRIORS=ONE), the current analysis is based on a common factor model. To use a common factor model, you specify PRIORS=SMC in the PROC FACTOR statement, as shown in the following:

ods graphics on;

proc factor data=SocioEconomics
   priors=smc msa residual
   rotate=promax reorder
   outstat=fact_all
   plots=(scree initloadings preloadings loadings);
run;
ods graphics off;

In the PROC FACTOR statement, you include several other options to help you analyze the results. To help determine whether the common factor model is appropriate, you request the Kaiser’s measure of sampling adequacy with the MSA option. You specify the RESIDUALS option to compute the residual correlations and partial correlations.

The ROTATE= and REORDER options are specified to enhance factor interpretability. The ROTATE=PROMAX option produces an orthogonal varimax prerotation (default) followed by an oblique Procrustes rotation, and the REORDER option reorders the variables according to their largest factor loadings. An OUTSTAT= data set is created by PROC FACTOR and displayed in Output 38.2.15.

PROC FACTOR can produce high-quality graphs that are very useful for interpreting the factor solutions. To request these graphs, ODS Graphics must be enabled. All ODS graphs in PROC FACTOR are requested with the PLOTS= option. In this example, you request a scree plot (SCREE) and loading plots for the factor matrix during the following three stages: initial unrotated solution (INITLOADINGS), prerotated (varimax) solution (PRELOADINGS), and promax-rotated solution (LOADINGS). The scree plot helps you determine the number of factors, and the loading plots help you visualize the patterns of factor loadings during various stages of analyses.

Principal Factor Analysis: Kaiser’s MSA and Factor Extraction Results

Output 38.2.1 displays the results of the partial correlations and Kaiser’s measure of sampling adequacy.

Output 38.2.1: Principal Factor Analysis: Partial Correlations and Kaiser’s MSA

Partial Correlations Controlling all other Variables
 PopulationSchoolEmploymentServicesHouseValue
Population1.00000-0.544650.970830.096120.15871
School-0.544651.000000.543730.049960.64717
Employment0.970830.543731.000000.06689-0.25572
Services0.096120.049960.066891.000000.59415
HouseValue0.158710.64717-0.255720.594151.00000

Kaiser's Measure of Sampling Adequacy: Overall MSA = 0.57536759
PopulationSchoolEmploymentServicesHouseValue
0.472078970.551588390.488511370.806643650.61281377


If the data are appropriate for the common factor model, the partial correlations (controlling all other variables) should be small compared to the original correlations. For example, the partial correlation between the variables School and HouseValue is 0.65, slightly less than the original correlation of 0.86 (see Output 38.1.3). The partial correlation between Population and School is –0.54, which is much larger in absolute value than the original correlation; this is an indication of trouble. Kaiser’s MSA is a summary, for each variable and for all variables together, of how much smaller the partial correlations are than the original correlations. Values of 0.8 or 0.9 are considered good, while MSAs below 0.5 are unacceptable. The variables Population, School, and Employment have very poor MSAs. Only the Services variable has a good MSA. The overall MSA of 0.58 is sufficiently poor that additional variables should be included in the analysis to better define the common factors. A commonly used rule is that there should be at least three variables per factor. In the following analysis, you determine that there are two common factors in these data. Therefore, more variables are needed for a reliable analysis.

Output 38.2.2 displays the results of the principal factor extraction.

Output 38.2.2: Principal Factor Analysis: Factor Extraction

Prior Communality Estimates: SMC
PopulationSchoolEmploymentServicesHouseValue
0.968591600.822285140.969180820.785724400.84701921

Eigenvalues of the Reduced Correlation Matrix: Total = 4.39280116 Average = 0.87856023
 EigenvalueDifferenceProportionCumulative
12.734300841.018232170.62250.6225
21.716068671.676505860.39071.0131
30.039562810.064086260.00901.0221
4-.024523450.04808427-0.00561.0165
5-.07260772 -0.01651.0000


The square multiple correlations are shown as prior communality estimates in Output 38.2.2. The PRIORS=SMC option basically replaces the diagonal of the original observed correlation matrix by these square multiple correlations. Because the square multiple correlations are usually less than one, the resulting correlation matrix for factoring is called the reduced correlation matrix. In the current example, the SMCs are all fairly large; hence, you expect the results of the principal factor analysis to be similar to those in the principal component analysis.

The first two largest positive eigenvalues of the reduced correlation matrix account for of the common variance. This is possible because the reduced correlation matrix, in general, is not necessarily positive definite, and negative eigenvalues for the matrix are possible. A pattern like this suggests that you might not need more than two common factors. The scree and variance explained plots of Output 38.2.3 clearly support the conclusion that two common factors are present. Showing in the left panel of Output 38.2.3 is the scree plot of the eigenvalues of the reduced correlation matrix. A sharp bend occurs at the third eigenvalue, reinforcing the conclusion that two common factors are present. These cumulative proportions of common variance explained by factors are plotted in the right panel of Output 38.2.3, which shows that the curve essentially flattens out after the second factor.

Output 38.2.3: Scree and Variance Explained Plots

Scree and Variance Explained Plots


Principal Factor Analysis: Initial Factor Solution

For the current analysis, PROC FACTOR retains two factors by certain default criteria. This decision agrees with the conclusion drawn by inspecting the scree plot. The principal factor pattern with the two factors is displayed in Output 38.2.4. This factor pattern is similar to the principal component pattern seen in Output 38.1.5 of Example 38.1. For example, the variable Services has the largest loading on the first factor, and the Population variable has the smallest. The variables Population and Employment have large positive loadings on the second factor, and the HouseValue and School variables have large negative loadings.

Output 38.2.4: Initial Factor Pattern Matrix and Communalities

Factor Pattern
 Factor1Factor2
Services0.87899-0.15847
HouseValue0.74215-0.57806
Employment0.714470.67936
School0.71370-0.55515
Population0.625330.76621

Variance Explained by Each
Factor
Factor1Factor2
2.73430081.7160687

Final Communality Estimates: Total = 4.450370
PopulationSchoolEmploymentServicesHouseValue
0.978113340.817563870.971999280.797743040.88494998


Comparing the current factor loading matrix in Output 38.2.4 with that in Output 38.1.5 in Example 38.1, you notice that the variables are arranged differently in the two output tables. This is due to the use of the REORDER option in the current analysis. The advantage of using this option might not be very obvious in Output 38.2.4, but you can see its value when looking at the rotated solutions, as shown in Output 38.2.7 and Output 38.2.11.

The final communality estimates are all fairly close to the priors (shown in Output 38.2.2). Only the communality for the variable HouseValue increased appreciably, from 0.847 to 0.885. Therefore, you are sure that all the common variance is accounted for.

Output 38.2.5 shows that the residual correlations (off-diagonal elements) are low, the largest being 0.03. The partial correlations are not quite as impressive, since the uniqueness values are also rather small. These results indicate that the squared multiple correlations are good but not quite optimal communality estimates.

Output 38.2.5: Residual and Partial Correlations

Residual Correlations With Uniqueness on the Diagonal
 PopulationSchoolEmploymentServicesHouseValue
Population0.02189-0.011180.005140.010630.00124
School-0.011180.182440.02151-0.023900.01248
Employment0.005140.021510.02800-0.00565-0.01561
Services0.01063-0.02390-0.005650.202260.03370
HouseValue0.001240.01248-0.015610.033700.11505

Root Mean Square Off-Diagonal Residuals: Overall = 0.01693282
PopulationSchoolEmploymentServicesHouseValue
0.008153070.018130270.013827640.021517370.01960158

Partial Correlations Controlling Factors
 PopulationSchoolEmploymentServicesHouseValue
Population1.00000-0.176930.207520.159750.02471
School-0.176931.000000.30097-0.124430.08614
Employment0.207520.300971.00000-0.07504-0.27509
Services0.15975-0.12443-0.075041.000000.22093
HouseValue0.024710.08614-0.275090.220931.00000

Root Mean Square Off-Diagonal Partials: Overall = 0.18550132
PopulationSchoolEmploymentServicesHouseValue
0.158508240.190258670.231818380.154470430.18201538


As displayed in Output 38.2.6, the unrotated factor pattern reveals two tight clusters of variables, with the variables HouseValue and School at the negative end of Factor2 axis and the variables Employment and Population at the positive end. The Services variable is in between but closer to the HouseValue and School variables. A good rotation would place the axes so that most variables would have zero loadings on most factors. As a result, the axes would appear as though they are put through the variable clusters.

Output 38.2.6: Unrotated Factor Loading Plot

Unrotated Factor Loading Plot


Principal Factor Analysis: Varimax Prerotation

In Output 38.2.7, the results of the varimax prerotation are shown. To yield the varimax-rotated factor loading (pattern), the initial factor loading matrix is postmultiplied by an orthogonal transformation matrix. This orthogonal transformation matrix is shown in Output 38.2.7, followed by the varimax-rotated factor pattern. This rotation or transformation leads to small loadings of Population and Employment on the first factor and small loadings of HouseValue and School on the second factor. Services appears to have a larger loading on the first factor than it has on the second factor, although both loadings are substantial. Hence, Services appears to be factorially complex.

With the REORDER option in effect, you can see the variable clusters clearly in the factor pattern. The first factor is associated more with the first three variables (first three rows of variables): HouseValue, School, and Services. The second factor is associated more with the last two variables (last two rows of variables): Population and Employment.

For orthogonal factor solutions such as the current varimax-rotated solution, you can also interpret the values in the factor loading (pattern) matrix as correlations. For example, HouseValue and Factor 1 have a high correlation at 0.94, while Population and Factor 1 have a low correlation at 0.02.

Output 38.2.7: Varimax Rotation: Transform Matrix and Rotated Pattern

Orthogonal Transformation Matrix
 12
10.788950.61446
2-0.614460.78895

Rotated Factor Pattern
 Factor1Factor2
HouseValue0.94072-0.00004
School0.904190.00055
Services0.790850.41509
Population0.022550.98874
Employment0.146250.97499

Variance Explained by Each
Factor
Factor1Factor2
2.34985672.1005128

Final Communality Estimates: Total = 4.450370
PopulationSchoolEmploymentServicesHouseValue
0.978113340.817563870.971999280.797743040.88494998


The variance explained by the factors are more evenly distributed in the varimax-rotated solution, as compared with that of the unrotated solution. Indeed, this is a typical fact for any kinds of factor rotation. In the current example, before the varimax rotation the two factors explain 2.73 and 1.72, respectively, of the common variance (see Output 38.2.4). After the varimax rotation the two rotated factors explain 2.35 and 2.10, respectively, of the common variance. However, the total variance accounted for by the factors remains unchanged after the varimax rotation. This invariance property is also observed for the communalities of the variables after the rotation, as evidenced by comparing the current communality estimates in Output 38.2.7 with those in Output 38.2.4.

Output 38.2.8 shows the graphical plot of the varimax-rotated factor loadings. Clearly, HouseValue and School cluster together on the Factor 1 axis, while Population and Employment cluster together on the Factor 2 axis. Service is closer to the cluster of HouseValue and School.

Output 38.2.8: Varimax-Rotated Factor Loadings

Varimax-Rotated Factor Loadings


An alternative to the scatter plot of factor loadings is the so-called vector plot of loadings, which is shown in Output 38.2.9. The vector plot is requested with the suboption VECTOR in the PLOTS= option. That is:

plots=preloadings(vector)

This generates the vector plot of loadings in Output 38.2.9.

Output 38.2.9: Varimax-Rotated Factor Loadings: Vector Plot

Varimax-Rotated Factor Loadings: Vector Plot


Principal Factor Analysis: Oblique Promax Rotation

For some researchers, the varimax-rotated factor solution in the preceding section might be good enough to provide them useful and interpretable results. For others who believe that common factors are seldom orthogonal, an obliquely rotated factor solution might be more desirable, or at least should be attempted.

PROC FACTOR provides a very large class of oblique factor rotations. The current example shows a particular one—namely, the promax rotation as requested by the ROTATE=PROMAX option.

The results of the promax rotation are shown in Output 38.2.10 and Output 38.2.11. The corresponding plot of factor loadings is shown in Output 38.2.12.

Output 38.2.10: Promax Rotation: Procrustean Target and Transformation

Target Matrix for Procrustean Transformation
 Factor1Factor2
HouseValue1.00000-0.00000
School1.000000.00000
Services0.694210.10045
Population0.000011.00000
Employment0.003260.96793

Procrustean Transformation Matrix
 12
11.04116598-0.0986534
2-0.10572260.96303019

Normalized Oblique Transformation
Matrix
 12
10.738030.54202
2-0.705550.86528


Output 38.2.10 shows the Procrustean target, to which the varimax factor pattern is rotated, followed by the display of the Procrustean transformation matrix. This is the matrix that transforms the varimax factor pattern so that the rotated pattern is as close as possible to the Procrustean target. However, because the variances of factors have to be fixed at 1 during the oblique transformation, a normalized version of the Procrustean transformation matrix is the one that is actually used in the transformation. This normalized transformation matrix is shown at the bottom of Output 38.2.10. Using this transformation matrix leads to the promax-rotated factor solution, as shown in Output 38.2.11.

Output 38.2.11: Promax Rotation: Factor Correlations and Factor Pattern

Inter-Factor Correlations
 Factor1Factor2
Factor11.000000.20188
Factor20.201881.00000

Rotated Factor Pattern (Standardized Regression Coefficients)
 Factor1Factor2
HouseValue0.95558485-0.0979201
School0.91842142-0.0935214
Services0.760532380.33931804
Population-0.07908321.00192402
Employment0.047990.97509085


After the promax rotation, the factors are no longer uncorrelated. As shown in Output 38.2.11, the correlation of the two factors is now 0.20. In the (initial) unrotated and the varimax solutions, the two factors are not correlated.

In addition to allowing the factors to be correlated, in an oblique factor solution you seek a pattern of factor loadings that is more "differentiated" (referred to as the "simple structures" in the literature). The more differentiated the loadings, the easier the interpretation of the factors.

For example, factor loadings of Services and Population on Factor 2 are 0.415 and 0.989, respectively, in the (orthogonal) varimax-rotated factor pattern (see Output 38.2.7). With the (oblique) promax rotation (see Output 38.2.11), these two loadings become even more differentiated with values 0.339 and 1.002, respectively. Overall, however, the factor patterns before and after the promax rotation do not seem to differ too much. This fact is confirmed by comparing the graphical plots of factor loadings. The plots in Output 38.2.12 (promax-rotated factor loadings) and Output 38.2.8 (varimax-rotated factor loadings) show very similar patterns.

Output 38.2.12: Promax Rotation: Factor Loading Plot

Promax Rotation: Factor Loading Plot


Unlike the orthogonal factor solutions where you can interpret the factor loadings as correlations between variables and factors, in oblique factor solutions such as the promax solution, you have to turn to the factor structure matrix for examining the correlations between variables and factors. Output 38.2.13 shows the factor structures of the promax-rotated solution.

Output 38.2.13: Promax Rotation: Factor Structures and Final Communalities

Factor Structure (Correlations)
 Factor1Factor2
HouseValue0.935820.09500
School0.899540.09189
Services0.829030.49286
Population0.123190.98596
Employment0.244840.98478

Variance Explained by Each
Factor Ignoring Other Factors
Factor1Factor2
2.44734952.2022803

Final Communality Estimates: Total = 4.450370
PopulationSchoolEmploymentServicesHouseValue
0.978113340.817563870.971999280.797743040.88494998


Basically, the factor structure matrix shown in Output 38.2.13 reflects a similar pattern to the factor pattern matrix shown in Output 38.2.11. The critical difference is that you can have the correlation interpretation only by using the factor structure matrix. For example, in the factor structure matrix shown in Output 38.2.13, the correlation between Population and Factor 2 is 0.986. The corresponding value shown in the factor pattern matrix in Output 38.2.11 is 1.002, which certainly cannot be interpreted as a correlation coefficient.

Common variance explained by the promax-rotated factors are 2.447 and 2.202, respectively, for the two factors. Unlike the orthogonal factor solutions (for example, the prerotated varimax solution), variance explained by these promax-rotated factors do not sum up to the total communality estimate 4.45. In oblique factor solutions, variance explained by oblique factors cannot be partitioned for the factors. Variance explained by a common factor is computed while ignoring the contributions from the other factors.

However, the communalities for the variables, as shown in the bottom of Output 38.2.13, do not change from rotation to rotation. They are still the same set of communalities in the initial, varimax-rotated, and promax-rotated solutions. This is a basic fact about factor rotations: they only redistribute the variance explained by the factors; the total variance explained by the factors for any variable (that is, the communality of the variable) remains unchanged.

In the literature of exploratory factor analysis, reference axes had been an important tool in factor rotation. Nowadays, rotations are seldom done through the uses of the reference axes. Despite that, results about reference axes do provide additional information for interpreting factor analysis results. For the current example of the promax rotation, PROC FACTOR shows the relevant results about the reference axes in Output 38.2.14.

Output 38.2.14: Promax Rotation: Reference Axis Correlations and Reference Structures

Reference Axis Correlations
 Factor1Factor2
Factor11.00000-0.20188
Factor2-0.201881.00000

Reference Structure (Semipartial Correlations)
 Factor1Factor2
HouseValue0.93591-0.09590
School0.89951-0.09160
Services0.744870.33233
Population-0.077450.98129
Employment0.047000.95501

Variance Explained by Each
Factor Eliminating Other
Factors
Factor1Factor2
2.24808922.0030200


To explain the results in the reference-axis system, some geometric interpretations of the factor axes are needed. Consider a single factor in a system of n common factors in an oblique factor solution. Taking away the factor under consideration, the remaining n – 1 factors span a hyperplane in the factor space of n – 1 dimensions. The vector that is orthogonal to this hyperplane is the reference axis (reference vector) of the factor under consideration. Using the same definition for the remaining factors, you have n reference vectors for n factors.

A factor in an oblique factor solution can be considered as the sum of two independent components: its associated reference vector and a component that is overlapped with all other factors. In other words, the reference vector of a factor is a unique part of the factor that is not predictable from all other factors. Thus, the loadings on a reference vector are the unique effects of the corresponding factor, partialling out the effects from all other factors. The variances explained by a reference vector are the unique variances explained by the corresponding factor, partialling out the variances explained by all other factors.

Output 38.2.14 shows the reference axis correlations. The correlation between the reference vectors is –0.20. Next, Output 38.2.14 shows the loadings on the reference vectors in the table entitled "Reference Structure (Semipartial Correlations)." As explained previously, loadings on a reference vector are also the unique effects of the corresponding factor, partialling out the effects from the all other factors. For example, the unique effect of Factor 1 on HouseValue is 0.936. Another important property of the reference vector system is that loadings on a reference vector are also correlations between the variables and the corresponding factor, partialling out the correlations between the variables and other factors. This means that the loading 0.936 in the reference structure table is the unique correlation between HouseValue and Factor 1, partialling out the correlation between HouseValue with Factor 2. Hence, as suggested by the title of table, all loadings reported in the "Reference Structure (Semipartial Correlations)" can be interpreted as semipartial correlations between variables and factors.

The last table shown in Output 38.2.14 are the variances explained by the reference vectors. As explained previously, these are also unique variances explained by the factors, partialling out the variances explained by all other factors (or eliminating all other factors, as suggested by the title of the table). In the current example, Factor 1 explains 2.248 of the variable variances, partialling out all variable variances explained by Factor 2.

Notice that factor pattern (shown in Output 38.2.11), factor structures (correlations, shown in Output 38.2.13), and reference structures (semipartial correlations, shown in Output 38.2.14) give you different information about the oblique factor solutions such as the promax-rotated solution. However, for orthogonal factor solutions such as the varimax-rotated solution, factor structures and reference structures are all the same as the factor pattern.

Principal Factor Analysis: Factor Rotations with Factor Pattern Input

The promax rotation is one of the many rotations that PROC FACTOR provides. You can specify many different rotation algorithms by using the ROTATE= options. In this section, you explore different rotated factor solutions from the initial principal factor solution. Specifically, you want to examine the factor patterns yielded by the quartimax transformation (an orthogonal transformation) and the Harris-Kaiser (an oblique transformation), respectively.

Rather than analyzing the entire problem again with new rotations, you can simply use the OUTSTAT= data set from the preceding factor analysis results.

First, the OUTSTAT= data set is printed using the following statements:

proc print data=fact_all;
run;

The output data set is displayed in Output 38.2.15.

Output 38.2.15: Output Data Set

Factor Output Data Set

Obs_TYPE__NAME_PopulationSchoolEmploymentServicesHouseValue
1MEAN 6241.6711.44172333.33120.83317000.00
2STD 3439.991.78651241.21114.9286367.53
3N 12.0012.000012.0012.00012.00
4CORRPopulation1.000.00980.970.4390.02
5CORRSchool0.011.00000.150.6910.86
6CORREmployment0.970.15431.000.5150.12
7CORRServices0.440.69140.511.0000.78
8CORRHouseValue0.020.86310.120.7781.00
9COMMUNAL 0.980.81760.970.7980.88
10PRIORS 0.970.82230.970.7860.85
11EIGENVAL 2.731.71610.04-0.025-0.07
12UNROTATEFactor10.630.71370.710.8790.74
13UNROTATEFactor20.77-0.55520.68-0.158-0.58
14RESIDUALPopulation0.02-0.01120.010.0110.00
15RESIDUALSchool-0.010.18240.02-0.0240.01
16RESIDUALEmployment0.010.02150.03-0.006-0.02
17RESIDUALServices0.01-0.0239-0.010.2020.03
18RESIDUALHouseValue0.000.0125-0.020.0340.12
19PRETRANSFactor10.79-0.6145...
20PRETRANSFactor20.610.7889...
21PREROTATFactor10.020.90420.150.7910.94
22PREROTATFactor20.990.00060.970.415-0.00
23TRANSFORFactor10.74-0.7055...
24TRANSFORFactor20.540.8653...
25FCORRFactor11.000.2019...
26FCORRFactor20.201.0000...
27PATTERNFactor1-0.080.91840.050.7610.96
28PATTERNFactor21.00-0.09350.980.339-0.10
29RCORRFactor11.00-0.2019...
30RCORRFactor2-0.201.0000...
31REFERENCFactor1-0.080.89950.050.7450.94
32REFERENCFactor20.98-0.09160.960.332-0.10
33STRUCTURFactor10.120.89950.240.8290.94
34STRUCTURFactor20.990.09190.980.4930.09


Various results from the previous factor analysis are saved in this data set, including the initial unrotated solution (its factor pattern is saved in observations with _TYPE_=UNROTATE), the prerotated varimax solution (its factor pattern is saved in observations with _TYPE_=PREROTAT), and the oblique promax solution (its factor pattern is saved in observations with _TYPE_=PATTERN).

When PROC FACTOR reads in an input data set with TYPE=FACTOR, the observations with _TYPE_=PATTERN are treated as the initial factor pattern to be rotated by PROC FACTOR. Hence, it is important that you provide the correct initial factor pattern for PROC FACTOR to read in.

In the current example, you need to provide the unrotated solution from the preceding analysis as the input factor pattern. The following statements create a TYPE=FACTOR data set fact2 from the preceding OUTSTAT= data set fact_all:

data fact2(type=factor);
   set fact_all;
   if _TYPE_ in('PATTERN' 'FCORR') then delete;
   if _TYPE_='UNROTATE' then _TYPE_='PATTERN';
run;

In these statements, you delete observations with _TYPE_=PATTERN or _TYPE_=FCORR, which are for the promax-rotated factor solution, and change observations with _TYPE_=UNROTATE to _TYPE_=PATTERN in the new data set fact2. In this way, the initial orthogonal factor pattern matrix is saved in the observations with _TYPE_=PATTERN.

You use this new data set and rotate the initial solution to another oblique solution with the ROTATE=QUARTIMAX option, as shown in the following statements:

proc factor data=fact2 rotate=quartimax reorder;
run;

As shown in Output 38.2.16, the new rotation uses a TYPE=FACTOR input data set.

Output 38.2.16: Quartimax Rotation With Input Factor Pattern

Quartimax Rotation From a TYPE=FACTOR Data Set

The FACTOR Procedure

Input Data TypeFACTOR
N Set/Assumed in Data Set12
N for Significance Tests12


The quartimax-rotated factor pattern is displayed in Output 38.2.17.

Output 38.2.17: Quartimax-Rotated Factor Pattern

Orthogonal Transformation Matrix
 12
10.801380.59815
2-0.598150.80138

Rotated Factor Pattern
 Factor1Factor2
HouseValue0.94052-0.01933
School0.90401-0.01799
Services0.799200.39878
Population0.042820.98807
Employment0.166210.97179

Variance Explained by Each
Factor
Factor1Factor2
2.36999412.0803754


The quartimax rotation produces an orthogonal transformation matrix shown at the top of Output 38.2.17. After the transformation, the factor pattern is shown next. Compared with the varimax-rotated factor pattern (see Output 38.2.7), the quartimax-rotated factor pattern shows some differences. The loadings of HouseValue and School on Factor 1 drop only slightly in the quartimax factor pattern, while the loadings of Services, Population, and Employment on Factor 1 gain relatively larger amounts. The total variance explained by Factor 1 in the varimax-rotated solution (see Output 38.2.7) is 2.350, while it is 2.370 after the quartimax-rotation. In other words, more variable variances are explained by the first factor in the quartimax factor pattern than in the varimax factor pattern. Although not very strongly demonstrated in the current example, this illustrates a well-known property about the quartimax rotation: it tends to produce a general factor for all variables.

Another oblique rotation is now explored. The Harris-Kaiser transformation weighted by the Cureton-Mulaik technique is applied to the initial factor pattern. To achieve this, you use the ROTATE=HK and NORM=WEIGHT options in the following PROC FACTOR statement:

ods graphics on;
proc factor data=fact2 rotate=hk norm=weight reorder plots=loadings;
run;
ods graphics off;

Output 38.2.18 shows the variable weights in the rotation.

Output 38.2.18: Harris-Kaiser Rotation: Weights

Variable Weights for Rotation
PopulationSchoolEmploymentServicesHouseValue
0.959827470.939454240.997463960.121947660.94007263


While all other variables have weights at least as large as 0.93, the weight for Services is only 0.12. This means that due to its small weight, Services is not as important as the other variables for determining the rotation (transformation). This makes sense when you look at the initial unrotated factor pattern plot in Output 38.2.6. In the plot, there are two main clusters of variables, and Services does not seem to fall into either of the clusters. In order to yield a Harris-Kaiser rotation (transformation) that would gear towards to two clusters, the Cureton-Mulaik weighting essentially downweights the contribution from Services in the factor rotation.

The results of the Harris-Kaiser factor solution are displayed in Output 38.2.19, with a graphical plot of rotated loadings displayed in Output 38.2.20.

Output 38.2.19: Harris-Kaiser Rotation: Factor Correlations and Factor Pattern

Inter-Factor Correlations
 Factor1Factor2
Factor11.000000.08358
Factor20.083581.00000

Rotated Factor Pattern (Standardized Regression Coefficients)
 Factor1Factor2
HouseValue0.940480.00279
School0.903910.00327
Services0.754590.41892
Population-0.063350.99227
Employment0.061520.97885


Because the Harris-Kaiser produces an oblique factor solution, you compare the current results with that of the promax (see Output 38.2.11), which also produces an oblique factor solution. The correlation between the factors in the Harris-Kaiser solution is 0.084; this value is much smaller than the same correlation in the promax solution, which is 0.201. However, the Harris-Kaiser rotated factor pattern shown in Output 38.2.19 is more or less the same as that of the promax-rotated factor pattern shown in Output 38.2.11. Which solution would you consider to be more reasonable or interpretable?

From the statistical point of view, the Harris-Kaiser and promax factor solutions are equivalent. They explain the observed variable relationships equally well. From the simplicity point of view, however, you might prefer to interpret the Harris-Kaiser solution because the factor correlation is smaller. In other words, the factors in the Harris-Kaiser solution do not overlap that much conceptually; hence they should be more distinctive to interpret. However, in practice simplicity in factor correlations might not the only principle to consider. Researchers might actually expect to have some factors to be highly correlated based on theoretical or substantive grounds.

Although the Harris-Kaiser and the promax factor patterns are very similar, the graphical plots of the loadings from the two solutions paint slightly different pictures. The plot of the promax-rotated loadings is shown in Output 38.2.12, while the plot of the loadings for the current Harris-Kaiser solution is shown in Output 38.2.20.

Output 38.2.20: Harris-Kaiser Rotation: Factor Loading Plot

Harris-Kaiser Rotation: Factor Loading Plot


The two factor axes in the Harris-Kaiser rotated pattern (Output 38.2.20) clearly cut through the centers of the two variable clusters, while the Factor 1 axis in the promax solution lies above a variable cluster (Output 38.2.12). The reason for this subtle difference is that in the Harris-Kaiser rotation, the Services is a "loner" that has been downweighted by the Cureton-Mulaik technique (see its relatively small weight in Output 38.2.18). As a result, the rotated axes are basically determined by the two variable clusters in the Harris-Kaiser rotation.

As far as the current discussion goes, it is not recommending one rotation method over another. Rather, it simply illustrates how you could control certain types of characteristics of factor rotation through the many options supported by PROC FACTOR. Should you prefer an orthogonal rotation to an oblique rotation? Should you choose the oblique factor solution with the smallest factor correlations? Should you use a weighting scheme that would enable you to find independent variable clusters? While PROC FACTOR enables you to explore all these alternatives, you must consult advanced textbooks and published articles to get satisfactory and complete answers to these questions.