EFA Procedure

Getting Started: EFA Procedure

(View the complete code for this example.)

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

This example demonstrates how you can use the EFA procedure to perform common factor analysis and factor rotation. The example uses the jobratings data set that is described in the examples for the FACTOR procedure in SAS/STAT software. In the original data set, 103 police officers were rated by their supervisors on 14 scales (variables). The current example excludes one of these scales—the overall rating variable.

The following DATA step creates the table jobratings:

options validvarname=any;
data mylib.jobratings;
   input ('Communication Skills'n
      'Problem Solving'n
      'Learning Ability'n
      'Judgment under Pressure'n
      'Observational Skills'n
      'Willingness to Confront Problems'n
      'Interest in People'n
      'Interpersonal Sensitivity'n
      'Desire for Self-Improvement'n
      'Appearance'n
      'Dependability'n
      'Physical Ability'n
      'Integrity'n) (1.);
   datalines;
2683885387986
7475887685766
5675786377587
6786977798899
9999779887888
8989789988879
8999988989979
8779479846888
3565233514311

   ... more lines ...   

9989989989989
7665639956748
;

These statements assume that your libref is named mylib, but you can substitute any appropriately defined libref.

You conduct a common factor analysis on these variables to see what latent factors are operating behind these ratings. As part of this analysis, you want to determine the number of latent factors that are likely to explain the data.

The following statements invoke the EFA procedure:

proc efa data=mylib.jobratings rotate=varimax;
   nfactors type=parallel alpha=0.05 seed=1031 nsimulations=20000;
   nfactors type=parallel alpha=0.05 seed=1128 nsimulations=20000;
   nfactors type=parallel alpha=0.10 seed=1031 nsimulations=20000;
   nfactors type=parallel alpha=0.10 seed=1128 nsimulations=20000;
   nfactors type=eigenvalue threshold=0.5  status=inactive;
   nfactors type=eigenvalue threshold=1.0  status=inactive;
   nfactors type=proportion threshold=0.9  status=inactive;
   nfactors type=proportion threshold=0.95 status=inactive;
   nfactors type=map2 status=inactive;
   nfactors type=map4 status=inactive;
run;

In the common factor model, you always assume that observed variables are functions of underlying latent factors. (For more information about the common factor model, see the section The Common Factor Model.) The goal of factor analysis is to extract and interpret a factor pattern that describes these relationships. By default, PROC EFA uses the principal factor extraction method, with prior communalities set to the squared multiple correlations (SMC) of each variable with all the other variables. You can use the METHOD= option to specify a different factor extraction method. You can use the PRIORS= option to specify different prior communality estimates. To facilitate interpretations, the ROTATE= option specifies the varimax orthogonal factor rotation.

PROC EFA supports various criteria that you can use to suggest the number of factors to extract. You use the NFACTORS statement to specify a criterion. You can use multiple NFACTORS statements to specify multiple criteria. In this example, you determine the number of factors to extract by using multiple parallel analyses, each with 20,000 simulations. You include two different thresholds for the parallel analyses and two different random number seeds. By default, the actual number of factors that are extracted is the minimum of the numbers that are suggested by the four parallel analyses. You can use the NFACTORS=method option to change this default behavior, where method is MIN, MAX, MEDIAN, or MEAN.

In this example, you specify multiple additional NFACTORS statements, all of which include the STATUS=INACTIVE option. When a criterion has an inactive status, it is computed and the results are displayed, but it is not used to determine the final number of factors that are extracted. This can be useful for investigating the robustness of the number of extracted factors to alternative criteria. The minimum eigenvalue criterion (TYPE=EIGENVALUE) computes the number of factors to extract by comparing the eigenvalues of the reduced correlation matrix to a threshold value. The proportion of common variance criterion (TYPE=PROPORTION) computes the number of factors to extract by determining the smallest number of factors such that the cumulative proportion of common variance explained exceeds a threshold value. The minimum average partial correlation analyses (TYPE=MAP2 and MAP4) compute the number of factors by determining the number of principal components that, when partialed out, results in the smallest average squared partial (or residual) correlations.

If you know the exact number of factors that you want to extract in an analysis, you can omit the NFACTORS statement and specify the NFACTORS=n option in the PROC EFA statement, where n is the number of factors that you want to extract.

The output from the factor analysis is displayed in Figure 1 through Figure 11.

The first outputs that PROC EFA produces are the "Model Information" table and a table that summarizes the number of observations in the input data table. These are shown in Figure 1. Together, these tables summarize important information about the input data that are analyzed as well as the primary analysis options that you specified. For this analysis, the numbers of observations that are used is the same as the number that are read. These numbers would be different if some observations were discarded from the analysis. PROC EFA uses only complete cases for analysis, so a record is discarded if any field contains a missing value.

Figure 1: Model Information and Number of Observations

Model Information
Data SourceJOBRATINGS
Data TypeRaw
Prior CommunalitiesSquared Multiple Correlations
Factor ExtractionPrincipal Factor Method
Factor RotationVarimax

Number of Observations Read103
Number of Observations Used103


Next, the "Simple Statistics" table in Figure 2 summarizes the means and standard deviations of the analysis variables.

Figure 2: Simple Statistics

Simple Statistics
VariableMeanStandard
Deviation
Communication Skills6.650491.76407
Problem Solving6.631071.59035
Learning Ability6.990291.33941
Judgment under Pressure6.737861.73183
Observational Skills6.932041.76158
Willingness to Confront Problems7.291261.52516
Interest in People6.708741.89235
Interpersonal Sensitivity6.621361.76077
Desire for Self-Improvement6.572821.72980
Appearance7.000001.79869
Dependability6.825241.91704
Physical Ability7.203881.55525
Integrity7.213591.84524


The summary output for the determination of the number of factors is shown in Figure 3. For most of the specified criteria, the results suggest that two latent factors should be retained to explain the data. The minimum eigenvalue criterion, with a threshold of 0.50, and the cumulative proportion criterion, with a threshold of 0.95, suggest the existence of a third latent factor. All four parallel analysis criteria suggest two latent factors, so the final number of latent factors that the analysis suggests is two.

Figure 3: Number of Factors

Determination of Number of Factors
CriterionDescriptionFactors
1Parallel Analysis (Alpha = 0.05, Seed = 1031, Sims = 20000)2
2Parallel Analysis (Alpha = 0.05, Seed = 1128, Sims = 20000)2
3Parallel Analysis (Alpha = 0.10, Seed = 1031, Sims = 20000)2
4Parallel Analysis (Alpha = 0.10, Seed = 1128, Sims = 20000)2
5Minimum Eigenvalue (Eigenvalue > 0.50)3
6Minimum Eigenvalue (Eigenvalue > 1.00)2
7Total Proportion (Threshold = 0.90)2
8Total Proportion (Threshold = 0.95)3
9Min Ave Squared Partial Corr (MAP2)2
10Min Ave 4th-powered Partial Corr (MAP4)2
Min = 2, Max = 2, Mean = 2, Median = 2


Additional information about the minimum eigenvalue and cumulative proportion criteria is presented in the "Eigenvalues of the Reduced Correlation Matrix" table in Figure 4. After the data correlation matrix is reduced by replacing its diagonal elements with the SMC prior communality estimates, the largest eigenvalue of the resulting matrix is approximately 6.18, and the corresponding latent factor explains approximately 77% of the common variance of the input data. The second eigenvalue is approximately 1.46, and the latent factors that correspond to these first two eigenvalues explain almost 95% of the common variance. A third eigenvalue has a value of approximately 0.56, and the latent factors that correspond to these three eigenvalues explain more than 100% of the common variance. Thus, two latent factors are sufficient to explain up to 90% of the common variance in the input data, and a third factor is needed to explain more than 95% of the common variance. Similarly, a minimum eigenvalue threshold of 1.00 produces an estimate of two latent factors, whereas a third latent factor is estimated if the minimum eigenvalue threshold is 0.50. (The fourth-largest eigenvalue for the reduced correlation matrix has a value of approximately 0.28.)

Figure 4: Eigenvalues and Common Variance Explained

Eigenvalues of the Reduced Correlation Matrix
 EigenvalueDifferenceProportionCumulative
16.1776054.7153190.76600.7660
21.4622860.9018330.18130.9473
30.5604530.2809390.06951.0168
40.2795130.0476600.03471.0515
50.2318530.1611340.02871.0802
60.0707190.0748960.00881.0890
7-0.0041770.033875-0.00051.0885
8-0.0380530.047765-0.00471.0838
9-0.0858180.024381-0.01061.0731
10-0.1101990.014527-0.01371.0595
11-0.1247260.023565-0.01551.0440
12-0.1482910.058236-0.01841.0256
13-0.206527 -0.02561.0000


The parallel analysis results are summarized in the tables in Figure 5. For each of the four specified parallel analysis criteria, correlation matrices are simulated from a Wishart distribution. These simulations represent correlation matrices that might be realized for a data-generating process in which all variables are independently and identically distributed according to the standard normal distribution. The eigenvalues of the simulated correlation matrices are computed and rank-ordered, and the eigenvalues of the input data correlation matrix are compared to the distributions of the eigenvalues at each position (largest eigenvalue, second-largest eigenvalue, and so on). When an eigenvalue of the input data correlation matrix exceeds the eigenvalue at some critical point (for example, the 95th percentile), this is understood to indicate the presence of a latent factor in the data. In all four parallel analyses in this example, the largest and second-largest eigenvalues of the input data correlation matrix exceed the simulated critical values. Thus all four parallel analyses suggest that two latent factors should be retained.

Figure 5: Parallel Analysis

Parallel Analysis (Criterion 1)
ComponentSimulated
Crit Val
Observed
Eigenvalue
11.79886.5474*
21.58091.7727*
31.43261.0052
41.31120.7431
51.20840.6783
61.11530.4514
71.03080.3822
80.95100.2978
90.87320.2744
100.79880.2623
110.72340.2446
120.64820.1978
130.56520.1427
* Retained Dimension (Obs > Crit)

Parallel Analysis (Criterion 2)
ComponentSimulated
Crit Val
Observed
Eigenvalue
11.79816.5474*
21.58121.7727*
31.43031.0052
41.31240.7431
51.20810.6783
61.11520.4514
71.02880.3822
80.95030.2978
90.87280.2744
100.79860.2623
110.72350.2446
120.64790.1978
130.56510.1427
* Retained Dimension (Obs > Crit)

Parallel Analysis (Criterion 3)
ComponentSimulated
Crit Val
Observed
Eigenvalue
11.75996.5474*
21.55281.7727*
31.41061.0052
41.29290.7431
51.19220.6783
61.10160.4514
71.01680.3822
80.93720.2978
90.86010.2744
100.78420.2623
110.70950.2446
120.63280.1978
130.54990.1427
* Retained Dimension (Obs > Crit)

Parallel Analysis (Criterion 4)
ComponentSimulated
Crit Val
Observed
Eigenvalue
11.76006.5474*
21.55331.7727*
31.41011.0052
41.29430.7431
51.19210.6783
61.10130.4514
71.01560.3822
80.93660.2978
90.85940.2744
100.78450.2623
110.70970.2446
120.63370.1978
130.54950.1427
* Retained Dimension (Obs > Crit)


The minimum average partial correlation analysis (MAP) results are summarized in the "Average Partial Correlations Controlling Principal Components" table in Figure 6. In this example, analyses are performed that use both squared (MAP2) and fourth-powered (MAP4) partial correlations. These two analyses are summarized in separate columns in the table. The number of latent factors that are estimated by a minimum average partial correlation analysis is the number of principal components that, when partialed out of the correlation matrix, produce the smallest average squared or fourth-powered partial correlations. As shown in Figure 6, the minimum value is obtained for both the MAP2 and the MAP4 analysis when two principal components are partialed out.

Figure 6: Minimum Average Partial Correlation Analysis

Average Partial Correlations Controlling
Principal Components
Number of
Components
SquaredFourth-
Powered
00.2290760.069057
10.0636550.011875
20.036328*0.003287*
30.0429700.004975
40.0571020.009267
50.0653260.014253
60.0833450.022845
70.1157060.037743
80.1534230.050831
90.2330630.120901
100.3091450.167278
110.5441630.416722
121.0000001.000000
* MAP = Minimum Values in Columns


Now that the number of factors is determined, it is time to extract and rotate a factor pattern that corresponds to a two-factor solution. For this analysis, the prior communality estimates are set to the squared multiple correlations. These are shown in Figure 7.

Figure 7: Prior Communalities

Prior Communality Estimates
Communication SkillsProblem SolvingLearning AbilityJudgment under
Pressure
Observational SkillsWillingness to
Confront Problems
Interest in PeopleInterpersonal SensitivityDesire for Self-ImprovementAppearanceDependabilityPhysical AbilityIntegrity
0.629813940.586574310.610098710.637660210.671875830.647798050.756415190.755848910.574601760.455053040.634490450.422453240.68195454


Figure 8 displays the initial factor pattern matrix. The factor pattern matrix represents standardized regression coefficients for predicting the variables by using the extracted factors. Because the initial factors are uncorrelated, the pattern matrix is also equal to the correlations between variables and the common factors.

The pattern matrix suggests that Factor1 represents general ability. All loadings for Factor1 in the Factor Pattern are at least 0.5. Factor2 consists of high positive loadings on certain task-related skills (Willingness to Confront Problems, Observational Skills, and Learning Ability) and high negative loadings on some interpersonal skills (Interpersonal Sensitivity, Interest in People, and Integrity). This factor measures individuals’ relative strength in these skills. Theoretically, individuals with high positive scores on this factor would exhibit better task-related skills than interpersonal skills. Individuals with high negative scores would exhibit better interpersonal skills than task-related skills. Individuals with scores near zero would have those skills balanced.

Figure 8: Initial Factor Pattern

Factor Pattern
 Factor1Factor2
Communication Skills0.754410.07707
Problem Solving0.685900.08026
Learning Ability0.659040.34808
Judgment under Pressure0.73391-0.21405
Observational Skills0.690390.45292
Willingness to Confront Problems0.664580.47460
Interest in People0.70770-0.53427
Interpersonal Sensitivity0.64668-0.61284
Desire for Self-Improvement0.738200.12506
Appearance0.571880.20052
Dependability0.79475-0.04516
Physical Ability0.512850.10251
Integrity0.74906-0.35091


Figure 9 displays the proportion of variance explained by each factor and the final communality estimates, including the total communality. The final communality estimates are the proportion of variance of the variables accounted for by the common factors. When the factors are orthogonal, the final communalities are calculated by taking the sum of squares of each row of the factor pattern matrix.

Figure 9: Initial Solution: Variance Explained and Final Communality Estimates

Variance Explained by Each
Factor
Factor1Factor2
6.17760551.4622860

Final Communality Estimates: Total = 7.639892
Communication SkillsProblem SolvingLearning AbilityJudgment under
Pressure
Observational SkillsWillingness to
Confront Problems
Interest in PeopleInterpersonal SensitivityDesire for Self-ImprovementAppearanceDependabilityPhysical AbilityIntegrity
0.575077750.476901280.555488540.584443010.681769860.666904090.786285880.793765210.560579570.367249940.633672860.273527120.68422640


Figure 10 displays the results of the varimax rotation of the two extracted factors and the corresponding orthogonal transformation matrix. The rotated factor pattern matrix is calculated by postmultiplying the original factor pattern matrix by the transformation matrix.

The rotated factor pattern matrix is somewhat easier to interpret. If a magnitude of at least 0.5 is required to indicate a salient variable-factor relationship, Factor1 now measures job enthusiasm and cognitive skills (Observational Skills, Willingness to Confront Problems, Learning Ability, Desire for Self-Improvement, Communication Skills, Problem Solving, Dependability, and Appearance) and Factor2 represents interpersonal skills (Interpersonal Sensitivity, Interest in People, Integrity, Judgment under Pressure, and Dependability).

Figure 10: Transformation Matrix and Rotated Factor Pattern

Orthogonal Transformation Matrix
 12
10.751280.65999
20.65999-0.75128

Rotated Factor Pattern
 Factor1Factor2
Communication Skills0.617640.44000
Problem Solving0.568270.39239
Learning Ability0.724850.17346
Judgment under Pressure0.410090.64519
Observational Skills0.817590.11538
Willingness to Confront Problems0.812510.08206
Interest in People0.179060.86846
Interpersonal Sensitivity0.081370.88721
Desire for Self-Improvement0.637130.39325
Appearance0.561980.22679
Dependability0.567270.55846
Physical Ability0.452950.26147
Integrity0.331160.75800


To make the rotated factor pattern matrix even easier to interpret, you can use the REORDER option to rearrange the rows of the factor pattern. When you specify the REORDER option, variables whose highest absolute loading (reference structure loading for oblique rotations) is on the first factor are displayed first, from largest to smallest loading, followed by variables whose highest absolute loading is on the second factor, and so on. For an example of how to use the REORDER option, see Example 10.2: Principal Factor Analysis. You can also specify the FUZZ=0.5 option, which causes all the entries of the factor pattern smaller than 0.5 in absolute value to be printed as missing. For an example of how to use the FUZZ= option, see Example 10.4: Extracting and Interpreting Solutions with Maximum Likelihood Factor Analysis.

Figure 11 displays the variance explained by each factor and the final communality estimates after the orthogonal rotation. Even though the variances explained by the rotated factors are different from those of the unrotated factors (compare to Figure 9), the cumulative variance explained by the common factors remains the same. Note also that the final communalities for variables, as well as the total communality, remain unchanged after rotation. Although rotating a factor solution does not increase or decrease the statistical quality of the factor model, it can simplify the interpretations of the factors and redistribute the variance explained by the factors.

Figure 11: Rotated Solution: Variance Explained and Final Communality Estimates

Variance Explained by Each
Factor
Factor1Factor2
4.12367933.5162122

Final Communality Estimates: Total = 7.639892
Communication SkillsProblem SolvingLearning AbilityJudgment under
Pressure
Observational SkillsWillingness to
Confront Problems
Interest in PeopleInterpersonal SensitivityDesire for Self-ImprovementAppearanceDependabilityPhysical AbilityIntegrity
0.575077750.476901280.555488540.584443010.681769860.666904090.786285880.793765210.560579570.367249940.633672860.273527120.68422640


Last updated: June 22, 2026