The MIXED Procedure

Example 78.7 Influence in Heterogeneous Variance Model

(View the complete code for this example.)

In this example from Snedecor and Cochran (1980, p. 216), a one-way classification model with heterogeneous variances is fit. The data, shown in the following DATA step, represent amounts of different types of fat absorbed by batches of doughnuts during cooking, measured in grams.

data absorb;
   input FatType Absorbed @@;
   datalines;
 1 164  1 172  1 168  1 177  1 156  1 195
 2 178  2 191  2 197  2 182  2 185  2 177
 3 175  3 193  3 178  3 171  3 163  3 176
 4 155  4 166  4 149  4 164  4 170  4 168
;

The statistical model for these data can be written as

where is the amount of fat absorbed by the jth batch of the ith fat type, and denotes the fat-type effects. A quick glance at the data suggests that observations 6, 9, 14, and 21 might be influential on the analysis, because these are extreme observations for the respective fat types.

The following SAS statements fit this model and request influence diagnostics for the fixed effects and covariance parameters. ODS Graphics is used to create plots of the influence diagnostics in addition to the tabular output. The ESTIMATES suboption requests plots of "leave-one-out" estimates for the fixed effects and group variances.

ods graphics on;

proc mixed data=absorb asycov;
   class FatType;
   model Absorbed = FatType / s
                    influence(iter=10 estimates);
   repeated / group=FatType;
   ods output Influence=inf;
run;

ods graphics off;

The "Influence" table is output to the SAS data set inf so that parameter estimates can be printed subsequently. Results from this analysis are shown in Output 78.7.1.

Output 78.7.1: Heterogeneous Variance Analysis

The Mixed Procedure

Model Information
Data SetWORK.ABSORB
Dependent VariableAbsorbed
Covariance StructureVariance Components
Group EffectFatType
Estimation MethodREML
Residual Variance MethodNone
Fixed Effects SE MethodModel-Based
Degrees of Freedom MethodBetween-Within

Covariance Parameter Estimates
Cov ParmGroupEstimate
ResidualFatType 1178.00
ResidualFatType 260.4000
ResidualFatType 397.6000
ResidualFatType 467.6000

Solution for Fixed Effects
EffectFatTypeEstimateStandard
Error
DFt ValuePr > |t|
Intercept 162.003.35662048.26<.0001
FatType110.00006.3979201.560.1337
FatType223.00004.6188204.98<.0001
FatType314.00005.2472202.670.0148
FatType40....


The fixed-effects solutions correspond to estimates of the following parameters:

You can easily verify that these estimates are simple functions of the arithmetic means in the groups. For example, , , and so forth. The covariance parameter estimates are the sample variances in the groups and are uncorrelated.

The variances in the four groups are shown in the "Covariance Parameter Estimates" table (Output 78.7.1). The estimated variance in the first group is two to three times larger than the variance in the other groups.

Output 78.7.2: Asymptotic Variances of Group Variance Estimates

Asymptotic Covariance Matrix of Estimates
RowCov ParmCovP1CovP2CovP3CovP4
1Residual12674   
2Residual 1459.26  
3Residual  3810.30 
4Residual   1827.90


In groups where the residual variance estimate is large, the precision of the estimate is also small (Output 78.7.2).

The following statements print the "leave-one-out" estimates for fixed effects and covariance parameters that were written to the inf data set with the ESTIMATES suboption (Output 78.7.3):

proc print data=inf label;
   var parm1-parm5 covp1-covp4;
run;

Output 78.7.3: Leave-One-Out Estimates

ObsInterceptFatType 1FatType 2FatType 3FatType 4Residual FatType 1Residual FatType 2Residual FatType 3Residual FatType 4
1162.0011.60023.00014.0000203.3060.40097.6067.600
2162.0010.00023.00014.0000222.4760.40097.6067.600
3162.0010.80023.00014.0000217.6860.40097.6067.600
4162.009.00023.00014.0000214.9960.40097.6067.600
5162.0013.20023.00014.0000145.7060.40097.6067.600
6162.005.40023.00014.000063.8060.40097.6067.600
7162.0010.00024.40014.0000178.0060.79597.6067.600
8162.0010.00021.80014.0000178.0064.69197.6067.600
9162.0010.00020.60014.0000178.0032.29697.6067.600
10162.0010.00023.60014.0000178.0072.79797.6067.600
11162.0010.00023.00014.0000178.0075.49097.6067.600
12162.0010.00024.60014.0000178.0056.28597.6067.600
13162.0010.00023.00014.2000178.0060.400121.6867.600
14162.0010.00023.00010.6000178.0060.40035.3067.600
15162.0010.00023.00013.6000178.0060.400120.7967.600
16162.0010.00023.00015.0000178.0060.400114.5067.600
17162.0010.00023.00016.6000178.0060.40071.3067.600
18162.0010.00023.00014.0000178.0060.400121.9867.600
19163.408.60021.60012.6000178.0060.40097.6069.799
20161.2010.80023.80014.8000178.0060.40097.6079.698
21164.607.40020.40011.4000178.0060.40097.6033.800
22161.6010.40023.40014.4000178.0060.40097.6083.292
23160.4011.60024.60015.6000178.0060.40097.6065.299
24160.8011.20024.20015.2000178.0060.40097.6073.677


The graphical displays in Output 78.7.4 and Output 78.7.5 are created when ODS Graphics is enabled. For general information about ODS Graphics, see Chapter 21: Statistical Graphics Using ODS. For specific information about the graphics available in the MIXED procedure, see the section ODS Graphics.

Output 78.7.4: Fixed-Effects Deletion Estimates

 Fixed-Effects Deletion Estimates


Output 78.7.5: Covariance Parameter Deletion Estimates

 Covariance Parameter Deletion Estimates


The estimate of the intercept is affected only when observations from the last group are removed. The estimate of the "FatType 1" effect reacts to removal of observations in the first and last group (Output 78.7.4).

While observations can affect one or more fixed-effects solutions in this model, they can affect only one covariance parameter, the variance in their group (Output 78.7.5). Observations 6, 9, 14, and 21, which are extreme in their group, reduce the group variance considerably.

Diagnostics related to residuals and predicted values are printed with the following statements:

proc print data=inf label;
   var observed predicted residual pressres
       student Rstudent;
run;

Output 78.7.6: Residual Diagnostics

ObsObserved
Value
Predicted MeanResidualPRESS ResidualInternally Studentized
Residual
Externally Studentized
Residual
1164172.0-8.000-9.600-0.6569-0.6146
2172172.00.0000.0000.00000.0000
3168172.0-4.000-4.800-0.3284-0.2970
4177172.05.0006.0000.41050.3736
5156172.0-16.000-19.200-1.3137-1.4521
6195172.023.00027.6001.88853.1544
7178185.0-7.000-8.400-0.9867-0.9835
8191185.06.0007.2000.84570.8172
9197185.012.00014.4001.69142.3131
10182185.0-3.000-3.600-0.4229-0.3852
11185185.00.000-0.0000.00000.0000
12177185.0-8.000-9.600-1.1276-1.1681
13175176.0-1.000-1.200-0.1109-0.0993
14193176.017.00020.4001.88503.1344
15178176.02.0002.4000.22180.1993
16171176.0-5.000-6.000-0.5544-0.5119
17163176.0-13.000-15.600-1.4415-1.6865
18176176.00.0000.0000.00000.0000
19155162.0-7.000-8.400-0.9326-0.9178
20166162.04.0004.8000.53290.4908
21149162.0-13.000-15.600-1.7321-2.4495
22164162.02.0002.4000.26650.2401
23170162.08.0009.6001.06591.0845
24168162.06.0007.2000.79940.7657


Observations 6, 9, 14, and 21 have large studentized residuals (Output 78.7.6). That the externally studentized residuals are much larger than the internally studentized residuals for these observations indicates that the variance estimate in the group shrinks when the observation is removed. Also important to note is that comparisons based on raw residuals in models with heterogeneous variance can be misleading. Observation 5, for example, has a larger residual but a smaller studentized residual than observation 21. The variance for the first fat type is much larger than the variance in the fourth group. A "large" residual is more "surprising" in the groups with small variance.

A measure of the overall influence on the analysis is the (restricted) likelihood distance, shown in Output 78.7.7. Observations 6, 9, 14, and 21 clearly displace the REML solution more than any other observations.

Output 78.7.7: Restricted Likelihood Distance

 Restricted Likelihood Distance


The following statements list the restricted likelihood distance and various diagnostics related to the fixed-effects estimates (Output 78.7.8):

proc print data=inf label;
   var leverage observed CookD DFFITS CovRatio RLD;
run;

Output 78.7.8: Restricted Likelihood Distance and Fixed-Effects Diagnostics

ObsLeverageObserved
Value
Cook's DDFFITSCOVRATIORestr. Likelihood
Distance
10.1671640.02157-0.274871.37060.1178
20.1671720.00000-0.000001.49980.1156
30.1671680.00539-0.132821.46750.1124
40.1671770.008430.167061.44940.1117
50.1671560.08629-0.649380.98220.5290
60.1671950.178311.410690.43015.8101
70.1671780.04868-0.439821.20780.1935
80.1671910.035760.365461.28530.1451
90.1671970.143051.034460.64162.2909
100.1671820.00894-0.172251.44630.1116
110.1671850.00000-0.000001.49980.1156
120.1671770.06358-0.522391.11830.2856
130.1671750.00061-0.044411.49610.1151
140.1671930.177661.401750.43405.7044
150.1671780.002460.089151.48510.1139
160.1671710.01537-0.228921.40780.1129
170.1671630.10389-0.754230.87660.8433
180.1671760.000000.000001.49980.1156
190.1671550.04349-0.410471.23900.1710
200.1671660.014200.219501.41480.1124
210.1671490.15000-1.095450.60002.7343
220.1671640.003550.107361.47860.1133
230.1671700.056800.485001.15920.2383
240.1671680.031950.342451.30790.1353


In this example, observations with large likelihood distances also have large values for Cook’s D and values of CovRatio far less than one (Output 78.7.8). The latter indicates that the fixed effects are estimated more precisely when these observations are removed from the analysis.

The following statements print the values of the D statistic and the CovRatio for the covariance parameters:

proc print data=inf label;
   var iter CookDCP CovRatioCP;
run;

The same conclusions as for the fixed-effects estimates hold for the covariance parameter estimates. Observations 6, 9, 14, and 21 change the estimates and their precision considerably (Output 78.7.9, Output 78.7.10). All iterative updates converged within at most four iterations.

Output 78.7.9: Covariance Parameter Diagnostics

ObsIterationsCook's D CovParmsCOVRATIO CovParms
130.050501.6306
230.156031.9520
330.124261.8692
430.107961.8233
540.082320.8375
641.029090.1606
710.000111.2662
820.012621.4335
930.541260.3573
1030.105311.8156
1130.156031.9520
1220.011601.0849
1330.152231.9425
1441.018650.1635
1530.141111.9141
1630.074941.7203
1730.181540.6671
1830.156031.9520
1920.002651.3326
2030.080081.7374
2110.625000.3125
2230.134721.8974
2320.002901.1663
2420.020201.4839


Output 78.7.10 displays the standard panel of influence diagnostics that is obtained when influence analysis is iterative. The Cook’s D and CovRatio statistics are displayed for each deletion set for both fixed-effects and covariance parameter estimates. This provides a convenient summary of the impact on the analysis for each deletion set, since Cook’s D statistic measures impact on the estimates and the CovRatio statistic measures impact on the precision of the estimates.

Output 78.7.10: Influence Diagnostics

 Influence Diagnostics


Observations 6, 9, 14, and 21 have considerable impact on estimates and precision of fixed effects and covariance parameters. This is not necessarily the case. Observations can be influential on only some aspects of the analysis, as shown in the next example.