The SANDWICH Procedure

Getting Started: SANDWICH Procedure

In health economics, large correlated data are common. These data usually contain information that is observed about large numbers of individuals, hospitals, or employers. The observations within individuals, hospitals, or employers are assumed to be correlated. There are two approaches to modeling such correlated data. One approach is to use random effects to introduce correlation in model estimation. Another approach is to use a robust variance estimator to adjust for correlation after the model is estimated. PROC SANDWICH provides a tool for analyzing large correlated data by using a robust variance estimator.

This example analyzes simulated health-care spending data by using a robust variance estimator. Suppose a number of health insurance choices are provided to 20,000 individuals each year over a period of four years. The goal of this analysis is to study the effect of choices of deductible and out-of-pocket payments on health care spending.

The following DATA step simulates 80,000 observations of annual health care spending from a linear model. The predictor variables are Individual, Year, Income, FamilySize, and choices of deductible and out-of-pocket payments for individual and family. DeductibleInd and DeductibleFam denote the five tiers of deductibles for an individual and a family, respectively. OutOfPocketInd and OutOfPocketFam denote the six tiers of out-of-pocket payments for an individual and a family, respectively. This DATA step assumes that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.


%let nDeductibleInd=5;
%let nDeductibleFam=5;
%let nOutOfPocketInd=6;
%let nOutOfPocketFam=6;
%let nFamSize=5;

data mycas.insure;
   call streaminit(1234);
   do Individual = 1 to 20000;
      rInd = 5 *rand("normal");
      Income = 3 + 2 *rand("normal");
      FamilySize = ceil(&nFamSize * rand("uniform"));
      do Year = 1 to 4;
         DeductibleInd = ceil(&nDeductibleInd * rand("uniform"));
         DeductibleFam = ceil(&nDeductibleFam * rand("uniform"));
         OutOfPocketInd= ceil(&nOutOfPocketInd * rand("uniform"));
         OutOfPocketFam= ceil(&nOutOfPocketFam * rand("uniform"));
         err   = 2*rand("normal");
         spend = 350 + Year + rInd + 3.2*Income + FamilySize +
                 DeductibleInd + DeductibleFam + OutOfPocketInd + OutOfPocketFam + err;
         output;
      end;
   end;
run;

Figure 1 below shows the spaghetti plot of health-care spending of 10 individuals over the four-year period. You can see that the intercepts of the trajectories of health-care spending vary from individual to individual. The variability from individual to individual is an indication of correlation within individuals.

Figure 1: Spaghetti Plot of Ten Individuals

Spaghetti Plot of Ten Individuals


The following SANDWICH procedure statements fit a linear regression model to the annual health-care spending data:

ods select ModelInfo Dimensions ModelAnova;
proc sandwich data=mycas.insure sparse;
   class individual year familysize DeductibleInd DeductibleFam
         OutOfPocketInd OutOfPocketFam;
   clusters individual;
   model spend =  Year Income|DeductibleInd|OutOfPocketInd
                  Income|FamilySize|OutOfPocketFam|DeductibleFam/ss3;
run;

In the preceding statements, the cluster effect, Individual, accounts for correlations among observations from the same individual. The SPARSE option tells PROC SANDWICH to invoke sparse matrix methods. All two-way and the three-way interaction effects among Income, DeductibleInd, and OutOfPocketInd are included to study how the effects of individual deductible, out-of-pocket payment, and income change with each other. Also, all two-way, all three-way, and the four-way interaction effects among Income, FamilySize, OutOfPocketFam, DeductibleFam are included to account for the interaction between family deductible, family out-of-pocket payment, family size, and income.

These interaction effects in addition to the main effects add up to a total number of 590 parameters, leading to a sum of squares and crossproducts matrix bold upper X prime bold upper X that is of dimension 590 times 590. Analyzing these data in PROC SURVEYREG requires the product bold upper X prime bold upper X for all 20,000 clusters. It takes about 2 GB of memory to store these matrices and a great deal of time to manipulate them. Such time-consuming and resource-draining problems can benefit from the distributed computing algorithms that are implemented in PROC SANDWICH.

Figure 2 shows the basic information of the model.

Figure 2: Model Information

The SANDWICH Procedure

Model Information
Data SourceINSURE
Response Variablespend
Design Matrix MethodSparse


Figure 3 shows that there are 23 effects in this model and these effects have 590 parameters.

Figure 3: Model Dimensions

Dimensions
Number of Effects23
Number of Parameters590
Number of Clusters20000


Figure 4 shows estimates of some of the parameters.

Figure 4: Part of the Parameter Estimates Table

The SANDWICH Procedure

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValuePr > |t|
Intercept1380.8571160.435528874.47<.0001
Year 11-2.9924920.020482-146.10<.0001
Year 21-2.0400890.020680-98.65<.0001
Year 31-0.9944460.020661-48.13<.0001
Year 400...
Income13.2578560.12070926.99<.0001
DeductibleInd 11-4.1216040.257734-15.99<.0001
DeductibleInd 21-2.8708250.261725-10.97<.0001
DeductibleInd 31-1.7941500.262846-6.83<.0001
DeductibleInd 41-0.8988970.261390-3.440.0006
DeductibleInd 500...
Income * DeductibleInd 110.0286170.0712370.400.6879
Income * DeductibleInd 21-0.0266240.070959-0.380.7075
Income * DeductibleInd 31-0.0662860.072775-0.910.3624
Income * DeductibleInd 41-0.0183520.071792-0.260.7982
Income * DeductibleInd 500...
 .....
 .....
 .....


In the "Parameter Estimates" table, the t statistics of the parameters are computed using the sandwich variance estimator, which accounts for clustering. The degrees of freedom for the t test is the number of clusters minus one. The t test results of five individual deductibles categories indicate that they are significant. The estimates of these five individual categories show how annual health-care spending decreases with the increase of the individual deductible.

Part of the "Type III Model ANOVA" table is shown in Figure 5. By construction, the true model consists of the main effects Year, Income, FamilySize, DeductibleInd, DeductibleFam, OutOfPocketInd, and OutOfPocketFam. F test results in Figure 5 show that only the effects in the true model are significant. The F statistics of the parameters are computed using the sandwich variance estimator that accounts for clustering. The degrees of freedom for the F test is the number of clusters minus one.

Figure 5: Part of the Type III Model ANOVA Table

The SANDWICH Procedure

Type III Model ANOVA
EffectDFSum of
Squares
Mean
Square
F ValuePr > F
Year3238937964.175737964.18<.0001
Income1323003230032299.8<.0001
DeductibleInd41774.85622443.71406443.71<.0001
Income*DeductibleInd45.179701.294921.290.2694
OutOfPocketInd52480.20898496.04180496.04<.0001
Income*OutOfPocketInd53.193000.638600.640.6703
DeductibleInd*OutOfPocketInd2017.044410.852220.850.6501
Income*DeductibleInd*OutOfPocketInd2027.232651.361631.360.1290
FamilySize4574.71260143.67815143.68<.0001
 .....
 .....
 .....


Last updated: September 17, 2021