CAEFFECT Procedure
Getting Started: CAEFFECT Procedure
(View the complete code for this example.)
Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
This section illustrates some of the basic features of the CAEFFECT procedure. The example uses the School data set available in the examples for the CAUSALTRT procedure in SAS/STAT software. The data for this example are a subset of data from the National Educational Longitudinal Study of 1988 in Murnane and Willett (2011). Demographic information was collected from the students and their families, along with standardized test scores in 1988 and in follow-up years.
As in the PROC CAUSALTRT example, the question of interest in this example is what effect attending a Catholic high school has on a student’s performance in mathematics, as measured by the student’s score on the mathematics portion of the standardized tests in 1992. The effect of interest is estimated in this example by using the targeted maximum likelihood estimation (TMLE) method, which is doubly robust.
The following DATA step creates the data table School in your CAS session:
data mylib.School;
input
Income $ FatherEd $ MotherEd $ Math BaseMath Catholic $;
datalines;
Middle HighSchool HighSchool 49.77 50.27 Yes
Middle Unknown Unknown 59.84 51.52 Yes
Middle NoHighSchool Postsecondary 50.38 47.56 Yes
Middle Unknown Unknown 45.03 46.60 Yes
High HighSchool HighSchool 54.26 60.22 Yes
... more lines ...
High College College 56.60 60.15 Yes
High College SomeSecondary 56.29 59.53 Yes
High College Postsecondary 58.16 57.06 Yes
High College SomeSecondary 63.57 63.51 Yes
;
These statements assume that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.
The School data table consists of records for 5,671 high school students whose 1988 total family income was at most $75,000. The variables in the data are as follows:
BaseMath: student’s score on the mathematics portion of a standardized test in 1988Catholic: Yes, if a student attended a Catholic high school; No, otherwiseFatherEd: highest level of education completed by the student’s father, with six possible levelsIncome: classification based on total family income, with values of Low, Middle, and HighMath: student’s score on the mathematics portion of a standardized test in 1992MotherEd: highest level of education completed by the student’s mother, with six possible levels
For this example, it is assumed that the variables BaseMath, FatherEd, MotherEd, and Income make up a valid adjustment set that you can use to obtain an unbiased estimate of the effect of interest. For more information about the causal assumptions that are required for valid effect estimation, see the section Potential Outcomes and Effect Definitions.
The TMLE method adjusts for the confounding variables by using models for both the treatment variable Catholic and the outcome variable Math. PROC CAEFFECT does not fit the required outcome and treatment models. Instead, the procedure performs adjustments that are based on predictions from previously fit models, enabling you to use the modeling technique of your choice to adjust for confounders.
In this example, the model for the treatment variable is fit by using the LOGSELECT procedure, as follows:
proc logselect data=mylib.School;
class FatherEd MotherEd Income;
model Catholic( event='Yes' ) = FatherEd MotherEd Income;
output out=mylib.schoolEst pred=pTrt copyvars=(_All_);
run;
The PRED= option in the OUTPUT statement is used to name the variable pTrt, which contains the predicted event probability of attending a Catholic high school. The _ALL_ keyword is used in the COPYVARS= option to copy all variables from the input data table to the output data table schoolEst. The following DATA step creates the variable pCnt, which contains the predicted probability of not attending a Catholic high school:
data mylib.schoolEst;
set mylib.schoolEst;
pCnt = 1 - pTrt;
run;
An alternative way to obtain the predicted treatment probabilities is to save the fitted treatment model as an analytic store and use the ASTORE procedure. This approach is demonstrated in Estimation by Inverse Probability Weighting.
For estimation methods that require modeling of the outcome variable, you can provide a previously fit model for the outcome in the form of an analytic store. PROC CAEFFECT uses this analytic store to obtain the required counterfactual predictions. The saved outcome model that you provide as input must use the confounding variables and the treatment variable as predictors. For this example, PROC BART is used to fit a Bayesian additive regression trees (BART) model for the outcome variable. BART models are nonparametric and use samples of a sum-of-trees ensemble to predict the conditional mean of a response variable. For more information about BART models, see the section Details: BART Procedure in Chapter 4, BART Procedure. The following statements use PROC BART to fit an outcome model that is saved in an analytic store named bartOutMod:
proc bart data=mylib.schoolEst nMC=200 nTree=100 seed=2455;
class Income FatherEd MotherEd Catholic;
model Math = BaseMath Income FatherEd MotherEd Catholic;
store out=mylib.bartOutMod;
run;
The following statements use PROC CAEFFECT to estimate potential outcome means and their difference by using the TMLE method:
proc caeffect data=mylib.schoolEst;
treatvar Catholic;
outcomevar Math;
outcomemodel restore=mylib.bartOutMod predname=P_Math;
pom treatlev='Yes' treatprob=pTrt;
pom treatlev='No' treatprob=pCnt;
difference evtLev='Yes';
run;
When you use PROC CAEFFECT, the TREATVAR statement and at least one POM statement are required. In the TREATVAR statement, you identify the treatment variable for the analysis. You use a POM statement to specify a potential outcome of interest and to specify any variables in the input data table that contain predicted values that the estimation method requires. You specify the outcome variable and any outcome variable options by using the OUTCOMEVAR statement. The TMLE method requires the observed outcome values.
You must use the TREATLEV= option in each POM statement. You use this option to specify the level of treatment that defines the potential outcome of interest. For estimation methods that require predicted treatment probabilities, you specify the variables that contain these values by using the TREATPROB= option.
For this example, you use the OUTCOMEMODEL statement to specify the BART model that has previously been fit to the data. PROC CAEFFECT uses the saved model to compute the predicted counterfactual outcome values for the outcome variable Math that the TMLE method requires. When you use the OUTCOMEMODEL statement, you use the RESTORE= option to specify the model that is saved as an analytic store, and you use the PREDNAME= option to specify the name of the variable that contains the model predictions. If you do not know the name of the variables that are created by the analytic store, you can use the DESCRIBE statement in PROC ASTORE and check the "Output Variables" table that it creates to find the variable names. As an alternative to using the OUTCOMEMODEL statement, you can use the PREDOUT= option in the POM statement to specify variables in the input data table that contain precomputed predictions for the counterfactual outcomes. An example of how you can use the alternative syntax is provided by Estimation by Regression Adjustment.
The DIFFERENCE statement requests a difference between the potential outcome mean estimates. In this example, a single difference is requested, and the EVTLEV= option specifies that the treatment level 'Yes' is used as the event level. In general, you can specify multiple DIFFERENCE statements, and in each DIFFERENCE statement you can specify reference levels, event levels, or both.
The output from this analysis is presented in Figure 1 through Figure 4.
The "Model Information" table in Figure 1 summarizes important information about the analysis that you perform. It includes the name of the input data source, the treatment variable, the estimation method, and the name of the analytic store that is used to score the predicted counterfactual outcomes.
Figure 1: Model Information
| Model Information | |
|---|---|
| Data Source | SCHOOLEST |
| Treatment Variable | Catholic |
| Outcome Variable | Math |
| Outcome Type | Continuous |
| Estimation Method | TMLE |
| Outcome Model Store | BARTOUTMOD |
Figure 2 displays two tables. The "Number of Observations" table displays the number of observations that are used in the analysis. The "Treatment Profile" table displays information about the levels of the treatment variable that appear in the input data table and the number of observations in each level.
Figure 2: Model Information
| Number of Observations Read | 5671 |
|---|---|
| Number of Observations Used | 5671 |
| Treatment Profile | |
|---|---|
| Catholic | Total Frequency |
| No | 5079 |
| Yes | 592 |
Figure 3 also displays two tables. The "POM Estimates" table indicates the level of treatment that defines the potential outcome and the estimate of its mean. The "POM Differences" table shows the requested difference in potential outcome mean estimates. The table also includes nonprinting columns that contain the potential outcome mean estimates for the event and reference levels; these columns are accessible if you save the table as an output data set by using the ODS OUTPUT statement or if you modify the table’s template. If the necessary causal assumptions are satisfied, then the difference in the potential outcome mean estimates is an unbiased estimate of the average treatment effect (ATE). The ATE estimate of 1.467 indicates that attending a Catholic high school has a small positive effect on math scores. For more information about the assumptions that are required for valid causal effect estimation within a potential outcomes framework, see the section Potential Outcomes and Effect Definitions.
Figure 3: Potential Outcome Mean and Difference Estimates
| POM Estimates | |
|---|---|
| Treatment Level | Estimate |
| Yes | 52.37357 |
| No | 50.90671 |
| POM Differences | ||
|---|---|---|
| Treatment Levels | Difference | |
| Event | Reference | |
| Yes | No | 1.4669 |
Finally, PROC CAEFFECT displays the table shown in Figure 4, which shows the amount of time (in seconds) that it took to perform the various tasks in the analysis.
Figure 4: Procedure Timing
| Task Timing | ||
|---|---|---|
| Task | Seconds | Percent |
| Parsing | 0.03 | 1.79% |
| Data Preparation | 0.03 | 1.63% |
| Outcome Prediction | 1.53 | 90.20% |
| Targeting | 0.10 | 5.66% |
| Estimation | 0.00 | 0.13% |
| Miscellaneous | 0.00 | 0.21% |
| Cleaning Up | 0.01 | 0.36% |
| Total | 1.70 | 100.00% |