CAEFFECT Procedure
Estimation Methods
PROC CAEFFECT implements several methods of estimating potential outcome means and their differences. The causal interpretation of these estimates depends on the plausibility of the causal assumptions that are described in the section Potential Outcomes and Effect Definitions.
These estimation methods differ primarily in how they make statistical adjustments for confounding variables. In general, you can adjust for confounding variables by modeling the treatment variable, the outcome variable, or both. PROC CAEFFECT is a model-agnostic tool for effect estimation. That is, the procedure itself does not support modeling of the treatment and outcome variables. Instead, as input to the procedure, you provide the quantities necessary for your chosen estimation method. Thus, you can adjust for the effects of confounding variables by using predictions from a previously fit model for the treatment variable, a previously fit model for the outcome variable, or models for both the treatment and outcome variables. The required inputs for each method that PROC CAEFFECT supports are summarized in Table 2.
Table 2: Required Inputs of Different Estimation Methods
| Method | Treatment Model | Outcome Model | Observed Outcome |
|---|---|---|---|
| Inverse probability weighting | X | X | |
| Regression adjustment | X | ||
| Augmented inverse probability weighting | X | X | X |
| Targeted maximum likelihood estimation | X | X | X |
The inverse probability weighting (IPW) method requires a model for the treatment variable. To use this method, you must specify variables in the input data table that contain the precomputed treatment assignment probabilities. For an example of how you can perform IPW estimation, see Estimation by Inverse Probability Weighting. For more information about estimation by inverse probability weighting, see the section Inverse Probability Weighting.
The regression adjustment method requires a model for the outcome variable. To use this method, you have two choices for the input that you must provide: you can specify variables in the input data table that contain the predicted potential outcomes for each specified level of treatment, or you can provide a previously fit model in the form of an analytic store by using the OUTCOMEMODEL statement. In the second case, the procedure uses the analytic store to compute the relevant predicted outcome values. When you provide an analytic store as input, the model must include the specified treatment variables as a predictor, and all the required predictor variables must appear in the input data table. For an example that shows how you can perform estimation by regression adjustment by providing the procedure with either an analytic store or precomputed predicted values, see Estimation by Regression Adjustment. For more information about estimation by regression adjustment, see the section Regression Adjustment.
The remaining two methods, augmented inverse probability weighting (AIPW) and targeted maximum likelihood estimation (TMLE), are so-called doubly robust estimation methods. An estimation method is considered doubly robust if it provides unbiased estimates whenever at least one of the models is correctly specified. For these methods, you must specify a model for both the treatment variable and the outcome variable. For an example of how you can use PROC CAEFFECT to perform doubly robust estimation, see Estimation by Doubly Robust Methods. For more information about estimation by augmented inverse probability weighting, see the section Augmented Inverse Probability Weighting. For more information about targeted maximum likelihood estimation, see the section Targeted Maximum Likelihood Estimation.
Supported Variable Types
PROC CAEFFECT supports the estimation of treatment effects for categorical treatment variables and for continuous, categorical, or binomial outcomes. In the case of categorical outcomes, you must specify a designated event level by using the EVENT= option in the OUTCOMEVAR statement, and the outcome model (specified as either predicted potential outcomes or an analytic store) must represent the probability of that event level. Similarly, for binomial data, the outcome model must represent the event probability. If you are modeling binomial data and the estimation method that you use requires observed outcome values, you must use the event/trial syntax in the OUTCOMEVAR statement. This syntax requires that you specify two variables that contain the counts from a binomial experiment. The event variable contains the number of positive responses, or events, and the trial variable contains the number of trials that are conducted.
Inverse Probability Weighting
When you estimate potential outcome means by inverse probability weighting (IPW), the observed outcomes for each subject are assigned weights inversely proportional to the probability of having received the observed level of treatment. Incorporating the IPW weights creates a pseudopopulation that corrects for the bias that the confounding variables introduce. You can obtain unbiased estimates of the potential outcome means from the weighted pseudopopulation. For more information about IPW estimation, see Lunceford and Davidian (2004), Hernán and Robins (2020), and references therein.
When estimating potential outcome means by IPW, you must provide as input
variables that contain the predicted treatment probabilities of a subject being assigned to the specified treatment levels. You specify the variables that contain these precomputed probabilities by using the TREATPROB= option in each potential outcome specification. For an example of using PROC CAEFFECT to implement the IPW estimation method, see Estimation by Inverse Probability Weighting.
The following list defines components of the estimating equations that you solve to obtain the IPW estimates of the potential outcome means:
The IPW estimates of the potential outcome means solve the equations
where
The solution to these equations has a closed-form expression given by
You can also use the IPW estimation method to estimate potential outcome means conditional on subjects receiving a particular treatment assignment. You specify the level of treatment to condition on by using the CONDEVENT= option in the TREATVAR statement. The level of treatment that you specify for the conditioning event must also be used in a potential outcome specification in a POM statement.
Let denote the level of treatment that you specify to condition on for computing conditional potential outcome mean estimates. The IPW estimates of the potential outcome means conditional on
solve a similar set of estimating equations, where the weights that are used to estimate the unconditional means (one over the observed treatment assignment probability) are multiplied by
, the probability of being assigned to the treatment level
. The score equation for each subject is then given by
The solution to the estimating equations for the conditional potential outcome means is then given by
If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the IPW estimation method, an observation from the input data table is not used if any of the following conditions occur:
A predicted treatment assignment probability is missing.
The value of the outcome variable is missing.
The observed treatment level is not used in a potential outcome specification.
The value of a frequency or weight variable, if specified, is missing or nonpositive.
Regression Adjustment
When you estimate potential outcome means by regression adjustment, you use predicted potential outcome values that are obtained from a model of the outcome variable. The predicted potential outcome values for each subject are obtained by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest and then by evaluating the outcome model for each treatment assignment in turn. The averages of the predicted values are used to estimate the average of the potential outcomes. For more information about estimation by regression adjustment, see Hernán and Robins (2020) and references therein.
To estimate potential outcome means by regression adjustment, you must provide as input either
variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the regression adjustment estimation method, see Estimation by Regression Adjustment.
To describe the estimating equations, you solve to obtain the regression adjustment estimates of the potential outcome means. Let denote the predicted counterfactual outcome value that you obtain by replacing the observed treatment assignment with
. The regression adjustment estimates of the potential outcome means solve the equations
where
The solution to these equations has a closed-form expression given by
You can also use the regression adjustment estimation method to estimate potential outcome means conditional on subjects receiving a particular treatment assignment. You specify the level of treatment to condition on by using the CONDEVENT= option in the TREATVAR statement. The level of treatment that you specify for the conditioning event must also be used in a potential outcome specification in a POM statement.
The regression adjustment estimates of the potential outcome means conditional on solve a similar set of estimating equations, where predicted counterfactual values are averaged only for subjects that have the treatment assignment that you are conditioning on. The following list defines components of the estimating equations that you solve to obtain the regression adjustment estimates of the potential outcome means conditional on a treatment level:
The score equation for each subject is then given by
The solution to these equations has a closed-form expression given by
If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the regression adjustment estimation method, an observation from the input data table is not used if either of the following conditions occurs:
A predicted potential outcome value is missing.
The value of a frequency or weight variable, if specified, is missing or nonpositive.
In addition to these conditions, if you are estimating potential outcome means conditional on a particular treatment assignment, only observations that have the specified treatment assignment are used.
Augmented Inverse Probability Weighting
When you estimate potential outcome means by augmented inverse probability weighting (AIPW), predictions from a model for the outcome variable and predictions from a model for the treatment variable are combined to form a doubly robust estimation method. An estimation method is said to be doubly robust if it provides unbiased estimates whenever at least one of the models that you use is correctly specified. The model for the outcome variable is used in the AIPW estimation method to predict potential outcome values for each subject by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest. The model for the treatment variable is used to compute weights that are inversely proportional to the probability that a subject has the observed treatment assignment. For more information about the AIPW estimation method, see Lunceford and Davidian (2004), Hernán and Robins (2020), and references therein.
To estimate potential outcome means by AIPW, you must provide as input
variables that contain the predicted probability of a subject being assigned to the specified treatment levels. You specify the variables that contain the precomputed treatment assignment probabilities by using the TREATPROB= option in each potential outcome specification. Also, you must specify either
variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the AIPW estimation method, see Estimation by Doubly Robust Methods.
The following list defines components of the estimating equations that you solve to obtain the AIPW estimates of the potential outcome means:
T denotes the treatment variable.
denotes the predicted potential outcome value that you obtain by replacing the observed treatment assignment with
.
denotes the indicator function for subject i being assigned to treatment level t.
denotes the conditional probability of subject i being assigned to treatment level t; that is,
.
The AIPW estimates of the potential outcome means solve the equations
where
The solution to these equations has a closed-form expression given by
If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the AIPW estimation method, an observation from the input data table is not used if any of the following conditions occur:
A predicted treatment assignment probability is missing.
The value of the outcome variable is missing.
A predicted counterfactual value is missing.
The value of a frequency or weight variable, if specified, is missing or nonpositive.
Targeted Maximum Likelihood Estimation
When you estimate potential outcome means by targeted maximum likelihood estimation (TMLE), predictions from a model for the outcome variable and predictions from a model for the treatment variable are combined to form a doubly robust estimation method. An estimation method is said to be doubly robust if it provides unbiased estimates whenever one of the models that you use is correctly specified. The model for the outcome variable is used in the TMLE method to predict potential outcome values for each subject by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest. The model for the treatment variable is used to compute covariates that a targeting step uses to reduce finite sample bias and model misspecification bias in the outcome model. For more information about the TMLE method, see Gruber and Van der Laan (2009), Van der Laan (2010a), Van der Laan (2010b), and Van der Laan and Rose (2011).
To estimate potential outcome means by TMLE, you must provide as input
variables that contain the predicted probability of a subject being assigned to the specified treatment levels. You specify the variables that contain the precomputed treatment assignment probabilities by using the TREATPROB= option in each potential outcome specification. Also, you must specify either
variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the TMLE method, see Estimation by Doubly Robust Methods.
The following list defines components of the equations that you use to obtain the TMLE estimates of the potential outcome means:
T denotes the treatment variable.
denotes the predicted potential outcome value that you obtain by replacing the observed treatment assignment with
.
denotes the predicted potential outcome value that corresponds to the observed treatment value.
denotes the predicted potential outcome value that you obtain by applying the targeting step to the initial predicted potential outcome value that corresponds to treatment level
.
denotes the indicator function for subject i being assigned to treatment level t.
denotes the conditional probability of subject i being assigned to treatment level t; that is,
.
The TMLE algorithm starts with the predicted potential outcomes that are obtained from the outcome model that you provide as input. The algorithm then uses a targeting step to refine those predictions. The targeting step is performed by regressing the observed outcome variable against a carefully chosen set of covariates; the initial predicted potential outcomes are specified as a regression offset.
There is one covariate, , for each potential outcome mean that you want to estimate (
). When you estimate the potential outcome means in the population, the values of the clever covariates are computed for each observation by the equation
Note that for each individual, only one of the covariates is nonzero. The covariates are used in the regression equation
In the targeting step, the TMLE algorithm computes estimates of the regression parameters
and then uses these estimates to update the predicted potential outcomes. These updated predictions are computed by the formula
When the outcome variable is binary, the estimated regression parameters are computed by logistic regression. When the outcome variable is continuous or binomial, the regression parameters are computed by a generalized linear model that has a binomial distribution and a logit link. In the latter case, you must rescale the observed outcomes and the predicted potential outcomes so that all values are contained in the open interval (0,1). In both cases, the regression is computed with no intercept and with
as the offset.
After the targeting step, the potential outcome means are given by
If you specify a weight variable by using the WEIGHT statement, this weight variable is applied to the regression model that defines the targeting step, and final computation of the potential outcome means is updated accordingly. When you use the TMLE estimation method, an observation from the input data table is not used if any of the following conditions occur:
A predicted treatment assignment probability is missing.
The value of the outcome variable is missing.
A predicted potential outcome value is missing.
The value of a frequency or weight variable, if specified, is missing or nonpositive.