CAEFFECT Procedure

Estimation Methods

PROC CAEFFECT implements several methods of estimating potential outcome means and their differences. The causal interpretation of these estimates depends on the plausibility of the causal assumptions that are described in the section Potential Outcomes and Effect Definitions.

These estimation methods differ primarily in how they make statistical adjustments for confounding variables. In general, you can adjust for confounding variables by modeling the treatment variable, the outcome variable, or both. PROC CAEFFECT is a model-agnostic tool for effect estimation. That is, the procedure itself does not support modeling of the treatment and outcome variables. Instead, as input to the procedure, you provide the quantities necessary for your chosen estimation method. Thus, you can adjust for the effects of confounding variables by using predictions from a previously fit model for the treatment variable, a previously fit model for the outcome variable, or models for both the treatment and outcome variables. The required inputs for each method that PROC CAEFFECT supports are summarized in Table 2.

Table 2: Required Inputs of Different Estimation Methods

Method Treatment Model Outcome Model Observed Outcome
Inverse probability weighting X X
Regression adjustment X
Augmented inverse probability weighting X X X
Targeted maximum likelihood estimation X X X


The inverse probability weighting (IPW) method requires a model for the treatment variable. To use this method, you must specify variables in the input data table that contain the precomputed treatment assignment probabilities. For an example of how you can perform IPW estimation, see Estimation by Inverse Probability Weighting. For more information about estimation by inverse probability weighting, see the section Inverse Probability Weighting.

The regression adjustment method requires a model for the outcome variable. To use this method, you have two choices for the input that you must provide: you can specify variables in the input data table that contain the predicted potential outcomes for each specified level of treatment, or you can provide a previously fit model in the form of an analytic store by using the OUTCOMEMODEL statement. In the second case, the procedure uses the analytic store to compute the relevant predicted outcome values. When you provide an analytic store as input, the model must include the specified treatment variables as a predictor, and all the required predictor variables must appear in the input data table. For an example that shows how you can perform estimation by regression adjustment by providing the procedure with either an analytic store or precomputed predicted values, see Estimation by Regression Adjustment. For more information about estimation by regression adjustment, see the section Regression Adjustment.

The remaining two methods, augmented inverse probability weighting (AIPW) and targeted maximum likelihood estimation (TMLE), are so-called doubly robust estimation methods. An estimation method is considered doubly robust if it provides unbiased estimates whenever at least one of the models is correctly specified. For these methods, you must specify a model for both the treatment variable and the outcome variable. For an example of how you can use PROC CAEFFECT to perform doubly robust estimation, see Estimation by Doubly Robust Methods. For more information about estimation by augmented inverse probability weighting, see the section Augmented Inverse Probability Weighting. For more information about targeted maximum likelihood estimation, see the section Targeted Maximum Likelihood Estimation.

Supported Variable Types

PROC CAEFFECT supports the estimation of treatment effects for categorical treatment variables and for continuous, categorical, or binomial outcomes. In the case of categorical outcomes, you must specify a designated event level by using the EVENT= option in the OUTCOMEVAR statement, and the outcome model (specified as either predicted potential outcomes or an analytic store) must represent the probability of that event level. Similarly, for binomial data, the outcome model must represent the event probability. If you are modeling binomial data and the estimation method that you use requires observed outcome values, you must use the event/trial syntax in the OUTCOMEVAR statement. This syntax requires that you specify two variables that contain the counts from a binomial experiment. The event variable contains the number of positive responses, or events, and the trial variable contains the number of trials that are conducted.

Inverse Probability Weighting

When you estimate potential outcome means by inverse probability weighting (IPW), the observed outcomes for each subject are assigned weights inversely proportional to the probability of having received the observed level of treatment. Incorporating the IPW weights creates a pseudopopulation that corrects for the bias that the confounding variables introduce. You can obtain unbiased estimates of the potential outcome means from the weighted pseudopopulation. For more information about IPW estimation, see Lunceford and Davidian (2004), Hernán and Robins (2020), and references therein.

When estimating script l potential outcome means by IPW, you must provide as input script l variables that contain the predicted treatment probabilities of a subject being assigned to the specified treatment levels. You specify the variables that contain these precomputed probabilities by using the TREATPROB= option in each potential outcome specification. For an example of using PROC CAEFFECT to implement the IPW estimation method, see Estimation by Inverse Probability Weighting.

The following list defines components of the estimating equations that you solve to obtain the IPW estimates of the potential outcome means:

  • T denotes the treatment variable.

  • y Subscript i denotes the observed outcome for subject i.

  • double struck 1 left parenthesis upper T Subscript i Baseline equals t right parenthesis denotes the indicator function for subject i being assigned to treatment level t.

  • e Subscript t i denotes the conditional probability of subject i being assigned to treatment level t; that is, probability left parenthesis upper T Subscript i Baseline equals t vertical bar bold italic upper X equals bold italic x Subscript i Baseline right parenthesis.

The IPW estimates of the potential outcome means solve the equations

bold upper S Subscript normal i normal p normal w Baseline left parenthesis bold italic mu right parenthesis equals sigma summation Underscript i equals 1 Overscript n Endscripts bold upper S Subscript normal i normal p normal w comma i Baseline left parenthesis bold italic mu right parenthesis equals 0

where

bold upper S Subscript normal i normal p normal w comma i Baseline left parenthesis bold italic mu right parenthesis equals Start 3 By 1 Matrix 1st Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t 1 right parenthesis left parenthesis StartFraction y Subscript i Baseline minus mu 1 Over e Subscript 1 i Baseline EndFraction right parenthesis 2nd Row  vertical ellipsis 3rd Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript script l Baseline right parenthesis left parenthesis StartFraction y Subscript i Baseline minus mu Subscript script l Baseline Over e Subscript script l i Baseline EndFraction right parenthesis EndMatrix

The solution to these equations has a closed-form expression given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript Baseline equals left parenthesis sigma summation Underscript i equals 1 Overscript n Endscripts StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction right parenthesis left parenthesis sigma summation Underscript i equals 1 Overscript n Endscripts StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis y Subscript i Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction right parenthesis normal f normal o normal r j equals 1 comma ellipsis comma script l

You can also use the IPW estimation method to estimate potential outcome means conditional on subjects receiving a particular treatment assignment. You specify the level of treatment to condition on by using the CONDEVENT= option in the TREATVAR statement. The level of treatment that you specify for the conditioning event must also be used in a potential outcome specification in a POM statement.

Let t Subscript normal t normal r normal t denote the level of treatment that you specify to condition on for computing conditional potential outcome mean estimates. The IPW estimates of the potential outcome means conditional on upper T equals t Subscript normal t normal r normal t solve a similar set of estimating equations, where the weights that are used to estimate the unconditional means (one over the observed treatment assignment probability) are multiplied by e Subscript t Sub Subscript normal t normal r normal t, the probability of being assigned to the treatment level t Subscript normal t normal r normal t. The score equation for each subject is then given by

bold upper S Subscript normal i normal p normal w comma i Baseline left parenthesis bold italic mu right parenthesis equals Start 3 By 1 Matrix 1st Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t 1 right parenthesis e Subscript t Sub Subscript normal t normal r normal t i Baseline left parenthesis StartFraction y Subscript i Baseline minus mu 1 Over e Subscript 1 i Baseline EndFraction right parenthesis 2nd Row  vertical ellipsis 3rd Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript script l Baseline right parenthesis e Subscript t Sub Subscript normal t normal r normal t i Baseline left parenthesis StartFraction y Subscript i Baseline minus mu Subscript script l Baseline Over e Subscript script l i Baseline EndFraction right parenthesis EndMatrix

The solution to the estimating equations for the conditional potential outcome means is then given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript vertical bar upper T equals t Sub Subscript normal t normal r normal t Subscript Baseline equals left parenthesis sigma summation Underscript i equals 1 Overscript n Endscripts StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis e Subscript t Sub Subscript normal t normal r normal t Subscript i Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction right parenthesis left parenthesis sigma summation Underscript i equals 1 Overscript n Endscripts StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis e Subscript t Sub Subscript normal t normal r normal t Subscript i Baseline y Subscript i Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction right parenthesis normal f normal o normal r j equals 1 comma ellipsis comma script l

If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the IPW estimation method, an observation from the input data table is not used if any of the following conditions occur:

  • A predicted treatment assignment probability is missing.

  • The value of the outcome variable is missing.

  • The observed treatment level is not used in a potential outcome specification.

  • The value of a frequency or weight variable, if specified, is missing or nonpositive.

Regression Adjustment

When you estimate potential outcome means by regression adjustment, you use predicted potential outcome values that are obtained from a model of the outcome variable. The predicted potential outcome values for each subject are obtained by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest and then by evaluating the outcome model for each treatment assignment in turn. The averages of the predicted values are used to estimate the average of the potential outcomes. For more information about estimation by regression adjustment, see Hernán and Robins (2020) and references therein.

To estimate script l potential outcome means by regression adjustment, you must provide as input either script l variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the regression adjustment estimation method, see Estimation by Regression Adjustment.

To describe the estimating equations, you solve to obtain the regression adjustment estimates of the potential outcome means. Let ModifyingAbove y With caret Subscript t i denote the predicted counterfactual outcome value that you obtain by replacing the observed treatment assignment with upper T equals t. The regression adjustment estimates of the potential outcome means solve the equations

bold upper S Subscript normal r normal e normal g normal a normal d normal j Baseline left parenthesis bold italic mu right parenthesis equals sigma summation Underscript i equals 1 Overscript n Endscripts bold upper S Subscript normal r normal e normal g normal a normal d normal j comma i Baseline left parenthesis bold italic mu right parenthesis equals 0

where

bold upper S Subscript normal r normal e normal g normal a normal d normal j comma i Baseline left parenthesis bold italic mu right parenthesis equals Start 3 By 1 Matrix 1st Row  ModifyingAbove y With caret Subscript 1 i Baseline minus mu 1 2nd Row  vertical ellipsis 3rd Row  ModifyingAbove y With caret Subscript script l i Baseline minus mu Subscript script l EndMatrix

The solution to these equations has a closed-form expression given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript Baseline equals StartFraction 1 Over n EndFraction sigma summation Underscript i equals 1 Overscript n Endscripts ModifyingAbove y With caret Subscript t Sub Subscript j Subscript i Baseline normal f normal o normal r j equals 1 comma ellipsis comma script l

You can also use the regression adjustment estimation method to estimate potential outcome means conditional on subjects receiving a particular treatment assignment. You specify the level of treatment to condition on by using the CONDEVENT= option in the TREATVAR statement. The level of treatment that you specify for the conditioning event must also be used in a potential outcome specification in a POM statement.

The regression adjustment estimates of the potential outcome means conditional on upper T equals t Subscript normal t normal r normal t solve a similar set of estimating equations, where predicted counterfactual values are averaged only for subjects that have the treatment assignment that you are conditioning on. The following list defines components of the estimating equations that you solve to obtain the regression adjustment estimates of the potential outcome means conditional on a treatment level:

  • t Subscript normal t normal r normal t denotes the level of treatment that you specify to define the conditioning event upper T equals t Subscript normal t normal r normal t.

  • double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript normal t normal r normal t Baseline right parenthesis denotes the indicator function for subject i being assigned to treatment level t Subscript normal t normal r normal t.

  • n Subscript t Sub Subscript normal t normal r normal t denotes the number of subjects that have the observed treatment assignment upper T Subscript i Baseline equals Subscript normal t normal r normal t Baseline.

The score equation for each subject is then given by

bold upper S Subscript normal r normal e normal g normal a normal d normal j comma i Baseline left parenthesis bold italic mu right parenthesis equals Start 3 By 1 Matrix 1st Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript normal t normal r normal t Baseline right parenthesis left parenthesis ModifyingAbove y With caret Subscript 1 i Baseline minus mu 1 right parenthesis 2nd Row  vertical ellipsis 3rd Row  double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript normal t normal r normal t Baseline right parenthesis left parenthesis ModifyingAbove y With caret Subscript script l i Baseline minus mu Subscript script l Baseline right parenthesis EndMatrix

The solution to these equations has a closed-form expression given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript vertical bar upper T equals t Sub Subscript normal t normal r normal t Subscript Baseline equals StartFraction 1 Over n Subscript t Sub Subscript normal t normal r normal t Subscript Baseline EndFraction sigma summation Underscript i equals 1 Overscript n Endscripts double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript normal t normal r normal t Baseline right parenthesis ModifyingAbove y With caret Subscript t Sub Subscript j Subscript i Baseline normal f normal o normal r j equals 1 comma ellipsis comma script l

If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the regression adjustment estimation method, an observation from the input data table is not used if either of the following conditions occurs:

  • A predicted potential outcome value is missing.

  • The value of a frequency or weight variable, if specified, is missing or nonpositive.

In addition to these conditions, if you are estimating potential outcome means conditional on a particular treatment assignment, only observations that have the specified treatment assignment are used.

Augmented Inverse Probability Weighting

When you estimate potential outcome means by augmented inverse probability weighting (AIPW), predictions from a model for the outcome variable and predictions from a model for the treatment variable are combined to form a doubly robust estimation method. An estimation method is said to be doubly robust if it provides unbiased estimates whenever at least one of the models that you use is correctly specified. The model for the outcome variable is used in the AIPW estimation method to predict potential outcome values for each subject by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest. The model for the treatment variable is used to compute weights that are inversely proportional to the probability that a subject has the observed treatment assignment. For more information about the AIPW estimation method, see Lunceford and Davidian (2004), Hernán and Robins (2020), and references therein.

To estimate script l potential outcome means by AIPW, you must provide as input script l variables that contain the predicted probability of a subject being assigned to the specified treatment levels. You specify the variables that contain the precomputed treatment assignment probabilities by using the TREATPROB= option in each potential outcome specification. Also, you must specify either script l variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the AIPW estimation method, see Estimation by Doubly Robust Methods.

The following list defines components of the estimating equations that you solve to obtain the AIPW estimates of the potential outcome means:

  • T denotes the treatment variable.

  • y Subscript i denotes the observed outcome for subject i.

  • ModifyingAbove y With caret Subscript t i denotes the predicted potential outcome value that you obtain by replacing the observed treatment assignment with upper T equals t.

  • double struck 1 left parenthesis upper T Subscript i Baseline equals t right parenthesis denotes the indicator function for subject i being assigned to treatment level t.

  • e Subscript t i denotes the conditional probability of subject i being assigned to treatment level t; that is, probability left parenthesis upper T Subscript i Baseline equals t vertical bar bold italic upper X equals bold italic x Subscript i Baseline right parenthesis.

The AIPW estimates of the potential outcome means solve the equations

bold upper S Subscript normal a normal i normal p normal w Baseline left parenthesis bold italic mu right parenthesis equals sigma summation Underscript i equals 1 Overscript n Endscripts bold upper S Subscript normal a normal i normal p normal w comma i Baseline left parenthesis bold italic mu right parenthesis equals 0

where

bold upper S Subscript normal a normal i normal p normal w comma i Baseline left parenthesis bold italic mu right parenthesis equals Start 3 By 1 Matrix 1st Row  StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t 1 right parenthesis y Subscript i Baseline Over e Subscript 1 i Baseline EndFraction minus ModifyingAbove y With caret Subscript 1 i Baseline left parenthesis StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t 1 right parenthesis minus e Subscript 1 i Baseline Over e Subscript 1 i Baseline EndFraction right parenthesis minus mu 1 2nd Row  vertical ellipsis 3rd Row  StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript script l Baseline right parenthesis y Subscript i Baseline Over e Subscript script l i Baseline EndFraction minus ModifyingAbove y With caret Subscript script l i Baseline left parenthesis StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript script l Baseline right parenthesis minus e Subscript script l i Baseline Over e Subscript script l i Baseline EndFraction right parenthesis minus mu Subscript script l EndMatrix

The solution to these equations has a closed-form expression given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript Baseline equals StartFraction 1 Over n EndFraction sigma summation Underscript i equals 1 Overscript n Endscripts StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis y Subscript i Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction minus ModifyingAbove y With caret Subscript t Sub Subscript j Subscript i Baseline left parenthesis StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis minus e Subscript t Sub Subscript j Subscript i Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction right parenthesis normal f normal o normal r j equals 1 comma ellipsis comma script l

If you specify a weight variable by using the WEIGHT statement, the score equation for each subject is multiplied by the value of this variable, and the solutions to the estimating equations are updated accordingly. When you use the AIPW estimation method, an observation from the input data table is not used if any of the following conditions occur:

  • A predicted treatment assignment probability is missing.

  • The value of the outcome variable is missing.

  • A predicted counterfactual value is missing.

  • The value of a frequency or weight variable, if specified, is missing or nonpositive.

Targeted Maximum Likelihood Estimation

When you estimate potential outcome means by targeted maximum likelihood estimation (TMLE), predictions from a model for the outcome variable and predictions from a model for the treatment variable are combined to form a doubly robust estimation method. An estimation method is said to be doubly robust if it provides unbiased estimates whenever one of the models that you use is correctly specified. The model for the outcome variable is used in the TMLE method to predict potential outcome values for each subject by replacing the observed treatment assignment with the treatment assignments that define the potential outcomes of interest. The model for the treatment variable is used to compute covariates that a targeting step uses to reduce finite sample bias and model misspecification bias in the outcome model. For more information about the TMLE method, see Gruber and Van der Laan (2009), Van der Laan (2010a), Van der Laan (2010b), and Van der Laan and Rose (2011).

To estimate script l potential outcome means by TMLE, you must provide as input script l variables that contain the predicted probability of a subject being assigned to the specified treatment levels. You specify the variables that contain the precomputed treatment assignment probabilities by using the TREATPROB= option in each potential outcome specification. Also, you must specify either script l variables that contain the predicted potential outcome values or a previously fit model for the outcome in the form of an analytic store. If the input that you provide consists of predicted potential outcome values, you must use the PREDOUT= option for each potential outcome specification to specify the variable that represents the predicted potential outcomes. If the input that you provide consists of a previously fit model, you use the OUTCOMEMODEL statement, specify the name of the analytic store by using the RESTORE= option, and specify the name of the variable that contains the model predictions by using the PREDNAME= option. When you specify a previously fit model, PROC CAEFFECT uses the model to compute the potential outcome predictions. For an example of using this procedure to implement the TMLE method, see Estimation by Doubly Robust Methods.

The following list defines components of the equations that you use to obtain the TMLE estimates of the potential outcome means:

  • T denotes the treatment variable.

  • y Subscript i denotes the observed outcome for subject i.

  • ModifyingAbove y With caret Subscript t i denotes the predicted potential outcome value that you obtain by replacing the observed treatment assignment with upper T equals t.

  • y overtilde Subscript i denotes the predicted potential outcome value that corresponds to the observed treatment value.

  • y Subscript t i Superscript asterisk denotes the predicted potential outcome value that you obtain by applying the targeting step to the initial predicted potential outcome value that corresponds to treatment level upper T equals t.

  • double struck 1 left parenthesis upper T Subscript i Baseline equals t right parenthesis denotes the indicator function for subject i being assigned to treatment level t.

  • e Subscript t i denotes the conditional probability of subject i being assigned to treatment level t; that is, probability left parenthesis upper T Subscript i Baseline equals t vertical bar bold italic upper X equals bold italic x Subscript i Baseline right parenthesis.

The TMLE algorithm starts with the predicted potential outcomes that are obtained from the outcome model that you provide as input. The algorithm then uses a targeting step to refine those predictions. The targeting step is performed by regressing the observed outcome variable against a carefully chosen set of covariates; the initial predicted potential outcomes are specified as a regression offset.

There is one covariate, upper H Subscript t Sub Subscript j, for each potential outcome mean that you want to estimate (j equals 1 comma ellipsis comma script l). When you estimate the potential outcome means in the population, the values of the clever covariates are computed for each observation by the equation

ModifyingAbove h With caret Subscript t Sub Subscript j Subscript i Baseline equals StartFraction double struck 1 left parenthesis upper T Subscript i Baseline equals t Subscript j Baseline right parenthesis Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction normal f normal o normal r j equals 1 comma ellipsis comma script l

Note that for each individual, only one of the script l covariates is nonzero. The covariates are used in the regression equation

normal l normal o normal g normal i normal t left parenthesis y right parenthesis equals normal l normal o normal g normal i normal t left parenthesis y overtilde right parenthesis plus sigma summation Underscript j equals 1 Overscript script l Endscripts epsilon Subscript j Baseline ModifyingAbove h With caret Subscript t Sub Subscript j

In the targeting step, the TMLE algorithm computes estimates ModifyingAbove epsilon With caret Subscript j of the regression parameters epsilon Subscript j and then uses these estimates to update the predicted potential outcomes. These updated predictions are computed by the formula

normal l normal o normal g normal i normal t left parenthesis y Subscript t Sub Subscript j Subscript i Superscript asterisk Baseline right parenthesis equals normal l normal o normal g normal i normal t left parenthesis ModifyingAbove y With caret Subscript t Sub Subscript j Subscript i Baseline right parenthesis plus StartFraction ModifyingAbove epsilon With caret Subscript j Baseline Over e Subscript t Sub Subscript j Subscript i Baseline EndFraction

When the outcome variable is binary, the estimated regression parameters ModifyingAbove epsilon With caret Subscript j are computed by logistic regression. When the outcome variable is continuous or binomial, the regression parameters are computed by a generalized linear model that has a binomial distribution and a logit link. In the latter case, you must rescale the observed outcomes and the predicted potential outcomes so that all values are contained in the open interval (0,1). In both cases, the regression is computed with no intercept and with normal l normal o normal g normal i normal t left parenthesis y overtilde right parenthesis as the offset.

After the targeting step, the potential outcome means are given by

ModifyingAbove mu With caret Subscript t Sub Subscript j Subscript Baseline equals StartFraction 1 Over n EndFraction sigma summation Underscript i equals 1 Overscript n Endscripts y Subscript t Sub Subscript j Subscript i Superscript asterisk Baseline normal f normal o normal r j equals 1 comma ellipsis comma script l

If you specify a weight variable by using the WEIGHT statement, this weight variable is applied to the regression model that defines the targeting step, and final computation of the potential outcome means is updated accordingly. When you use the TMLE estimation method, an observation from the input data table is not used if any of the following conditions occur:

  • A predicted treatment assignment probability is missing.

  • The value of the outcome variable is missing.

  • A predicted potential outcome value is missing.

  • The value of a frequency or weight variable, if specified, is missing or nonpositive.

Last updated: June 22, 2026