CAEFFECT Procedure

Example 6.1 Estimation by Inverse Probability Weighting

(View the complete code for this example.)

This example uses the SmokingWeight data set available in the examples for the CAUSALTRT procedure in SAS/STAT software. These data are a subset of data from the NHANES I Epidemiologic Follow-Up Study (NHEFS) in Hernán and Robins (2020). For the study, medical and behavioral information were collected during an initial physical examination and again at follow-up interviews about 10 years later. As in the PROC CAUSALTRT examples, the question of interest here is what effect does quitting smoking, as recorded in the variable Quit, have on an individual’s change in weight (measured in kilograms), as recorded in the variable Change.

To create the data table SmokingWeight in your CAS session, you can modify the DATA step that is used to create the data set in PROC CAUSALTRT to create the data table in your CAS engine libref. For this example, it is assumed that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.

In addition to the treatment variable Quit and outcome variable Change, the data include the following variables:

  • Activity: level of daily activity, with values 0, 1, and 2

  • Age: age in 1971

  • BaseWeight: weight in kilograms in 1971

  • Education: level of education, with values 0, 1, 2, 3, and 4

  • Exercise: amount of regular recreational exercise, with values 0, 1, and 2

  • PerDay: number of cigarettes smoked per day in 1971

  • Race: 0 for white; 1 otherwise

  • Sex: 0 for male; 1 for female

  • Weight: weight in kilograms at the follow-up interview

  • YearsSmoke: number of years an individual has smoked

For this example, it is assumed that the variables Activity, Age, BaseWeight, Education, Exercise, PerDay, Race, Sex, and YearsSmoke make up a valid adjustment set that you can use to obtain unbiased estimates of the potential outcome means.

The following statements use the LOGSELECT procedure to fit a logistic model for the treatment variable Quit:

proc logselect data=mylib.SmokingWeight;
   class Sex Race Education Exercise Activity;
   model Quit = Sex Race Education Exercise
                Activity Age YearsSmoke PerDay;
   store out=mylib.logTrtModel;
run;

The model for the treatment variable is specified in the MODEL statement. The STORE statement creates an analytic store that contains the fitted model in the logTrtModel data table.

The following statements use the SCORE statement in the ASTORE procedure to obtain the predicted probabilities of a subject being assigned to each treatment level:

proc astore;
   score data=mylib.SmokingWeight out=mylib.logScored
         rstore=mylib.logTrtModel copyvars=(Quit Change);
run;

You use the DATA= option to specify the data table to score, and you use the RSTORE= to specify the stored model to use for scoring. To avoid data duplication for large data tables, the variables in the input data table are not included in the output data table unless you specify them in the COPYVARS= option. To estimate the potential outcome means by inverse probability weighting, you need the treatment variable and observed outcome values, so the variables Quit and Change are specified in the COPYVARS= option.

The following statements use the CAEFFECT procedure to estimate potential outcome means by using inverse probability weighting (IPW):

proc caeffect data=mylib.logScored inference;
   treatvar Quit;
   outcomevar Change;
   pom treatlev = 1 treatProb = P_Quit1;
   pom treatlev = 0 treatProb = P_Quit0;
   difference reflev=0;
run;

When you use the CAEFFECT procedure, you must specify the TREATVAR statement and at least one POM statement. You specify the input data table to use for the potential outcome means estimation in the DATA= option. The TREATVAR statement specifies information about the treatment variable. Each POM statement specifies a potential outcome of interest for your analysis.

For each potential outcome specification, you must use the TREATLEV= option to identify the level of treatment that defines the potential outcome. In this example, two potential outcomes are specified. One outcome is defined by the level 1, which corresponds to a subject quitting smoking, and the second outcome is defined by the level 0, which corresponds to a subject not quitting. For each potential outcome specification, you must also identify any variables in the input data table that are needed for the estimation. For IPW estimation, you must specify variables in the input data table that contain the previously computed treatment assignment probabilities; you do this by using the TREATPROB= option. In this example, the names of those variables are created by PROC ASTORE. If you do not know the names of the variables created by PROC ASTORE that contain the predicted values, you can use the DESCRIBE statement in PROC ASTORE and check the "Output Variables" table that it creates to find the variable names.

For IPW estimation, the observed outcome values are required, and therefore you must also specify the OUTCOMEVAR statement. Finally, in this example, you specify a DIFFERENCE statement to request a difference between the potential outcome mean estimates. A single difference is requested, and the REFLEV= option specifies that the treatment level 0 is used as the reference level. In general, you can specify multiple DIFFERENCE statements, and in each DIFFERENCE statement you can specify reference levels, event levels, or both.

The output from this analysis is displayed by default. It is presented in Output 6.1.1 through Output 6.1.4.

The "Model Information" table in Output 6.1.1 summarizes important information, including the data source, treatment variable name, estimation method, and type of outcome variable.

Output 6.1.1: Model Information

The CAEFFECT Procedure

Model Information
Data SourceLOGSCORED
Treatment VariableQuit
Outcome VariableChange
Outcome TypeContinuous
Estimation MethodInverse Probability Weighting


Output 6.1.2 displays the "Number of Observations" table. In this analysis, a number of observations are not used, because there are missing predicted treatment assignment probabilities. These probabilities are missing because there are missing values in the predictor variables.

Output 6.1.2: Number of Observations

Number of Observations Read1746
Number of Observations Used1566


Output 6.1.3 displays the "Treatment Profile" table. This table includes information about the levels of the treatment variable that appear in the input data table and the number of observations in each level that are used to estimate the potential outcome means.

Output 6.1.3: Treatment Profile

Treatment Profile
QuitTotal
Frequency
01163
1403


Output 6.1.4 displays the potential outcome mean estimates and their difference. The "POM Estimates" table indicates the level of treatment that defines the potential outcome and the estimate of its mean. The "POM Differences" table shows the requested difference in potential outcome mean estimates. The table also includes nonprinting columns that contain the potential outcome mean estimates for the event and reference levels; these columns are accessible if you save the table as an output data set by using the ODS OUTPUT statement or if you modify the table’s template. If the relevant causal assumptions hold true for the data-generating process that produced these data, then the difference in potential outcome mean estimates corresponds to an estimate of the average treatment effect (ATE) of a weight change that is approximately 3.25 kilograms greater because the subjects quit smoking.

Output 6.1.4: Potential Outcome Mean and Difference Estimates

POM Estimates
Treatment
Level
EstimateStandard
Error
95% Confidence LimitsZPr > |Z|
15.051460.472544.125305.9776210.69<.0001
01.800990.219381.371002.230978.21<.0001

POM Differences
Treatment LevelsDifferenceStandard
Error
 ZPr > |Z|
EventReference
95% Confidence Limits
103.25050.520982.229364.271586.24<.0001


By default, PROC CAEFFECT does not produce standard errors or confidence limits, because the validity of the covariance matrix estimate for the potential outcome mean estimates depends on factors outside the procedure invocation. In this example, where maximum likelihood estimation is used to estimate the treatment model and inverse probability weighting is used to estimate the potential outcome means, the variance estimation method that the procedure uses is known to produce valid but conservative statistical inference even without the use of data splitting. For this reason, standard errors and confidence limits are produced by using the INFERENCE option. For more information about the factors that determine the validity of the inferential statistics that the procedure produces, see the section Standard Errors and Confidence Limits.

You can also use the IPW estimation method in PROC CAEFFECT to produce estimates of the potential outcome means that are conditional on subjects receiving a particular level of treatment. To specify the level of treatment to condition on, you can use the CONDEVENT= option in the TREATVAR statement. The following statements request an analysis that uses the same data, but in this case for potential outcome mean estimates conditional on subjects quitting smoking:

proc caeffect data=mylib.logScored;
   treatvar Quit / condevent=1;
   outcomevar Change;
   pom treatlev = 1 treatProb = P_Quit1;
   pom treatlev = 0 treatProb = P_Quit0;
   difference reflev=0;
run;

The potential outcome mean estimates conditional on subjects quitting smoking and their difference are displayed in Output 6.1.5. The "POM Estimates" and "POM Differences" tables now include a column for the level of treatment that defines the conditioning event; this column is not displayed but is accessible if you save the table as an output data set by specifying the ODS OUTPUT statement. In this case, because the potential outcome that defines the event level in the difference comparison is the same as the level that defines the conditioning event, the estimated difference would correspond to an estimate of the average treatment effect for the treated (ATT).

Output 6.1.5: Conditional Potential Outcome Mean and Difference Estimates

The CAEFFECT Procedure

POM Estimates
Treatment
Level
Estimate
14.52491
01.28412

POM Differences
Treatment LevelsDifference
EventReference
103.2408


Last updated: June 22, 2026