The CATMOD Procedure

Example 32.1 Linear Response Function, r=2 Responses

(View the complete code for this example.)

In an example from Ries and Smith (1963), the choice of detergent brand (Brand = M or X) is related to three other categorical variables: the softness of the laundry water (Softness = soft, medium, or hard), the temperature of the water (Temperature = high or low), and whether the subject was a previous user of Brand M (Previous = yes or no). The linear response function, which could also be specified as RESPONSE MARGINALS, yields one probability, Pr(brand preference=M), as the response function to be analyzed. Two models are fit in this example: the first model is a saturated one, containing all of the main effects and interactions, while the second is a reduced model containing only the main effects. The following statements produce Output 32.1.1 through Output 32.1.4:

data detergent;
   input Softness $ Brand $ Previous $ Temperature $ Count @@;
   datalines;
soft X yes high 19   soft X yes low 57
soft X no  high 29   soft X no  low 63
soft M yes high 29   soft M yes low 49
soft M no  high 27   soft M no  low 53
med  X yes high 23   med  X yes low 47
med  X no  high 33   med  X no  low 66
med  M yes high 47   med  M yes low 55
med  M no  high 23   med  M no  low 50
hard X yes high 24   hard X yes low 37
hard X no  high 42   hard X no  low 68
hard M yes high 43   hard M yes low 52
hard M no  high 30   hard M no  low 42
;
title 'Detergent Preference Study';
proc catmod data=detergent;
   response 1 0;
   weight Count;
   model Brand=Softness|Previous|Temperature / freq prob;
   title2 'Saturated Model';
run;

The "Data Summary" table (Output 32.1.1) indicates that you have two response levels and twelve populations.

Output 32.1.1: Detergent Preference Study: Linear Model Analysis

Detergent Preference Study
Saturated Model

The CATMOD Procedure

Data Summary
ResponseBrandResponse Levels2
Weight VariableCountPopulations12
Data SetDETERGENTTotal Frequency1008
Frequency Missing0Observations24


The "Population Profiles" table in Output 32.1.2 displays the ordering of independent variable levels as used in the table of parameter estimates.

Output 32.1.2: Population Profiles

Population Profiles
SampleSoftnessPreviousTemperatureSample Size
1hardnohigh72
2hardnolow110
3hardyeshigh67
4hardyeslow89
5mednohigh56
6mednolow116
7medyeshigh70
8medyeslow102
9softnohigh56
10softnolow116
11softyeshigh48
12softyeslow106


Since Brand M is the first level in the "Response Profiles" table (Output 32.1.3), the RESPONSE statement causes Pr(Brand=M) to be the single response function modeled.

Output 32.1.3: Response Profiles, Frequencies, and Probabilities

Response Profiles
ResponseBrand
1M
2X

Response Frequencies
SampleResponse Number
12
13042
24268
34324
45237
52333
65066
74723
85547
92729
105363
112919
124957

Response Probabilities
SampleResponse Number
12
10.416670.58333
20.381820.61818
30.641790.35821
40.584270.41573
50.410710.58929
60.431030.56897
70.671430.32857
80.539220.46078
90.482140.51786
100.456900.54310
110.604170.39583
120.462260.53774


The "Analysis of Variance" table in Output 32.1.4 shows that all of the interactions are nonsignificant.

Output 32.1.4: Analysis of Variance

Analysis of Variance
SourceDF Chi-SquarePr > ChiSq
Intercept1983.13<.0001
Softness20.090.9575
Previous122.68<.0001
Softness*Previous23.850.1457
Temperature13.670.0555
Softness*Temperature20.230.8914
Previous*Temperature12.260.1324
Softnes*Previou*Temperat20.760.6850
Residual0..


Therefore, a main-effects model is fit with the following statements:

   model Brand=Softness Previous Temperature
       / clparm noprofile design;
   title2 'Main-Effects Model';
run;
quit;

The PROC CATMOD statement is not required due to the interactive capability of the CATMOD procedure. The NOPROFILE option suppresses the redisplay of the "Response Profiles" table. The CLPARM option produces 95% confidence limits for the parameter estimates. Output 32.1.5 through Output 32.1.7 are produced.

The design matrix in Output 32.1.5 displays the results of the differential-effects modeling used in PROC CATMOD.

Output 32.1.5: Main-Effects Design Matrix

Detergent Preference Study
Main-Effects Model

The CATMOD Procedure

Data Summary
ResponseBrandResponse Levels2
Weight VariableCountPopulations12
Data SetDETERGENTTotal Frequency1008
Frequency Missing0Observations24

Response Functions and Design Matrix
SampleResponse
Function
Design Matrix
1 2 3 4 5
10.4166711011
20.381821101-1
30.64179110-11
40.58427110-1-1
50.4107110111
60.431031011-1
70.67143101-11
80.53922101-1-1
90.482141-1-111
100.456901-1-11-1
110.604171-1-1-11
120.462261-1-1-1-1


The analysis of variance table in Output 32.1.6 shows that previous use of Brand M, together with the temperature of the laundry water, is a significant factor in whether a subject prefers Brand M laundry detergent. The table also shows that the additive model fits since the goodness-of-fit statistic (the residual chi-square) is nonsignificant.

Output 32.1.6: ANOVA Table for the Main-Effects Model

Analysis of Variance
SourceDF Chi-SquarePr > ChiSq
Intercept11004.93<.0001
Softness20.240.8859
Previous120.96<.0001
Temperature13.950.0468
Residual78.260.3100


The chi-square test in Output 32.1.7 shows that the Softness parameters are not significantly different from zero; as expected, the Wald confidence limits for these two estimates contain zero. So softness of the water is not a factor in choosing Brand M.

Output 32.1.7: WLS Estimates for the Main-Effects Model

Analysis of Weighted Least Squares Estimates
Parameter Estimate Standard
Error
Chi-
Square
Pr > ChiSq95% Confidence Limits
Intercept 0.50800.01601004.93<.00010.47660.5394
Softnesshard-0.002560.02180.010.9066-0.04540.0402
 med0.01040.02180.230.6342-0.03230.0530
Previousno-0.07110.015520.96<.0001-0.1015-0.0407
Temperaturehigh0.03190.01613.950.04680.0004460.0634


The negative coefficient for Previous (–0.0711) indicates that the first level of Previous (which is shown to be 'no') is associated with a smaller probability of preferring Brand M than the second level of Previous (with coefficient constrained to be 0.0711 since the parameter estimates for a given effect must sum to zero). In other words, previous users of Brand M are much more likely to prefer it than those who have never used it before.

Similarly, the positive coefficient for Temperature indicates that the first level of Temperature (which, from the "Population Profiles" table, is 'high') has a larger probability of preferring Brand M than the second level of Temperature. In other words, those who do their laundry in hot water are more likely to prefer Brand M than those who do their laundry in cold water.