LOGSELECT Procedure
Example 18.4 Ordinal Logistic Regression
(View the complete code for this example.)
Consider a study of the effects of various cheese additives on taste. Researchers tested four cheese additives and obtained 52 response ratings for each additive. Each response was measured on a scale of nine categories ranging from strong dislike (1) to excellent taste (9). The data, given in McCullagh and Nelder (1989, p. 175) in the form of a two-way frequency table of additive by rating, are saved in the data table mylib.Cheese by using the following program. The variable y contains the response rating. The variable Additive specifies the cheese additive (1, 2, 3, or 4). The variable freq gives the frequency with which each additive received each rating.
data mylib.Cheese;
do Additive = 1 to 4;
do y = 1 to 9;
input freq @@;
output;
end;
end;
label y='Taste Rating';
datalines;
0 0 1 7 8 8 19 8 1
6 9 12 11 7 6 1 0 0
1 1 6 8 23 7 5 1 0
0 0 0 1 3 7 14 16 11
;
The response variable y is ordinally scaled. A cumulative logit model is used to investigate the effects of the cheese additives on taste. The following statements invoke PROC LOGSELECT to fit this model with y as the response variable and three indicator variables as explanatory variables, with the fourth additive as the reference level. With this parameterization, each Additive parameter compares an additive to the fourth additive.
proc logselect data=mylib.Cheese;
freq freq;
class Additive(ref='4') / param=ref;
model y=Additive;
run;
Results from the logistic analysis are shown in Output 18.4.1 through Output 18.4.3.
The "Response Profile" table in Output 18.4.1 shows that the strong dislike (y=1) end of the rating scale is associated with lower OrderedValue values in the "Response Profile" table; hence the probability of disliking the additives is modeled.
Output 18.4.1: Proportional Odds Model Regression Analysis
| Model Information | |
|---|---|
| Data Source | CHEESE |
| Response Variable | y |
| Number of Response Levels | 9 |
| Frequency Variable | freq |
| Distribution | Multinomial |
| Link Type | Cumulative |
| Link Function | Logit |
| Optimization Technique | Newton-Raphson with Ridging |
| Number of Observations Read | 36 |
|---|---|
| Number of Observations Used | 28 |
| Sum of Frequencies Read | 208 |
| Sum of Frequencies Used | 208 |
| Response Profile | ||
|---|---|---|
| Ordered Value | y | Total Frequency |
| 1 | 1 | 7 |
| 2 | 2 | 10 |
| 3 | 3 | 19 |
| 4 | 4 | 27 |
| 5 | 5 | 41 |
| 6 | 6 | 28 |
| 7 | 7 | 39 |
| 8 | 8 | 25 |
| 9 | 9 | 12 |
| Probabilities modeled are cumulated over the lower Ordered Values. |
| Class Level Information | ||
|---|---|---|
| Class | Levels | Values |
| Additive | 4 | 1 2 3 4 |
Output 18.4.2: Proportional Odds Model Regression Analysis
| Convergence criterion (GCONV=1E-8) satisfied. |
| Dimensions | |
|---|---|
| Columns in Design | 11 |
| Number of Effects | 2 |
| Max Effect Columns | 8 |
| Rank of Design | 11 |
| Parameters in Optimization | 11 |
| Testing Global Null Hypothesis: BETA=0 | |||
|---|---|---|---|
| Test | DF | Chi-Square | Pr > ChiSq |
| Likelihood Ratio | 3 | 148.4539 | <.0001 |
| Fit Statistics | |
|---|---|
| -2 Log Likelihood | 711.34790 |
| AIC (smaller is better) | 733.34790 |
| AICC (smaller is better) | 734.69484 |
| SBC (smaller is better) | 770.06082 |
The positive value (1.6128) for the parameter estimate for Additive=1 in Output 18.4.3 indicates a tendency toward the lower-numbered categories of the first cheese additive relative to the fourth. In other words, the fourth additive tastes better than the first additive. Similarly, the second and third additives are both less favorable than the fourth additive. The relative magnitudes of these slope estimates imply the preference ordering: fourth, first, third, second.
Output 18.4.3: Proportional Odds Model Regression Analysis
| Parameter Estimates | ||||||
|---|---|---|---|---|---|---|
| Parameter | y | DF | Estimate | Standard Error | Chi-Square | Pr > ChiSq |
| Intercept | 1 | 1 | -7.080166 | 0.564010 | 157.5844 | <.0001 |
| Intercept | 2 | 1 | -6.024980 | 0.476431 | 159.9230 | <.0001 |
| Intercept | 3 | 1 | -4.925416 | 0.425651 | 133.8992 | <.0001 |
| Intercept | 4 | 1 | -3.856801 | 0.388022 | 98.7968 | <.0001 |
| Intercept | 5 | 1 | -2.520552 | 0.345268 | 53.2940 | <.0001 |
| Intercept | 6 | 1 | -1.568538 | 0.312208 | 25.2408 | <.0001 |
| Intercept | 7 | 1 | -0.066875 | 0.273819 | 0.0596 | 0.8071 |
| Intercept | 8 | 1 | 1.492974 | 0.335696 | 19.7794 | <.0001 |
| Additive 1 | 1 | 1.612791 | 0.380544 | 17.9617 | <.0001 | |
| Additive 2 | 1 | 4.964640 | 0.476721 | 108.4546 | <.0001 | |
| Additive 3 | 1 | 3.322683 | 0.421830 | 62.0444 | <.0001 | |
Multinomial-Trial Syntax for Grouped Data
These data can be grouped according to the value of the Additive variable in the following DATA step:
data mylib.Cheese;
do Additive = 1 to 4;
input y1-y9;
output;
end;
datalines;
0 0 1 7 8 8 19 8 1
6 9 12 11 7 6 1 0 0
1 1 6 8 23 7 5 1 0
0 0 0 1 3 7 14 16 11
;
The analysis proceeds as before with multinomial-trial syntax:
proc logselect data=mylib.Cheese;
class Additive(ref='4') / param=ref;
model y1-y9=Additive;
run;
Parameter estimates and tests are the same as for the preceding analysis and are not displayed here. The tables displayed in Output 18.4.4 and Output 18.4.5 are different from those produced by the preceding analysis: the "Model Information" and "Response Profile" tables include the nine response variables, there are only four observations in the data, and the fit statistics are modified by the inclusion of the constant term in the log likelihood.
Output 18.4.4: Information Tables with Multinomial-Trial Syntax
| Model Information | |
|---|---|
| Data Source | CHEESE |
| Response Variable | y1 |
| Response Variable | y2 |
| Response Variable | y3 |
| Response Variable | y4 |
| Response Variable | y5 |
| Response Variable | y6 |
| Response Variable | y7 |
| Response Variable | y8 |
| Response Variable | y9 |
| Number of Response Levels | 9 |
| Distribution | Multinomial |
| Link Type | Cumulative |
| Link Function | Logit |
| Optimization Technique | Newton-Raphson with Ridging |
| Number of Observations Read | 4 |
|---|---|
| Number of Observations Used | 4 |
| Response Profile | ||
|---|---|---|
| Ordered Value | Response | Total Frequency |
| 1 | y1 | 7 |
| 2 | y2 | 10 |
| 3 | y3 | 19 |
| 4 | y4 | 27 |
| 5 | y5 | 41 |
| 6 | y6 | 28 |
| 7 | y7 | 39 |
| 8 | y8 | 25 |
| 9 | y9 | 12 |
| Probabilities modeled are cumulated over the lower Ordered Values. |
Output 18.4.5: Fit Statistics with Multinomial-Trial Syntax
| Fit Statistics | |
|---|---|
| -2 Log Likelihood | 95.33994 |
| AIC (smaller is better) | 117.33994 |
| AICC (smaller is better) | 118.68687 |
| SBC (smaller is better) | 154.05285 |