The BCHOICE Procedure

A Logit Model Example with Random Effects

(View the complete code for this example.)

Choice models that have random effects (or random coefficients) provide solutions to create individual-level or group-specific utilities. Because people have different preferences, it can be misleading to roll the whole sample together into a single set of utilities. The desire to account for individual differences, instead of treating all respondents alike, provides challenges in marketing research. For logit models that have random effects, using frequentist methods to optimize of the likelihood function can be numerically difficult. Bayesian methods are ideally suited for analysis with random effects.

Choice models that have random effects generalize the standard choice models to incorporate individual-level effects. Let the utility that individual i obtains from alternative j in choice situation t ( ) be

where is the observed choice for individual i and alternative j in choice situation t; is the fixed design vector for individual i and alternative j in choice situation t; are the fixed coefficients; is the random design vector for individual i and alternative j in choice situation t; and are the random coefficients for individual i corresponding to .

It is assumed that each is drawn from a superpopulation and that this superpopulation is normal, . An additional stage is added to the model in which a prior for is specified:

The covariance matrix characterizes the extent of heterogeneity among individuals. Large diagonal elements of indicate substantial heterogeneity in part-worths. Off-diagonal elements indicate patterns in the evaluation of attribute levels in pairs.

Consider a study that estimates the market demand for kitchen trash cans (Rossi 2013). There are four attributes, and each has two levels: touchless opening (Yes/No), material (Steel/Plastic), automatic trash bag replacement (Yes/No), and price (40). The number of all possible hypothetical types of trash cans is . Including more attributes and more levels can easily become unmanageable. The study uses a fractional factorial design, in which the first three factors are set up to be a full factorial design and the fourth is generated as the product of the first three. This design confounds the three-way interaction with the effect of the fourth factor, shown in Table 27.1.

Table 27.1: Design for the Trash Can Study

Obs

Touchless

Steel

AutoBag

Price80

1

–1

–1

–1

–1

2

–1

–1

1

1

3

–1

1

–1

1

4

–1

1

1

–1

5

1

–1

–1

1

6

1

–1

1

–1

7

1

1

–1

–1

8

1

1

1

1


In Table 27.1, 1 means "Yes" and –1 means "No." This is a balanced design, in which each level appears the same number of times. This study assigns only two alternatives to a choice set by randomly sampling two rows from the previous table and giving each individual 10 choice sets (or choice tasks) to pick from. For more information about how to design a choice model efficiently, see Kuhfeld (2010).

Data were obtained by enrolling 104 people and assigning 10 choice tasks to each of them: for each task, the participants stated their preference between two types of trash cans. The following steps read in the data:

data Trashcan;
   input ID Task Choice Index Touchless Steel AutoBag Price80 @@;
   datalines;
1 1 1 1 0 1 1 0 1 1 0 2 1 1 0 0 1 2 0 1 0 0 0 0 1 2 1 2 1 1 1 1 1 3 0 1
0 0 0 0 1 3 1 2 1 1 0 0 1 4 0 1 0 1 0 1 1 4 1 2 1 0 0 1 1 5 1 1 0 1 1 0
1 5 0 2 1 0 0 1 1 6 0 1 0 0 1 1 1 6 1 2 1 1 1 1 1 7 0 1 1 0 0 1 1 7 1 2
1 1 0 0 1 8 0 1 0 0 1 1 1 8 1 2 0 1 1 0 1 9 1 1 0 0 1 1 1 9 0 2 1 0 0 1
1 10 0 1 0 1 1 0 1 10 1 2 1 0 1 0 2 1 1 1 0 1 1 0 2 1 0 2 1 1 0 0 2 2 0
1 0 0 0 0 2 2 1 2 1 1 1 1 2 3 0 1 0 0 0 0 2 3 1 2 1 1 0 0 2 4 0 1 0 1 0
1 2 4 1 2 1 0 0 1 2 5 1 1 0 1 1 0 2 5 0 2 1 0 0 1 2 6 1 1 0 0 1 1 2 6 0
2 1 1 1 1 2 7 1 1 1 0 0 1 2 7 0 2 1 1 0 0 2 8 1 1 0 0 1 1 2 8 0 2 0 1 1
0 2 9 1 1 0 0 1 1 2 9 0 2 1 0 0 1 2 10 1 1 0 1 1 0 2 10 0 2 1 0 1 0 3 1

   ... more lines ...   

2 1 2 1 1 1 1 104 3 0 1 0 0 0 0 104 3 1 2 1 1 0 0 104 4 0 1 0 1 0 1 104
4 1 2 1 0 0 1 104 5 1 1 0 1 1 0 104 5 0 2 1 0 0 1 104 6 0 1 0 0 1 1 104
6 1 2 1 1 1 1 104 7 0 1 1 0 0 1 104 7 1 2 1 1 0 0 104 8 0 1 0 0 1 1 104
8 1 2 0 1 1 0 104 9 0 1 0 0 1 1 104 9 1 2 1 0 0 1 104 10 0 1 0 1 1 0 104
10 1 2 1 0 1 0
;
proc print data=Trashcan (obs=8);
run;

The data for the first four choice tasks are shown in Figure 27.8.

Figure 27.8: Data for the First Four Choice Tasks

ObsIDTaskChoiceIndexTouchlessSteelAutoBagPrice80
111110110
211021100
312010000
412121111
513010000
613121100
714010101
814121001


In the data, ID is the individual’s ID number, and Task indexes the number of choice tasks. The response is Choice, which states each individual’s choice for each choice task. Touchless, Steel, AutoBag, and Price80 are the attribute variables; for each of them, 1 means "Yes" and 0 means "No." In the data, 0 replaces the –1 values that are shown in the design matrix in Table 27.1.

The following statements fit a logit model with random effects:

proc bchoice data=Trashcan seed=1 nmc=30000 thin=2 nthreads=4;
   class ID Task;
   model Choice = Touchless Steel AutoBag Price80 / choiceset=(ID Task);
   random Touchless Steel AutoBag Price80 / sub=ID monitor=(1 to 5) type=un;
run;

The NTHREADS option in the PROC BCHOICE statement specifies the number of threads to be used for running analytic computations and simulation simultaneously. Using four threads at the same time enhances the efficiency and reduces the run time. If you do not specify the NTHREADS option, the default number is 1. The maximum number of threads should not exceed the total number of CPUs on the host where the analytic computations execute.

The choice set is specified by ID (which identifies the participants) and by Task (which identifies each of the 10 choice tasks that are assigned to each participant). The variables ID and Task are needed in the CLASS statement because they define the choice set in the MODEL statement.

In addition to the MODEL statement for fixed effects, the RANDOM statement is added for random effects. Note that Touchless, Steel, AutoBag, and Price80 are listed as both fixed and random effects, so that their average part-worth values in the population are estimated via fixed effects and the deviation from the overall mean for each individual is presented through random effects. The SUB=ID argument in the RANDOM statement defines ID as a subject index for the random effects grouping, so that each person with a different ID has his own random effects. The MONITOR=(1 to 5) option requests to display the summary and diagnostics statistics of the random-effects parameter estimates for the first five subjects. By default, PROC BCHOICE does not print the summary and diagnostics statistics for any individual-level random-effects to save time and space. In models that have a large number of individual random effects (for example, tens of thousands individuals), it may take a long time to display the summary, diagnostics statistics, and plots for all the individual-level parameters, so be cautious when using the MONITOR option without providing a subset of the random-effects parameters.

The TYPE=UN option in the RANDOM statement specifies an unstructured covariance matrix for the random effects. The unstructured type provides a mechanism for estimating the correlation between the random effects. The TYPE=VC (variance components) option, which is a diagonal matrix and is the default structure, models a different variance component for each random effect.

Summary statistics for the fixed coefficients (), the covariance of the random coefficients (), and the random coefficients () for the first five individuals are shown in Figure 27.9.

Figure 27.9: Posterior Summary Statistics

The BCHOICE Procedure

Posterior Summaries and Intervals
ParameterSubjectNMeanStandard
Deviation
95% HPD Interval
Touchless 150001.70430.27091.16452.2322
Steel 150001.05160.26080.53761.5691
AutoBag 150002.17220.35831.48862.8948
Price80 15000-4.63210.6720-5.9688-3.4046
RECov Touchless, Touchless 150003.11451.06131.34415.2576
RECov Steel, Touchless 15000-0.47270.8796-2.18841.2334
RECov Steel, Steel 150002.63120.95651.10354.5817
RECov AutoBag, Touchless 15000-0.64730.8811-2.38031.1974
RECov AutoBag, Steel 150000.05950.7078-1.43181.3748
RECov AutoBag, AutoBag 150003.55471.58370.96116.7248
RECov Price80, Touchless 15000-1.30161.1857-3.80460.7154
RECov Price80, Steel 15000-1.59351.2052-4.14280.6170
RECov Price80, AutoBag 15000-2.21901.7156-5.70730.7349
RECov Price80, Price80 150007.74523.43622.228714.3546
TouchlessID 1150000.52441.0525-1.33662.7890
SteelID 115000-0.32981.0799-2.30571.9751
AutoBagID 1150001.62161.3842-0.92074.4860
Price80ID 115000-0.53942.2018-5.22513.4324
TouchlessID 215000-1.74980.9341-3.59650.0610
SteelID 215000-1.43430.8572-3.21760.2048
AutoBagID 2150000.47671.3208-2.02393.1743
Price80ID 2150004.68731.49551.92237.7535
TouchlessID 315000-0.61910.9695-2.40241.4490
SteelID 3150000.57790.9524-1.27922.4187
AutoBagID 315000-1.87371.2107-4.28780.4407
Price80ID 3150003.75071.67470.47957.0672
TouchlessID 415000-0.59210.9679-2.52221.2694
SteelID 4150000.31521.0416-1.71342.3286
AutoBagID 4150000.97701.4250-1.71033.9276
Price80ID 415000-1.90682.3717-6.42562.8233
TouchlessID 515000-1.26891.1491-3.47200.9828
SteelID 5150001.65811.2220-0.47314.2409
AutoBagID 5150001.26951.5708-1.75064.4762
Price80ID 515000-0.37062.3810-5.05934.2134


The fixed effects (Touchless, Steel, AutoBag, and Price80) are shown in the first four rows. Across all the respondents in the data, the average part-worths for touchless opening, steel material, and automatic trash bag replacement are all positive, indicating that most people favor those features; the average part-worth for having to pay USD80 for a trash can instead of USD40 is negative (–4.6), which is very intuitive, because spending more money is usually unfavorable.

The covariance estimate of the random coefficients () is displayed by the parameters whose label begins with "RECov":

where the dots refer to the corresponding elements in the lower part of the symmetric covariance matrix. The covariance estimate of the random coefficients () characterizes the variability of part-worths across respondents. Some of the diagonal elements of the matrix are large. For example, the variance for price (labeled "RECov Price80, Price80") is quite large, indicating substantial unexplained difference in response to price. Off-diagonal elements of the matrix illustrate attribute levels that tend to be evaluated similarly (positive covariance) or differently (negative covariance) across all the respondents. The covariances between each of the attributes (Touchless, Steel, and AutoBag) and Price80 are all negative, implying that the respondents who prefer some of the new features are those who are also unwilling to pay a higher price for the trash can. Therefore, offering a discounted price might be a particularly effective method of introducing the new features to customers.

The next set of parameters that are displayed are the estimates for the individual-level random effects for the first five respondents (see Figure 27.9). These estimates are the deviation from the overall means (which are estimated via the fixed effects). The part-worth for touchless opening for the first respondent (who is labeled "ID 1" in the Subject column) is 1.7 + 0.5 = 2.2.

Allenby and Rossi (1999) and Rossi, Allenby, and McCulloch (2005) propose a hierarchical Bayesian random-effects model that is set up in a different way such that there are no fixed effects but only random effects. For more information about this type of model, see the section Random Effects and a follow-up example in A Random-Effects-Only Logit Model.