The KCLUS Procedure
Example 9.3 Clustering Nominal Variables
In this example, PROC KCLUS clusters nominal variables in the Baseball data set. The Baseball data set includes 322 observations, and each observation has 24 variables. Among these 24 variables, the 5 nominal ones are selected as the input data to show an example of running k-modes clustering on a nominal data set. You can load the Sashelp.Baseball data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step.
data mycas.baseball;
Set sashelp.baseball;
Keep Team League Division Position Div;
Run;
These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
The following statements run the k-modes clustering algorithm with a frequency-based distance measure (DISTANCENOM=RELATIVEFREQ) and verify whether the clusters that the procedure obtains match the labels of the observations in the data table:
proc kclus data=mycas.baseball maxiter=10 maxc=5 DISTANCENOM=RELATIVEFREQ
outstat(outiter)=mycas.kclusOutstat2;
input Team League Division Position Div / level=nominal;
score out=mycas.kclusOut2 copyvars=(Team League Division Position Div);
ods output FreqNom=FreqNom1;
run;
Output 9.3.1 shows the cluster summary table that is produced for five clusters.
Output 9.3.1: Cluster Summary Table for Five Clusters
| Cluster Summary for Nominal Variables | |||||||
|---|---|---|---|---|---|---|---|
| Cluster | Frequency | Distance from Cluster Centroid to Observation | Within Cluster Distance | Nearest Cluster | Distance to Nearest Cluster Centroid | ||
| Minimum | Maximum | Average | |||||
| 1 | 72 | 1.6806 | 2.0000 | 1.9466 | 140.2 | 4 | 3.8824 |
| 2 | 75 | 1.6800 | 2.0000 | 1.9474 | 146.1 | 1 | 4.0000 |
| 3 | 90 | 1.7111 | 2.0000 | 1.9580 | 176.2 | 4 | 3.8824 |
| 4 | 85 | 1.7059 | 2.0000 | 1.9550 | 166.2 | 3 | 3.8667 |
Output 9.3.2 shows the frequencies of levels for the nominal input variable Team and information about how the levels of variables are distributed in each cluster; this information is important for revealing intracluster similarity. The following statement prints the observations from the frequency table, as shown in Output 9.3.2:
proc print noobs data=FreqNom1(obs=12);
run;
Output 9.3.2: Frequencies for Nominal Variables
| Variable | Level | FrequencyRead | _1 | _2 | _3 | _4 |
|---|---|---|---|---|---|---|
| Team | Atlanta | 11 | 0 | 11 | 0 | 0 |
| Team | Baltimore | 15 | 0 | 0 | 0 | 15 |
| Team | Boston | 10 | 0 | 0 | 0 | 10 |
| Team | California | 13 | 0 | 0 | 13 | 0 |
| Team | Chicago | 24 | 11 | 0 | 13 | 0 |
| Team | Cincinnati | 12 | 0 | 12 | 0 | 0 |
| Team | Cleveland | 12 | 0 | 0 | 0 | 12 |
| Team | Detroit | 12 | 0 | 0 | 0 | 12 |
| Team | Houston | 11 | 0 | 11 | 0 | 0 |
| Team | Kansas City | 14 | 0 | 0 | 14 | 0 |
| Team | Los Angeles | 14 | 0 | 14 | 0 | 0 |
| Team | Milwaukee | 14 | 0 | 0 | 0 | 14 |