The KCLUS Procedure

Example 9.3 Clustering Nominal Variables

In this example, PROC KCLUS clusters nominal variables in the Baseball data set. The Baseball data set includes 322 observations, and each observation has 24 variables. Among these 24 variables, the 5 nominal ones are selected as the input data to show an example of running k-modes clustering on a nominal data set. You can load the Sashelp.Baseball data set into your CAS session by naming your CAS engine libref in the first statement of the following DATA step.

data mycas.baseball;
  Set sashelp.baseball;
  Keep Team League Division Position Div;
Run;

These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

The following statements run the k-modes clustering algorithm with a frequency-based distance measure (DISTANCENOM=RELATIVEFREQ) and verify whether the clusters that the procedure obtains match the labels of the observations in the data table:


proc kclus data=mycas.baseball maxiter=10 maxc=5 DISTANCENOM=RELATIVEFREQ
   outstat(outiter)=mycas.kclusOutstat2;
input Team League Division Position Div / level=nominal;
score out=mycas.kclusOut2 copyvars=(Team League Division Position Div);
ods output FreqNom=FreqNom1;
run;

Output 9.3.1 shows the cluster summary table that is produced for five clusters.

Output 9.3.1: Cluster Summary Table for Five Clusters

The KCLUS Procedure

Cluster Summary for Nominal Variables
ClusterFrequencyDistance from Cluster Centroid
to Observation
Within
Cluster
Distance
Nearest
Cluster
Distance
to
Nearest
Cluster
Centroid
MinimumMaximumAverage
1721.68062.00001.9466140.243.8824
2751.68002.00001.9474146.114.0000
3901.71112.00001.9580176.243.8824
4851.70592.00001.9550166.233.8667


Output 9.3.2 shows the frequencies of levels for the nominal input variable Team and information about how the levels of variables are distributed in each cluster; this information is important for revealing intracluster similarity. The following statement prints the observations from the frequency table, as shown in Output 9.3.2:

proc print noobs data=FreqNom1(obs=12);
run;

Output 9.3.2: Frequencies for Nominal Variables

VariableLevelFrequencyRead_1_2_3_4
TeamAtlanta1101100
TeamBaltimore1500015
TeamBoston1000010
TeamCalifornia1300130
TeamChicago24110130
TeamCincinnati1201200
TeamCleveland1200012
TeamDetroit1200012
TeamHouston1101100
TeamKansas City1400140
TeamLos Angeles1401400
TeamMilwaukee1400014


Last updated: April 08, 2021