The CANDISC Procedure

Example 31.1 Analyzing Iris Data by Using PROC CANDISC

(View the complete code for this example.)

The iris data that were published by Fisher (1936) have been widely used for examples in discriminant analysis and cluster analysis. The sepal length, sepal width, petal length, and petal width are measured in millimeters in 50 iris specimens from each of three species: Iris setosa, I. versicolor, and I. virginica. The iris data set is available from the Sashelp library.

This example is a canonical discriminant analysis that creates an output data set that contains scores on the canonical variables and plots the canonical variables.

The following statements produce Output 31.1.1 through Output 31.1.6:

title 'Fisher (1936) Iris Data';

proc candisc data=sashelp.iris out=outcan distance anova;
   class Species;
   var SepalLength SepalWidth PetalLength PetalWidth;
run;

PROC CANDISC first displays information about the observations and the classes in the data set in Output 31.1.1.

Output 31.1.1: Iris Data: Summary Information

Fisher (1936) Iris Data

The CANDISC Procedure

Total Sample Size150DF Total149
Variables4DF Within Classes147
Classes3DF Between Classes2

Number of Observations Read150
Number of Observations Used150

Class Level Information
SpeciesVariable
Name
FrequencyWeightProportion
SetosaSetosa5050.00000.333333
VersicolorVersicolor5050.00000.333333
VirginicaVirginica5050.00000.333333


The DISTANCE option in the PROC CANDISC statement displays squared Mahalanobis distances between class means. Results from the DISTANCE option are shown in Output 31.1.2.

Output 31.1.2: Iris Data: Squared Mahalanobis Distances and Distance Statistics

Fisher (1936) Iris Data

The CANDISC Procedure

Squared Distance to Species
From SpeciesSetosa VersicolorVirginica
Setosa089.86419179.38471
Versicolor89.86419017.20107
Virginica179.3847117.201070

F Statistics, NDF=4, DDF=144 for Squared Distance
to Species
From SpeciesSetosa VersicolorVirginica
Setosa0550.188891098
Versicolor550.188890105.31265
Virginica1098105.312650

Prob > Mahalanobis Distance for Squared Distance
to Species
From SpeciesSetosa VersicolorVirginica
Setosa1.0000<.0001<.0001
Versicolor<.00011.0000<.0001
Virginica<.0001<.00011.0000


Output 31.1.3 displays univariate and multivariate statistics. The ANOVA option uses univariate statistics to test the hypothesis that the class means are equal. The resulting R-square values range from 0.4008 for SepalWidth to 0.9414 for PetalLength, and each variable is significant at the 0.0001 level. The multivariate test for differences between the classes (which is displayed by default) is also significant at the 0.0001 level; you would expect this from the highly significant univariate test results.

Output 31.1.3: Iris Data: Univariate and Multivariate Statistics

Fisher (1936) Iris Data

The CANDISC Procedure

Univariate Test Statistics
F Statistics, Num DF=2, Den DF=147
VariableLabelTotal
Standard
Deviation
Pooled
Standard
Deviation
Between
Standard
Deviation
R-SquareR-Square
/ (1-RSq)
F ValuePr > F
SepalLengthSepal Length (mm)8.28075.14797.95060.61871.6226119.26<.0001
SepalWidthSepal Width (mm)4.35873.39693.36820.40080.668849.16<.0001
PetalLengthPetal Length (mm)17.65304.303320.90700.941416.05661180.16<.0001
PetalWidthPetal Width (mm)7.62242.04658.96730.928913.0613960.01<.0001

Average R-Square
Unweighted0.7224358
Weighted by Variance0.8689444

Multivariate Statistics and F Approximations
S=2 M=0.5 N=71
StatisticValueF ValueNum DFDen DFPr > F
Wilks' Lambda0.02343863199.158288<.0001
Pillai's Trace1.1918988353.478290<.0001
Hotelling-Lawley Trace32.47732024582.208203.4<.0001
Roy's Greatest Root32.191929201166.964145<.0001
NOTE: F Statistic for Roy's Greatest Root is an upper bound.
NOTE: F Statistic for Wilks' Lambda is exact.


Output 31.1.4 displays canonical correlations and eigenvalues. The R square between Can1 and the CLASS variable, 0.969872, is much larger than the corresponding R square for Can2, 0.222027.

Output 31.1.4: Iris Data: Canonical Correlations and Eigenvalues

Fisher (1936) Iris Data

The CANDISC Procedure

 Canonical
Correlation
Adjusted
Canonical
Correlation
Approximate
Standard
Error
Squared
Canonical
Correlation
Eigenvalues of Inv(E)*H
= CanRsq/(1-CanRsq)
Test of H0: The canonical correlations in the current row and all that follow are zero
 EigenvalueDifferenceProportionCumulativeLikelihood
Ratio
Approximate
F Value
Num DFDen DFPr > F
10.9848210.9845080.0024680.96987232.191931.90650.99120.99120.02343863199.158288<.0001
20.4711970.4614450.0637340.2220270.2854 0.00881.00000.7779733713.793145<.0001


Output 31.1.5 displays correlations between canonical and original variables.

Output 31.1.5: Iris Data: Correlations between Canonical and Original Variables

Fisher (1936) Iris Data

The CANDISC Procedure

Total Canonical Structure
VariableLabelCan1Can2
SepalLengthSepal Length (mm)0.7918880.217593
SepalWidthSepal Width (mm)-0.5307590.757989
PetalLengthPetal Length (mm)0.9849510.046037
PetalWidthPetal Width (mm)0.9728120.222902

Between Canonical Structure
VariableLabelCan1Can2
SepalLengthSepal Length (mm)0.9914680.130348
SepalWidthSepal Width (mm)-0.8256580.564171
PetalLengthPetal Length (mm)0.9997500.022358
PetalWidthPetal Width (mm)0.9940440.108977

Pooled Within Canonical Structure
VariableLabelCan1Can2
SepalLengthSepal Length (mm)0.2225960.310812
SepalWidthSepal Width (mm)-0.1190120.863681
PetalLengthPetal Length (mm)0.7060650.167701
PetalWidthPetal Width (mm)0.6331780.737242


Output 31.1.6 displays canonical coefficients. The raw canonical coefficients for the first canonical variable, Can1, show that the classes differ most widely on the linear combination of the centered variables: .

Output 31.1.6: Iris Data: Canonical Coefficients

Fisher (1936) Iris Data

The CANDISC Procedure

Total-Sample Standardized Canonical Coefficients
VariableLabelCan1Can2
SepalLengthSepal Length (mm)-0.6867795330.019958173
SepalWidthSepal Width (mm)-0.6688250750.943441829
PetalLengthPetal Length (mm)3.885795047-1.645118866
PetalWidthPetal Width (mm)2.1422387152.164135931

Pooled Within-Class Standardized Canonical Coefficients
VariableLabelCan1Can2
SepalLengthSepal Length (mm)-.42695484860.0124075316
SepalWidthSepal Width (mm)-.52124167580.7352613085
PetalLengthPetal Length (mm)0.9472572487-.4010378190
PetalWidthPetal Width (mm)0.57516077190.5810398645

Raw Canonical Coefficients
VariableLabelCan1Can2
SepalLengthSepal Length (mm)-.08293776420.0024102149
SepalWidthSepal Width (mm)-.15344730680.2164521235
PetalLengthPetal Length (mm)0.2201211656-.0931921210
PetalWidthPetal Width (mm)0.28104603090.2839187853


Output 31.1.7 displays class level means on canonical variables.

Output 31.1.7: Iris Data: Canonical Means

Class Means on Canonical Variables
SpeciesCan1Can2
Setosa-7.6075999270.215133017
Versicolor1.825049490-0.727899622
Virginica5.7825504370.512766605


The TEMPLATE and SGRENDER procedures are used to create a plot of the first two canonical variables. The following statements produce Output 31.1.8:

proc template;
   define statgraph scatter;
      begingraph / attrpriority=none;
         entrytitle 'Fisher (1936) Iris Data';
         layout overlayequated / equatetype=fit
            xaxisopts=(label='Canonical Variable 1')
            yaxisopts=(label='Canonical Variable 2');
            scatterplot x=Can1 y=Can2 / group=species name='iris'
                                        markerattrs=(size=3px);
            layout gridded / autoalign=(topright topleft);
               discretelegend 'iris' / border=false opaque=false;
            endlayout;
         endlayout;
      endgraph;
   end;
run;

proc sgrender data=outcan template=scatter;
run;

Output 31.1.8: Iris Data: Plot of First Two Canonical Variables

Iris Data: Plot of First Two Canonical Variables


The plot of canonical variables in Output 31.1.8 shows that of the two canonical variables, Can1 has more discriminatory power.