The CORRESP Procedure

Example 36.1 Simple and Multiple Correspondence Analysis of Automobiles and Their Owners

(View the complete code for this example.)

In this example, PROC CORRESP creates a contingency table from categorical data and performs a simple correspondence analysis. The data are from a sample of individuals who were asked to provide information about themselves and their automobiles. The questions included origin of the automobile (American, Japanese, or European) and family status (single, married, single living with children, or married living with children).

The first steps read the input data and assign formats:

title1 'Automobile Owners and Auto Attributes';
title2 'Simple Correspondence Analysis';

proc format;
   value Origin  1 = 'American' 2 = 'Japanese' 3 = 'European';
   value Size    1 = 'Small'    2 = 'Medium'   3 = 'Large';
   value Type    1 = 'Family'   2 = 'Sporty'   3 = 'Work';
   value Home    1 = 'Own'      2 = 'Rent';
   value Sex     1 = 'Male'     2 = 'Female';
   value Income  1 = '1 Income' 2 = '2 Incomes';
   value Marital 1 = 'Single with Kids' 2 = 'Married with Kids'
                 3 = 'Single'           4 = 'Married';
run;

data Cars;
   missing a;
   input (Origin Size Type Home Income Marital Kids Sex) (1.) @@;
   * Check for End of Line;
   if n(of Origin -- Sex) eq 0 then do; input; return; end;
   marital = 2 * (kids le 0) + marital;
   format Origin Origin. Size Size. Type Type. Home Home.
          Sex Sex. Income Income. Marital Marital.;
   output;
   datalines;
131112212121110121112201131211011211221122112121131122123211222212212201
121122023121221232211101122122022121110122112102131112211121110112311101
211112113211223121122202221122111311123131211102321122223221220221221101
122122022121220211212201221122021122110132112202213112111331226122221101

   ... more lines ...   

212122011211122131221101121211022212220212121101
;

PROC CORRESP is used to perform the simple correspondence analysis. The ALL option displays all tables, including the contingency table, chi-square information, profiles, and all results of the correspondence analysis. The CHI2P option displays the chi-square p-value. The TABLES statement specifies the row and column categorical variables. The results are displayed with ODS Graphics.

The following statements produce Output 36.1.1:

ods graphics on;

* Perform Simple Correspondence Analysis;
proc corresp data=Cars all chi2p;
   tables Marital, Origin;
run;

Correspondence analysis locates all the categories in a Euclidean space. The first two dimensions of this space are plotted to examine the associations among the categories.[29] Because the smallest dimension of this table is three, there is no loss of information when only two dimensions are plotted. The plot should be thought of as two different overlaid plots, one for each categorical variable. Distances between points within a variable have meaning, but distances between points from different variables do not.

Output 36.1.1: Simple Correspondence Analysis

Automobile Owners and Auto Attributes
Simple Correspondence Analysis

The CORRESP Procedure

Contingency Table
 AmericanEuropeanJapaneseSum
Married371451102
Married with Kids521544111
Single331563111
Single with Kids61815
Sum12845166339

Chi-Square Statistic Expected Values
 AmericanEuropeanJapanese
Married38.513313.539849.9469
Married with Kids41.911514.734554.3540
Single41.911514.734554.3540
Single with Kids5.66371.99127.3451

Observed Minus Expected Values
 AmericanEuropeanJapanese
Married-1.51330.46021.0531
Married with Kids10.08850.2655-10.3540
Single-8.91150.26558.6460
Single with Kids0.3363-0.99120.6549

Contributions to the Total Chi-Square Statistic
 AmericanEuropeanJapaneseSum
Married0.059460.015640.022200.09730
Married with Kids2.428400.004781.972354.40553
Single1.894820.004781.375313.27492
Single with Kids0.019970.493370.058390.57173
Sum4.402650.518583.428258.34947

Row Profiles
 AmericanEuropeanJapanese
Married0.3627450.1372550.500000
Married with Kids0.4684680.1351350.396396
Single0.2972970.1351350.567568
Single with Kids0.4000000.0666670.533333

Column Profiles
 AmericanEuropeanJapanese
Married0.2890630.3111110.307229
Married with Kids0.4062500.3333330.265060
Single0.2578130.3333330.379518
Single with Kids0.0468750.0222220.048193


crse1ab

Row Coordinates
 Dim1Dim2
Married-0.02780.0134
Married with Kids0.19910.0064
Single-0.17160.0076
Single with Kids-0.0144-0.1947

Summary Statistics for the Row Points
 QualityMassInertia
Married1.00000.30090.0117
Married with Kids1.00000.32740.5276
Single1.00000.32740.3922
Single with Kids1.00000.04420.0685

Partial Contributions to Inertia for the Row
Points
 Dim1Dim2
Married0.01020.0306
Married with Kids0.56780.0076
Single0.42170.0108
Single with Kids0.00040.9511

Indices of the Coordinates That Contribute Most to Inertia for the Row Points
 Dim1Dim2Best
Married002
Married with Kids101
Single101
Single with Kids022

Squared Cosines for the Row Points
 Dim1Dim2
Married0.81210.1879
Married with Kids0.99900.0010
Single0.99800.0020
Single with Kids0.00540.9946

Column Coordinates
 Dim1Dim2
American0.1847-0.0166
European0.00130.1073
Japanese-0.1428-0.0163

Summary Statistics for the Column
Points
 QualityMassInertia
American1.00000.37760.5273
European1.00000.13270.0621
Japanese1.00000.48970.4106

Partial Contributions to Inertia
for the Column Points
 Dim1Dim2
American0.56340.0590
European0.00000.8672
Japanese0.43660.0737

Indices of the Coordinates That Contribute
Most to Inertia for the Column Points
 Dim1Dim2Best
American101
European022
Japanese101

Squared Cosines for the Column
Points
 Dim1Dim2
American0.99200.0080
European0.00010.9999
Japanese0.98710.0129

crse1g

To interpret the plot, start by interpreting the row points separately from the column points. The European point is near and to the left of the centroid, so it makes a relatively small contribution to the chi-square statistic (because it is near the centroid), it contributes almost nothing to the inertia of dimension one (because its coordinate on dimension one has a small absolute value relative to the other column points), and it makes a relatively large contribution to the inertia of dimension two (because its coordinate on dimension two has a large absolute value relative to the other column points). Its squared cosines for dimension one and two, approximately 0 and 1, respectively, indicate that its position is almost completely determined by its location on dimension two. Its quality of display is 1.0, indicating perfect quality, because the table is two-dimensional after the centering. The American and Japanese points are far from the centroid, and they lie along dimension one. They make relatively large contributions to the chi-square statistic and the inertia of dimension one. The horizontal dimension seems to be largely determined by Japanese versus American automobile ownership.

In the row points, the Married point is near the centroid, and the Single with Kids point has a small coordinate on dimension one that is near zero. The horizontal dimension seems to be largely determined by the Single versus the Married with Kids points. The two interpretations of dimension one show the association with being Married with Kids and owning an American auto, and being single and owning a Japanese auto. The fact that the Married with Kids point is close to the American point and the fact that the Japanese point is near the Single point should be ignored. Distances between row and column points are not defined. The plot shows that more people who are married with kids than you would expect if the rows and columns were independent drive an American auto, and more people who are single than you would expect if the rows and columns were independent drive a Japanese auto.

In the second part of this example, PROC CORRESP creates a Burt table from categorical data and performs a multiple correspondence analysis. The variables used in this example are Origin, Size, Type, Income, Home, Marital, and Sex. MCA specifies multiple correspondence analysis, OBSERVED displays the Burt table. The TABLES statement with only a single variable list and no comma creates the Burt table.

The following statements produce Output 36.1.2:

title2 'Multiple Correspondence Analysis';

* Perform Multiple Correspondence Analysis;
proc corresp mca observed data=Cars;
   tables Origin Size Type Income Home Marital Sex;
run;

Output 36.1.2: Multiple Correspondence Analysis

Automobile Owners and Auto Attributes
Multiple Correspondence Analysis

The CORRESP Procedure

Burt Table
 AmericanEuropeanJapaneseLargeMediumSmallFamilySportyWork1 Income2 IncomesOwnRentMarriedMarried
with Kids
SingleSingle with
Kids
FemaleMale
American125003660298124205867933237503265867
European04404202017234182638613151512123
Japanese0016526110276593074911115451446287095
Large364242003011120223579211111725
Medium6020610141089391357841063542514087071
Small29201020015155663073781015050375866289
Family811776308955174006910513044507935108391
Sporty24235913966010605551713535125724462
Work2043011133000542628411316181732232
1 Income581874205773695526150080701027991447103
2 Incomes6726912284781055128018416222918210110282
Own933811135106101130714180162242076106528114128
Rent326547355044351370220922535773557
Married37135194250503516109176251010005348
Married with Kids501544215137791218278210630109004861
Single321562114058355717991052570010903574
Single with Kids61818610231418700015132
Female5821701770628344224710211435534835131490
Male672395257189916232103821285748617420185


crse1cc

Column Coordinates
 Dim1Dim2
American-0.40350.8129
European-0.0568-0.5552
Japanese0.3208-0.4678
Large-0.69491.5666
Medium-0.25620.0965
Small0.4326-0.5258
Family-0.42010.3602
Sporty0.6604-0.6696
Work0.05750.1539
1 Income0.82510.5472
2 Incomes-0.6727-0.4461
Own-0.3887-0.0943
Rent1.02250.2480
Married-0.4169-0.7954
Married with Kids-0.82000.3237
Single1.14610.2930
Single with Kids0.43730.8736
Female-0.3365-0.2057
Male0.27100.1656

Summary Statistics for the Column Points
 QualityMassInertia
American0.49250.05350.0521
European0.04730.01880.0724
Japanese0.31410.07060.0422
Large0.42240.01800.0729
Medium0.05480.06030.0482
Small0.38250.06460.0457
Family0.33300.07440.0399
Sporty0.41120.04530.0569
Work0.00520.02310.0699
1 Income0.79910.06420.0459
2 Incomes0.79910.07870.0374
Own0.42080.10350.0230
Rent0.42080.03930.0604
Married0.34960.04320.0581
Married with Kids0.37650.04660.0561
Single0.67800.04660.0561
Single with Kids0.04490.00640.0796
Female0.12530.06370.0462
Male0.12530.07910.0372

Partial Contributions to Inertia for the Column
Points
 Dim1Dim2
American0.02680.1511
European0.00020.0248
Japanese0.02240.0660
Large0.02680.1886
Medium0.01220.0024
Small0.03730.0764
Family0.04050.0413
Sporty0.06100.0870
Work0.00020.0023
1 Income0.13480.0822
2 Incomes0.10990.0670
Own0.04820.0039
Rent0.12690.0103
Married0.02320.1169
Married with Kids0.09670.0209
Single0.18890.0171
Single with Kids0.00380.0209
Female0.02230.0115
Male0.01790.0093

Indices of the Coordinates That Contribute Most to Inertia for the Column Points
 Dim1Dim2Best
American022
European002
Japanese022
Large022
Medium001
Small022
Family202
Sporty222
Work002
1 Income111
2 Incomes111
Own101
Rent101
Married022
Married with Kids101
Single101
Single with Kids002
Female001
Male001

Squared Cosines for the Column Points
 Dim1Dim2
American0.09740.3952
European0.00050.0468
Japanese0.10050.2136
Large0.06950.3530
Medium0.04800.0068
Small0.15440.2281
Family0.19190.1411
Sporty0.20270.2085
Work0.00060.0046
1 Income0.55500.2441
2 Incomes0.55500.2441
Own0.39750.0234
Rent0.39750.0234
Married0.07530.2742
Married with Kids0.32580.0508
Single0.63640.0416
Single with Kids0.00900.0359
Female0.09120.0341
Male0.09120.0341

crse1m

Multiple correspondence analysis locates all the categories in a Euclidean space. The first two dimensions of this space are plotted to examine the associations among the categories. The top-right quadrant of the plot shows that the categories Single, Single with Kids, 1 Income, and Rent are associated. Proceeding clockwise, the categories Sporty, Small, and Japanese are associated. The bottom-left quadrant shows the association between being married, owning your own home, and having two incomes. Having children is associated with owning a large American family auto. Such information could be used in market research to identify target audiences for advertisements.

This interpretation is based on points found in approximately the same direction from the origin and in approximately the same region of the space. Distances between points do not have a straightforward interpretation in multiple correspondence analysis. The geometry of multiple correspondence analysis is not a simple generalization of the geometry of simple correspondence analysis (Greenacre and Hastie 1987; Greenacre 1988).

If you want to perform a multiple correspondence analysis and get scores for the individuals, you can specify the BINARY option to analyze the binary table, as in the following statements. In the interest of space, only the first 10 rows of coordinates are printed in Output 36.1.3.

title2 'Binary Table';

* Perform Multiple Correspondence Analysis;
proc corresp data=Cars binary;
   ods select RowCoors;
   tables Origin Size Type Income Home Marital Sex;
run;

Output 36.1.3: Correspondence Analysis of a Binary Table

Automobile Owners and Auto Attributes
Binary Table
 
The CORRESP Procedure
 
Row Coordinates

 Dim1Dim2
1-0.40931.0878
20.8198-0.2221
3-0.2193-0.5328
40.43821.1799
5-0.67500.3600
6-0.17780.1441
7-0.93750.6846
8-0.7405-0.1539
9-0.3027-0.2749
10-0.7263-0.0803




[29] In this analysis, the chi-square statistic is not significantly different from 0. Hence, you would not reject the null hypothesis that the rows and columns are independent. If your goal is hypothesis testing, you might stop at this point and not proceed to interpret the graphical and tabular results. If your goal is exploratory data analysis, you might proceed and interpret the results. This example will proceed, because it is intended to be a small, simple teaching example.

Last updated: February 13, 2019