SPARSEML Procedure

Getting Started: SPARSEML Procedure

Note: Input data must be in a CAS table that is accessible in your CAS session. You must refer to this table by using a two-level name. The first level must be a CAS engine libref, and the second level must be the table name. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.

The following DATA step creates the data table smldata. This data set contains one variable named vars and 15 observations; vars is a character variable. Each observation contains the data information, which includes a target and a sequence of combined column indexes and column values. The target value is either 1 or –1. The column index starts at 1.

data smldata;
    length vars $ 20.;
    input vars $ 1-20;
    datalines;
1 1:-1 2:3
1
1 1:1 2:1
1 1:2 2:2
1 1:3 2:3
1 1:4 2:4
1 1:5 2:5
-1 2:2
-1 1:1 2:3
-1 1:2 2:4
-1 1:3 2:5
1 1:0 2:-5
1 1:5
-1 1:0 2:5
-1 1:2 2:8
;
run;

You can load the smldata data set into your CAS session in the following DATA step:

data mylib.smldata;
    set smldata;
run;

These statements assume that your CAS engine libref is named mylib, but you can substitute any appropriately defined CAS engine libref.

The following statements use PROC SPARSEML to run the sparse machine learning algorithm on the mylib.smldata data table:

proc sparseml data= mylib.smldata;
    input vars;
run;

The INPUT statement defines the input variable vars, which contains both target and sparse input values in sparse string format.

PROC SPARSEML generates several ODS tables, some of which are shown in Figure 1 through Figure 3.

The "Data Information" table in Figure 1 shows that the number of observations is 15, the number of features is 2, and the number of sparse elements is 26.

Figure 1: Sparse Data Information

The SPARSEML Procedure

Data Information
Number of Rows15
Number of Features2
Number of Sparse Elements26


The "Misclassification Matrix" table in Figure 2 shows that among the total of fifteen observations, nine observations are classified as 1, and six observations are classified as –1. The number of correctly predicted 1 observations is eight, and the number of correctly predicted –1 observations is six. Thus the accuracy is 93.33%, as indicated in the "Fit Statistics" table in Figure 3.

Figure 2: Misclassification Matrix

Misclassification Matrix
ObservedTraining Prediction
1-1Total
1819
-1066
Total8715


Figure 3: Fit Statistics

Fit Statistics
StatisticTraining
Accuracy0.9333
Error0.0667
Sensitivity0.8889
Specificity1.0000


A relatively good model means that misclassification is low and both sensitivity and specificity are high.

Last updated: September 04, 2026