t-Distributed Stochastic Neighbor Embedding Action Set

Example 23.1 Computing a Two-Dimensional Embedding of Input Data

Iris Data

This section contains PROC CAS code.

Note: Input data must be in a CAS table that is accessible in your CAS session. This table has a two-level name; the first level is your CAS engine libref and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 3: Shared Concepts in SAS Visual Data Mining and Machine Learning 8.3: Procedures Guide. For more information about PROC CAS and programming in CASL, see Getting Started with CASL, SAS Cloud Analytic Services: CAS Procedure Programming Guide and Reference, and SAS Viya: System Programming Guide.

This example shows how to use the tSne action to compute a two-dimensional embedding of observations in an input data table. The example uses the Iris data from Fisher (1936), which contain morphological measurements of 50 specimens from each of three different species of iris flowers: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data. The tSne action returns a two-dimensional representation of each observation. The analysis uses four variables: SepalLength, SepalWidth, PetalLength, and PetalWidth. The remaining variables in the data table are not used.

You can load the Iris data into your CAS session by using the following DATA step. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.

data mycas.iris;
	set sashelp.iris;
	id=_n_;
run;

The following statements run the tSne action and create a two-dimensional embedding of the Iris data:

 proc cas;
   loadactionset "tSne";
   action tSne result=R / table={name="iris"},
       inputs        = {"SepalLength", "SepalWidth", "PetalLength", "PetalWidth"},
       nDimensions   = 2,
       perplexity    = 5,
       learningRate  = 100,
       maxIters      = 500,
       output        = {casOut={name="tsne_out", replace="TRUE"},
                        copyvars={"id", "species"}};
 run;

The table parameter names the input data table. The inputs parameter specifies the input variables to be used. The nDimensions parameter requests that the model return two embedding dimensions. The perplexity parameter specifies 5 as the value of the perplexity. The learningRate parameter specifies 100 as the learning rate for the optimization. The maxIters parameter specifies 500 as the maximum number of iterations. The output parameter outputs the embedding to the tsne_out data table and copies the ID and Species variables to the output.

Computing a Two-Dimensional Embedding of Input Data

This section contains Lua code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be loaded into CAS. One way to do this is to convert the Iris data to the comma-separated-value (CSV) file Iris.csv and then use the following code to load the CSV file into CAS:

   s:loadtable{casLib="casuser", path="Iris.csv"}

For more information about coding in Lua, see Getting Started with SAS Viya for Lua and SAS Viya: System Programming Guide.

This example shows how to use the tSne action to compute a two-dimensional embedding of observations in an input data table. The example uses the Iris data from Fisher (1936), which contain morphological measurements of 50 specimens from each of three different species of iris flowers: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data. The tSne action returns a two-dimensional representation of each observation. The analysis uses four variables: SepalLength, SepalWidth, PetalLength, and PetalWidth. The remaining variables in the data table are not used.

The following statements run the tSne action and create a two-dimensional embedding of the Iris data:

   s:loadactionset{actionset="tSne"}
   s:tSne{table={name="iris"},
     inputs        = {"SepalLength", "SepalWidth", "PetalLength", "PetalWidth"},
     nDimensions   = 2,
     perplexity    = 5,
     learningRate  = 100,
     maxIters      = 500,
     output        = {casOut={name="tsne_out", replace="TRUE"},
                      copyvars={"id", "species"}}};

   s:fetch{table={name="tsne_out"}}

The table parameter names the input data table. The inputs parameter specifies the input variables to be used. The nDimensions parameter requests that the model return two embedding dimensions. The perplexity parameter specifies 5 as the value of the perplexity. The learningRate parameter specifies 100 as the learning rate for the optimization. The maxIters parameter specifies 500 as the maximum number of iterations. The output parameter outputs the embedding to the tsne_out data table and copies the ID and Species variables to the output.

Computing a Two-Dimensional Embedding of Input Data

This section contains Python code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be loaded into CAS. One way to do this is to convert the Iris data to the comma-separated-value (CSV) file Iris.csv and then use the following code to load the CSV file into CAS:

   s.upload_file('Iris.csv')
  

For more information about coding in Python, see Getting Started with SAS Viya for Python and SAS Viya: System Programming Guide.

This example shows how to use the tSne action to compute a two-dimensional embedding of observations in an input data table. The example uses the Iris data from Fisher (1936), which contain morphological measurements of 50 specimens from each of three different species of iris flowers: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data. The tSne action returns a two-dimensional representation of each observation. The analysis uses four variables: SepalLength, SepalWidth, PetalLength, and PetalWidth. The remaining variables in the data table are not used.

The following statements run the tSne action and create a two-dimensional embedding of the Iris data:

	s.loadactionset{actionset="tSne"}
	s.tSne(output        = {"casOut":{"name":"out", "replace":True},
                            "copyVars":{"id", "species"}},
         inputs        = {"SepalLength", "SepalWidth", "PetalLength", "PetalWidth"},
         nDimensions   = 2,
         perplexity    = 5,
         learningRate  = 100,
         maxIters      = 500,
	         table         = {"name":"iris"})

	tsne_out=s.CASTable('out')
	pprint(tsne_out.fetch(to=20))

The table parameter names the input data table. The inputs parameter specifies the input variables to be used. The nDimensions parameter requests that the model return two embedding dimensions. The perplexity parameter specifies 5 as the value of the perplexity. The learningRate parameter specifies 100 as the learning rate for the optimization. The maxIters parameter specifies 500 as the maximum number of iterations. The output parameter outputs the embedding to the tsne_out data table and copies the ID and Species variables to the output.

Computing a Two-Dimensional Embedding of Input Data

This section contains R code for the analysis in the CASL version of this example, which contains details about the results.

Note: In order to run this code, the data that are described in the CASL version need to be loaded into CAS. One way to do this is to convert the Iris data to the comma-separated-value (CSV) file Iris.csv and then use the following code to load the CSV file into CAS:

   m <- cas.read.csv(s, "Iris.csv", casOut=list(name="Iris"))
  

For more information about coding in R, see Getting Started with SAS Viya for R and SAS Viya: System Programming Guide.

This example shows how to use the tSne action to compute a two-dimensional embedding of observations in an input data table. The example uses the Iris data from Fisher (1936), which contain morphological measurements of 50 specimens from each of three different species of iris flowers: Iris setosa, I. versicolor, and I. virginica. Mezzich and Solomon (1980) discuss a variety of cluster analyses that use the Iris data. The tSne action returns a two-dimensional representation of each observation. The analysis uses four variables: SepalLength, SepalWidth, PetalLength, and PetalWidth. The remaining variables in the data table are not used.

The following statements run the tSne action and create a two-dimensional embedding of the Iris data:

 loadActionSet(s,'tSne')
 rs <- cas.tSne.tSne(s,
    table         = list(name = "hmeq"),
    inputs        = list("SepalLength", "SepalWidth", "PetalLength", "PetalWidth"),
    nDimensions   = 2,
    perplexity    = 5,
    learningRate  = 100,
    maxIters      = 500,
    output        = list(casout = list(name = "tsne_out",
                                       replace = TRUE))),

The table parameter names the input data table. The inputs parameter specifies the input variables to be used. The nDimensions parameter requests that the model return two embedding dimensions. The perplexity parameter specifies 5 as the value of the perplexity. The learningRate parameter specifies 100 as the learning rate for the optimization. The maxIters parameter specifies 500 as the maximum number of iterations. The output parameter outputs the embedding to the tsne_out data table and copies the ID and Species variables to the output.

Last updated: June 07, 2018