The NETWORK Procedure

Example 2.21 Learning Word Embeddings from Synonym Pairs

This example illustrates how to learn word embeddings by using node similarity. Consider an undirected graph where the nodes are a set of five terms and the links are synonym relationships between pairs of terms. The assigned weight on each link represents the strength of the synonym relationship. You can compute word embeddings for these terms by using vector node similarity. The graph is shown in Figure 240.

Figure 240: Synonym Pairs Graph

Synonym Pairs Graph


You can represent the graph by using the following links data table, mycas.LinkSetIn:

data mycas.SynonymLinks;
   input from $ to $ weight;
   datalines;
book mp3 0.1
dvd music 0.1
dvd video 0.95
mp3 music 0.9
mp3 video 0.1
music video 0.3
;

In practice, this input graph would typically be based on co-occurrence frequency in a large corpus and could contain millions of nodes. Word embeddings provide a compressed representation of the information in the original graph and are often inputs to low-dimensional visualizations and machine learning pipelines.

The following statements use the VECTOR=TRUE option to generate word embeddings for the terms in the data set:

%let nDim=5;
%let nSamples=100000;
%let convergenceThreshold=0.001;
proc network
   links                   = mycas.SynonymLinks
   outNodes                = mycas.WordEmbeddings;
   nodeSimilarity
      vector               = true
      jaccard              = false
      proximityOrder       = first
      nDimensions          = &nDim
      nSamples             = &nSamples
      convergenceThreshold = &convergenceThreshold
      outSimilarity        = mycas.OutSimilarity;
run;

The output data table mycas.WordEmbeddings contains the five-element embedding vector for each node. The word embedding vector for each of the five terms is shown in Output 2.21.1.

Output 2.21.1: Embedding Vectors Output

nodevec_1vec_2vec_3vec_4vec_5
book-0.21486-0.422980.14155-0.848030.18904
dvd-0.43874-0.46665-0.485340.13360-0.57995
mp3-0.34548-0.11227-0.485670.119790.78601
music-0.08957-0.04459-0.789930.047400.60312
video-0.35923-0.31128-0.712430.01254-0.51608


The output data table mycas.OutSimilarity contains the pairwise vector similarity between terms, which excludes redundant node pairs where source comes lexicographically after sink. For display purposes, it is helpful to create a symmetric results table that also includes pairs where source comes lexicographically after sink. You can create the table mycas.OutSimilaritySymmetric, shown in Output 2.21.2, by using the following DATA step:

data mycas.OutSimilaritySymmetric;
   set mycas.OutSimilarity;
   output;
   if source < sink then do;
      temp = source;
      source = sink;
      sink = temp;
      output;
   end;
   drop temp;
run;

Output 2.21.2: Vector Similarity Output

sinklinkvector
book01.00000
mp310.54999
music00.50006
dvd00.50001
video00.49991

sinklinkvector
dvd01.00000
video10.97481
music10.55002
book00.50001
mp300.49992

sinklinkvector
mp301.00000
music10.94967
video10.55046
book10.54999
dvd00.49992

sinklinkvector
music01.00000
mp310.94967
video10.64908
dvd10.55002
book00.50006

sinklinkvector
video01.00000
dvd10.97481
music10.64908
mp310.55046
book00.49991


To validate the quality of the word embeddings, you can use the scalar product between embedding vector pairs to reconstruct the original graph. Note that although vector scalar products range from –1 to 1, the vector similarity variable is normalized to lie in the range 0 to 1. Unassociated words, therefore, are assigned a vector similarity score of 0.5.

To reconstruct an approximation of the original graph, you can scale the scores back to the range –1 to 1 and discard any pairs whose weight is less than 0.001. You can create the table mycas.ReconstructedLinks by using the following DATA step:

data mycas.ReconstructedLinks;
   set mycas.OutSimilarity(where=(source < sink));
   weight = vector*2-1;
   keep source sink weight;
   if weight >= 0.001;
   rename source=from sink=to;
run;

For ease of comparison, you can join the original link weights to the mycas.ReconstructedLinks table, shown in Output 2.21.3. The following statements use PROC FEDSQL to perform this join:

proc fedsql sessref=mysess;
   create table ReconstructedLinks {options replace=true} as
   select a.*, b.weight as "originalWeight"
   from ReconstructedLinks as a
   join SynonymLinks as b
   on a.from = b.from and a.to = b.to;
quit;

Output 2.21.3: Links Reconstructed from Vector Similarity

fromtoweightoriginalWeight
bookmp30.099970.10
dvdmusic0.100040.10
dvdvideo0.949610.95
mp3music0.899330.90
mp3video0.100920.10
musicvideo0.298160.30


The reconstructed graph is shown in Figure 241.

Figure 241: Reconstructed Synonym Pairs Graph

Reconstructed Synonym Pairs Graph


Last updated: November 22, 2022