The NETWORK Procedure
Example 2.21 Learning Word Embeddings from Synonym Pairs
This example illustrates how to learn word embeddings by using node similarity. Consider an undirected graph where the nodes are a set of five terms and the links are synonym relationships between pairs of terms. The assigned weight on each link represents the strength of the synonym relationship. You can compute word embeddings for these terms by using vector node similarity. The graph is shown in Figure 240.
Figure 240: Synonym Pairs Graph

You can represent the graph by using the following links data table, mycas.LinkSetIn:
data mycas.SynonymLinks;
input from $ to $ weight;
datalines;
book mp3 0.1
dvd music 0.1
dvd video 0.95
mp3 music 0.9
mp3 video 0.1
music video 0.3
;
In practice, this input graph would typically be based on co-occurrence frequency in a large corpus and could contain millions of nodes. Word embeddings provide a compressed representation of the information in the original graph and are often inputs to low-dimensional visualizations and machine learning pipelines.
The following statements use the VECTOR=TRUE option to generate word embeddings for the terms in the data set:
%let nDim=5;
%let nSamples=100000;
%let convergenceThreshold=0.001;
proc network
links = mycas.SynonymLinks
outNodes = mycas.WordEmbeddings;
nodeSimilarity
vector = true
jaccard = false
proximityOrder = first
nDimensions = &nDim
nSamples = &nSamples
convergenceThreshold = &convergenceThreshold
outSimilarity = mycas.OutSimilarity;
run;
The output data table mycas.WordEmbeddings contains the five-element embedding vector for each node. The word embedding vector for each of the five terms is shown in Output 2.21.1.
Output 2.21.1: Embedding Vectors Output
| node | vec_1 | vec_2 | vec_3 | vec_4 | vec_5 |
|---|---|---|---|---|---|
| book | -0.21486 | -0.42298 | 0.14155 | -0.84803 | 0.18904 |
| dvd | -0.43874 | -0.46665 | -0.48534 | 0.13360 | -0.57995 |
| mp3 | -0.34548 | -0.11227 | -0.48567 | 0.11979 | 0.78601 |
| music | -0.08957 | -0.04459 | -0.78993 | 0.04740 | 0.60312 |
| video | -0.35923 | -0.31128 | -0.71243 | 0.01254 | -0.51608 |
The output data table mycas.OutSimilarity contains the pairwise vector similarity between terms, which excludes redundant node pairs where source comes lexicographically after sink. For display purposes, it is helpful to create a symmetric results table that also includes pairs where source comes lexicographically after sink. You can create the table mycas.OutSimilaritySymmetric, shown in Output 2.21.2, by using the following DATA step:
data mycas.OutSimilaritySymmetric;
set mycas.OutSimilarity;
output;
if source < sink then do;
temp = source;
source = sink;
sink = temp;
output;
end;
drop temp;
run;
Output 2.21.2: Vector Similarity Output
| sink | link | vector |
|---|---|---|
| book | 0 | 1.00000 |
| mp3 | 1 | 0.54999 |
| music | 0 | 0.50006 |
| dvd | 0 | 0.50001 |
| video | 0 | 0.49991 |
| sink | link | vector |
|---|---|---|
| dvd | 0 | 1.00000 |
| video | 1 | 0.97481 |
| music | 1 | 0.55002 |
| book | 0 | 0.50001 |
| mp3 | 0 | 0.49992 |
| sink | link | vector |
|---|---|---|
| mp3 | 0 | 1.00000 |
| music | 1 | 0.94967 |
| video | 1 | 0.55046 |
| book | 1 | 0.54999 |
| dvd | 0 | 0.49992 |
| sink | link | vector |
|---|---|---|
| music | 0 | 1.00000 |
| mp3 | 1 | 0.94967 |
| video | 1 | 0.64908 |
| dvd | 1 | 0.55002 |
| book | 0 | 0.50006 |
| sink | link | vector |
|---|---|---|
| video | 0 | 1.00000 |
| dvd | 1 | 0.97481 |
| music | 1 | 0.64908 |
| mp3 | 1 | 0.55046 |
| book | 0 | 0.49991 |
To validate the quality of the word embeddings, you can use the scalar product between embedding vector pairs to reconstruct the original graph. Note that although vector scalar products range from –1 to 1, the vector similarity variable is normalized to lie in the range 0 to 1. Unassociated words, therefore, are assigned a vector similarity score of 0.5.
To reconstruct an approximation of the original graph, you can scale the scores back to the range –1 to 1 and discard any pairs whose weight is less than 0.001. You can create the table mycas.ReconstructedLinks by using the following DATA step:
data mycas.ReconstructedLinks;
set mycas.OutSimilarity(where=(source < sink));
weight = vector*2-1;
keep source sink weight;
if weight >= 0.001;
rename source=from sink=to;
run;
For ease of comparison, you can join the original link weights to the mycas.ReconstructedLinks table, shown in Output 2.21.3. The following statements use PROC FEDSQL to perform this join:
proc fedsql sessref=mysess;
create table ReconstructedLinks {options replace=true} as
select a.*, b.weight as "originalWeight"
from ReconstructedLinks as a
join SynonymLinks as b
on a.from = b.from and a.to = b.to;
quit;
Output 2.21.3: Links Reconstructed from Vector Similarity
| from | to | weight | originalWeight |
|---|---|---|---|
| book | mp3 | 0.09997 | 0.10 |
| dvd | music | 0.10004 | 0.10 |
| dvd | video | 0.94961 | 0.95 |
| mp3 | music | 0.89933 | 0.90 |
| mp3 | video | 0.10092 | 0.10 |
| music | video | 0.29816 | 0.30 |
The reconstructed graph is shown in Figure 241.
Figure 241: Reconstructed Synonym Pairs Graph
