The NETWORK Procedure

Example 2.12 Connected Components for US Patent Citations

This example looks at the structural relationship of US patent citations by using a large data set that is maintained by the Stanford Network Analysis Project (SNAP) (Leskovec 2014). The citation graph includes over 16 million citations made to patents between 1975 and 1999.

The following statements construct the links data table mycas.Patents from a local copy of the raw patent citation data:

filename in 'cit-Patents.txt';
data mycas.Patents;
   infile in firstobs=5 dlm='09'X;
   input from to;
run;

The following statements find the connected components of the citation graph by using a distributed union-find algorithm. This algorithm takes advantage of all the machines in your configured session.

proc network
   links        = mycas.Patents
   outNodes     = mycas.NodeSetOut;
   connectedComponents
      out       = mycas.ConCompOut
      algorithm = parallel;
run;
%put &_NETWORK_;

The progress of the procedure is shown in Output 2.12.1.

Output 2.12.1: PROC NETWORK Log: Connected Components for US Patent Citations

NOTE: ------------------------------------------------------------------------------------------
NOTE: Running NETWORK.                                                                          
NOTE: ------------------------------------------------------------------------------------------
NOTE: The number of nodes in the input graph is 3774768.                                        
NOTE: The number of links in the input graph is 16518948.                                       
NOTE: Processing connected components using 4 threads across 4 machines.                        
NOTE: The graph has 3627 connected components.                                                  
NOTE: Processing connected components used 0.62 (cpu: 1.54) seconds.                            
NOTE: The Cloud Analytic Services server processed the request in 2.285761 seconds.             
NOTE: The data set MYCAS.NODESETOUT has 3774768 observations and 2 variables.                   
NOTE: The data set MYCAS.CONCOMPOUT has 3627 observations and 2 variables.                      
STATUS=OK  PROBLEM_TYPE=CONNECTEDCOMPONENTS  SOLUTION_STATUS=OK  NUM_COMPONENTS=3627            
CPU_TIME=23.42  REAL_TIME=2.29                                                                  


The 10 biggest components are shown in Output 2.12.2. It is interesting to note that the vast majority of patents (over 99%) are all contained in the same component. This is not too surprising, because many of the seminal patent claims are required in order to understand subsequent inventions.

Output 2.12.2: Ten Largest Components for US Patent Citations

Obsconcompnodes
113764117
26919
329516
430115
54314
6322214
716314
8268014
919413
10128313


Last updated: June 11, 2021