Text Mining Action Set

Example 22.3 Apply the SVD and Topic Discovery after Parsing

Apply the SVD and Topic Discovery after Parsing

This section contains PROC CAS code.

Note: Input data must be in a CAS table that is accessible in your CAS session. This table has a two-level name; the first level is your CAS engine libref and the second level is the table name. You refer to this table in the CAS procedure by specifying only the second level. For more information about two-level names, see Chapter 3: Shared Concepts in SAS Visual Data Mining and Machine Learning 8.3: Procedures Guide. For more information about PROC CAS and programming in CASL, see Getting Started with CASL, SAS Cloud Analytic Services: CAS Procedure Programming Guide and Reference, and SAS Viya: System Programming Guide.

Finally, the tmSvd action can be called separately if you decide to derive topics after you have initially parsed them. The following PROC CAS step calls the tmSvd action and uses the Parent and Terms table output from the tmMine action to discover the topics.

   proc cas;
       loadtable caslib="ReferenceData" path="en_stoplist.sashdat";
   run;
quit;
 proc cas;
    loadactionset "textMining";
    action tmSvd;
    param
      parent={ name="parent"}
      terms={name="terms"}
      k=3
      u ={ name="svdu", replace=TRUE}
      numLabels=3
      topics={name="topicsSVD",replace=TRUE}
 ;
 action table.fetch /table="topicsSVD"; run;
 run;
 quit;

The singular value decomposition and topic discovery can be computed alone after parsing has been done. Output 22.3.1 displays the contents of the mycas.topicsvd table, which contain the results of the discovered topics.

Output 22.3.1: Discovered Topics

Results from table.fetch

Selected Rows from Table TOPICSSVD
_Index_Topic IDTopicTerm Cutoff
11book, plot, read0.381
22movie, +bore, +watch0.385
33tv, phone, resolution0.385


Apply the SVD and Topic Discovery after Parsing

This section contains Lua code.

- Load action sets

-- Upload data
s:upload{'reviews.csv', casout={name='reviews'}}
s.loadtable(caslib="ReferenceData",path="en_stoplist.sashdat")

-- Load action sets
s:loadactionset{actionset='textmining'}

-- Discover topics and Doc Projections
s:tmMine{
	         docid='did', 
           docpro={name='docpro',replace=True},
           documents='reviews', 
           k=3,
           nounGroups=false, 
           numLabels=3,
           offset={name='offset',replace=true}, 
           parent={name='parent',replace=true}, 
           parseConfig={name='config',replace=true}, 
           reduce=2,
           stopList='en_stoplist',
           tagging=true, 
           terms={name='terms',replace=true}, 
           text='text',  
           topicDecision=true,
           topics={name='topics',replace=True},
           u={name='svdu',replace=True},
           
                      
}

  -- Finally, if you did not calculate the SVD or topics initially,
  -- you can do it with the parent and term tables as input.
  
s:tmSvd{
        k=3,
        numLabels =3,
        parent='parent',
        terms='terms',
        topics={name='topicsSVD',replace=True},
        u={name='svduS',replace=True}
}
  -- topic table
  r=s.fetch{table='topicsSVD'}
  print(r.Fetch)

Apply the SVD and Topic Discovery after Parsing

This section contains Python code.

   import swat

# Create training data
from io import StringIO                
reviews = StringIO('''text,positive,category,did                   
"This is the greatest phone ever! love it! It can replace my tv!",1,electronics,1
"The phone's battery life is too short and screen resolution is low.",0,electronics,2
"The screen resolution is low, but I love this tv.  Good viewing.",1,electronics,3
"The movie itself is great and I liked watching it. Good acting!",1,movies,4
"The movie's story is boring and the acting is poor.",0,movies,5
"I watched this movie but it was boring..",0,movies,6
"The book has a terrific plot!",1,books,7
"The book's plot was suspenseful. Good read.",1,books,8
"I love the author, but this book is a waste of time to read.",0,books,9''')


handler = dmh.CSV(reviews, skipinitialspace=True)
s.addtable(table='reviews', **handler.args.addtable)
s.loadtable(caslib="ReferenceData",path="en_stoplist.sashdat")

# Discover topics and Doc Projections
s.loadactionset(actionset='textmining')
s.tmMine(docId="did",                                           # 2
        docPro={"name":"docpro", "replace":True},
        documents={"name":"reviews"},
        k=3,
        nounGroups=False,
        numLabels=3,
        offset={"name":"offset", "replace":True},
        parent={"name":"parent", "replace":True},
        parseConfig={"name":"config", "replace":True},
        reduce=2,
        stopList={"name":"en_stopList"},
        tagging=True,
        terms={"name":"terms", "replace":True},
        text="text",
        topicDecision=True,
        topics={"name":"topics", "replace":True},
        u={"name":"svdu", "replace":True}
        )

# If you didn't calculate the SVD or topics initially, you can
# do it with the parent and term tables as input.
s.tmSvd(k=3,                              
        numLabels=3,
        parent={"name":"parent"},
        terms={"name":"terms"},
        topics={"name":"topicsSVD", "replace":True},
        u={"name":"svdu", "replace":True}                   
        )
# Topic table
pprint(s.fetch('topicsSVD'))

Apply the SVD and Topic Discovery after Parsing

This example is not available for the R programming language.

Last updated: June 07, 2018