MTLEARN Procedure

PROC MTLEARN Statement

  • PROC MTLEARN <options>;

The PROC MTLEARN statement invokes the procedure. Table 1 summarizes the options available in the PROC MTLEARN statement.

Table 1: PROC MTLEARN Statement Options

Option Description
Input Data Table Options
DATA= Specifies the input data table
GRAPHTABLE= Specifies the user-defined graph table
Multitask Learning Options
GRAPHTYPE= Specifies the type of graph table
MAXITER= Specifies the maximum number of iterations
NTHREADS= Specifies the number of threads to use on each computation node
REGL1= Specifies the script l 1 (LASSO) penalization weight
REGL2= Specifies the script l 2 graph penalization weight
SEED= Specifies the seed to use for pseudorandom number generation
TOLERANCE= Specifies the optimization tolerance (absolute script l 2 difference of solution) as a stopping criterion
Output Data Table Options
GRAPHOUT= Specifies the output data table in which to save the graph table
MODELOUT= Specifies the output data table in which to save the estimated multitask regression weights


You can specify the following options:

DATA=libref.data-table

names the input data table for PROC MTLEARN to use. The default is the most recently created data table. libref.data-table is a two-level name, where

libref

refers to a collection of information that is defined in the LIBNAME statement and includes the library, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about libref, see the section Using CAS Sessions and CAS Engine Librefs.

data-table

specifies the name of the input data table.

The data table must include one variable named id as the row ID variable.

GRAPHOUT=libref.data-table

specifies the output data table in which to save the graph table. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

GRAPHTABLE=libref.data-table

specifies the user-defined graph table. The first column of the graph table must contain the row ID variable. The graph table must also include all target variables that you specify in the TARGET statement. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

GRAPHTYPE=CLUSTER |CUSTOM |FUSE |INDEP

specifies the type of graph table.

You can specify one of the following values:

CLUSTER

generates a cluster graph table, where all tasks are connected (to a virtual mean task).

CUSTOM

uses the graph table that you specify.

FUSE

generates a fuse graph table, where each task is connected to the next task in the target list.

INDEP

generates an independent graph table, where all tasks are independent.

If you specify the GRAPHTABLE= option, the value of the GRAPHTYPE= option is automatically overwritten with CUSTOM.

By default, GRAPHTYPE=CLUSTER.

MAXITER=number

specifies the maximum number of iterations for the algorithm to perform, where number is an integer greater than or equal to 1.

By default, MAXITER=100. You can tune this value by using the AUTOTUNE statement.

MODELOUT=libref.data-table

specifies the output data table in which to save the estimated multitask regression weights. libref.data-table is a two-level name, where libref refers to the library, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

NTHREADS=number

specifies the number of threads to use for the computation, where number is an integer greater than 0. The default value is the maximum number of available threads per computer.

REGL1=number

specifies the script l 1 (LASSO) penalization weight, where number is greater than or equal to 0. The default value is 0.01. You can tune this value by using the AUTOTUNE statement.

REGL2=number

specifies the script l 2 graph penalization weight, where number is greater than or equal to 0. The default value is 0.01. You can tune this value by using the AUTOTUNE statement.

SEED=random-seed

specifies an integer that is used to start the pseudorandom number generator. This option enables you to reproduce the same sample output, but only when NTHREADS=1. If you do not specify a seed, or if you specify a value less than or equal to 0, the seed is generated from reading the time of day from the computer’s clock.

TOLERANCE=number

specifies the optimization tolerance (absolute script l 2 difference of solution) as a stopping criterion, where number is greater than 0. The default value is 1E–6. You can tune this value by using the AUTOTUNE statement.

Last updated: July 11, 2024