Reinforcement Learning Action Set: Syntax

Provides actions for building, training, and scoring reinforcement learning models

rlTrainPolicyGradient Action

Do reinforcement learning training using Policy Gradient algorithm.

AliasrlTrainPG
reinforcementLearn.rlTrainPolicyGradient <result=results> <status=rc> /
required parameter actorModel
={{
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
}, {...}},
actorOptimizer={method="ADAM" | "SGD", method-specific-parameters},
checkPointBestOut
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
checkPointIn
={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | greenplum-parameters | hadoop-parameters | hana-parameters | hdfs-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | netezza-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
checkPointOut
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
criticModel
={{
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
}, {...}},
criticOptimizer={method="ADAM" | "SGD", method-specific-parameters},
required parameter environment={type="BUILTIN" | "REMOTE", type-specific-parameters},
gamma=double,
gpu
={
deterministic=TRUE | FALSE,
device=integer,
enable=TRUE | FALSE
},
required parameter modelOut
={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
nThreads=integer,
numEpisodes=integer,
required parameter pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters},
seed=double,
testInterval=integer
;
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

 checkPointIn

—

Specifies a table that includes a saved model for warm starting.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 checkPointBestOut

—

Specifies a table that to save the best observed model.

 checkPointOut

—

Specifies a table that is used to save the model.

required parametermodelOut

—

Specifies a table that is used to save model structure and weights.

Parameter Descriptions

* actorModel={{nnParams-1} <, {nnParams-2}, ...>}

Specifies the properties of the layers of the neural net used to build the actor network.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

actorOptimizer={method="ADAM" | "SGD", method-specific-parameters}

The optimization method for the actor model: SGD or ADAM.

AliasactorOptimization

The value that you specify for the method parameter determines the other parameters that apply.

checkPointBestOut={casouttable}

Specifies a table that to save the best observed model.

For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.

checkPointFreq=integer

Specifies how often to save training parameters.

Range1–100000000

checkPointIn={castable}

Specifies a table that includes a saved model for warm starting.

For more information about specifying the checkPointIn parameter, see the common castable parameter.

checkPointOut={casouttable}

Specifies a table that is used to save the model.

For more information about specifying the checkPointOut parameter, see the common casouttable parameter.

criticModel={{nnParams-1} <, {nnParams-2}, ...>}

Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

criticOptimizer={method="ADAM" | "SGD", method-specific-parameters}

The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.

AliascriticOptimization

The value that you specify for the method parameter determines the other parameters that apply.

* environment={type="BUILTIN" | "REMOTE", type-specific-parameters}

The reinforcement learning environment.

Aliasenv

The value that you specify for the type parameter determines the other parameters that apply.

gamma=double

Discount factor used to compute long term value.

Default0.99
Range0–1

gpu={gpuOptions}

Specifies GPU related options.

The gpuOptions value can be one or more of the following:

deterministic=TRUE | FALSE

Specifies that the action will use deterministic GPU functions (may be slower).

DefaultFALSE
device=integer

Specifies a specific GPU device id.

Minimum value0
enable=TRUE | FALSE

Specifies that action will use a GPU to perform model calculations.

DefaultFALSE

maxNormOfGradient=double

Adjust the magnitude of the gradient so that it is less than this value.

Range1–100000000

* modelOut={casouttable}

Specifies a table that is used to save model structure and weights.

For more information about specifying the modelOut parameter, see the common casouttable parameter.

nThreads=integer

Number of CPU threads to use.

Range0–1024

numEpisodes=integer

Number of episodes to train before terminating.

Default10
Range1–100000000

numTestEpisodes=integer

Specifies number of episodes to run when running test.

Default100
Range1–100000000

* pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}

The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C

AliaspgAlg

The value that you specify for the type parameter determines the other parameters that apply.

seed=double

Seed for random number generator.

Minimum value0

testInterval=integer

Specifies how often to run test.

Default100
Range1–100000000

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=TRUE | FALSE

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFALSE

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=TRUE | FALSE

Enables Nesterov momentum.

DefaultFALSE

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=TRUE | FALSE

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFALSE

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=TRUE | FALSE

Enables Nesterov momentum.

DefaultFALSE

Parameters for type="BUILTIN"

* name="string"

Parameters for type="REMOTE"

imageStack=integer

Specifies number of consecutive observations to be stacked as an input observation.

Default1
Range1–10000
* name="string"

The environment name used in RL.

render=TRUE | FALSE

Whether render the environment display or not.

DefaultFALSE
renderFreq=integer

Specifies how frequently the environment is rendered.

AliasrenderFrequency
Default1
Range1–10000000
renderSleep=double

Time to sleep (in seconds) between rendering frames.

Default0
seed=integer

Seed of the environment.

Default0
* url="string"

The url:port used for in RL training.

Default"127.0.0.1:8080"
wrapper={"string-1" <, "string-2", ...>}

A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.

Parameters for type="ACTORCRITIC"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
GAELambda=double

Specifies the value of the lambda in generalized advantage estimation (GAE).

Default0.9
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024
useGAELoss=TRUE | FALSE

Use GAE Loss.

DefaultFALSE

Parameters for type="REINFORCE"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

Parameters for type="REINFORCERTG"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

rlTrainPolicyGradient Action

Do reinforcement learning training using Policy Gradient algorithm.

AliasrlTrainPG
results, info = s:reinforcementLearn_rlTrainPolicyGradient{
required parameter actorModel
={{
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
}, {...}},
actorOptimizer={method="ADAM" | "SGD", method-specific-parameters},
checkPointBestOut
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
checkPointIn
={
caslib="string",
computedOnDemand=true | false,
computedVars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
computedVarsProgram="string",
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression",
whereTable
={
casLib="string"
dataSourceOptions={adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | greenplum-parameters | hadoop-parameters | hana-parameters | hdfs-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | netezza-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
importOptions={fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter name="table-name"
vars
={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}}
where="where-expression"
}
},
checkPointOut
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
criticModel
={{
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
}, {...}},
criticOptimizer={method="ADAM" | "SGD", method-specific-parameters},
required parameter environment={type="BUILTIN" | "REMOTE", type-specific-parameters},
gamma=double,
gpu
={
deterministic=true | false,
device=integer,
enable=true | false
},
required parameter modelOut
={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
nThreads=integer,
numEpisodes=integer,
required parameter pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters},
seed=double,
testInterval=integer
}
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

 checkPointIn

—

Specifies a table that includes a saved model for warm starting.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 checkPointBestOut

—

Specifies a table that to save the best observed model.

 checkPointOut

—

Specifies a table that is used to save the model.

required parametermodelOut

—

Specifies a table that is used to save model structure and weights.

Parameter Descriptions

* actorModel={{nnParams-1} <, {nnParams-2}, ...>}

Specifies the properties of the layers of the neural net used to build the actor network.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

actorOptimizer={method="ADAM" | "SGD", method-specific-parameters}

The optimization method for the actor model: SGD or ADAM.

AliasactorOptimization

The value that you specify for the method parameter determines the other parameters that apply.

checkPointBestOut={casouttable}

Specifies a table that to save the best observed model.

For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.

checkPointFreq=integer

Specifies how often to save training parameters.

Range1–100000000

checkPointIn={castable}

Specifies a table that includes a saved model for warm starting.

For more information about specifying the checkPointIn parameter, see the common castable parameter.

checkPointOut={casouttable}

Specifies a table that is used to save the model.

For more information about specifying the checkPointOut parameter, see the common casouttable parameter.

criticModel={{nnParams-1} <, {nnParams-2}, ...>}

Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

criticOptimizer={method="ADAM" | "SGD", method-specific-parameters}

The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.

AliascriticOptimization

The value that you specify for the method parameter determines the other parameters that apply.

* environment={type="BUILTIN" | "REMOTE", type-specific-parameters}

The reinforcement learning environment.

Aliasenv

The value that you specify for the type parameter determines the other parameters that apply.

gamma=double

Discount factor used to compute long term value.

Default0.99
Range0–1

gpu={gpuOptions}

Specifies GPU related options.

The gpuOptions value can be one or more of the following:

deterministic=true | false

Specifies that the action will use deterministic GPU functions (may be slower).

Defaultfalse
device=integer

Specifies a specific GPU device id.

Minimum value0
enable=true | false

Specifies that action will use a GPU to perform model calculations.

Defaultfalse

maxNormOfGradient=double

Adjust the magnitude of the gradient so that it is less than this value.

Range1–100000000

* modelOut={casouttable}

Specifies a table that is used to save model structure and weights.

For more information about specifying the modelOut parameter, see the common casouttable parameter.

nThreads=integer

Number of CPU threads to use.

Range0–1024

numEpisodes=integer

Number of episodes to train before terminating.

Default10
Range1–100000000

numTestEpisodes=integer

Specifies number of episodes to run when running test.

Default100
Range1–100000000

* pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}

The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C

AliaspgAlg

The value that you specify for the type parameter determines the other parameters that apply.

seed=double

Seed for random number generator.

Minimum value0

testInterval=integer

Specifies how often to run test.

Default100
Range1–100000000

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=true | false

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

Defaultfalse

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=true | false

Enables Nesterov momentum.

Defaultfalse

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=true | false

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

Defaultfalse

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=true | false

Enables Nesterov momentum.

Defaultfalse

Parameters for type="BUILTIN"

* name="string"

Parameters for type="REMOTE"

imageStack=integer

Specifies number of consecutive observations to be stacked as an input observation.

Default1
Range1–10000
* name="string"

The environment name used in RL.

render=true | false

Whether render the environment display or not.

Defaultfalse
renderFreq=integer

Specifies how frequently the environment is rendered.

AliasrenderFrequency
Default1
Range1–10000000
renderSleep=double

Time to sleep (in seconds) between rendering frames.

Default0
seed=integer

Seed of the environment.

Default0
* url="string"

The url:port used for in RL training.

Default"127.0.0.1:8080"
wrapper={"string-1" <, "string-2", ...>}

A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.

Parameters for type="ACTORCRITIC"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
GAELambda=double

Specifies the value of the lambda in generalized advantage estimation (GAE).

Default0.9
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024
useGAELoss=true | false

Use GAE Loss.

Defaultfalse

Parameters for type="REINFORCE"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

Parameters for type="REINFORCERTG"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

rlTrainPolicyGradient Action

Do reinforcement learning training using Policy Gradient algorithm.

AliasrlTrainPG
results=s.reinforcementLearn.rlTrainPolicyGradient(
required parameter actorModel
=[{
"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH",
"n":integer,
"size":integer,
"stride":integer,
"type":"CONV" | "FC"
}<, {...}>],
actorOptimizer={"method":"ADAM" | "SGD", method-specific-parameters},
checkPointBestOut
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
checkPointIn
={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"groupByMode":"NOSORT" | "REDISTRIBUTE",
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression",
"whereTable"
:{
"casLib":"string"
"dataSourceOptions":{adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | greenplum-parameters | hadoop-parameters | hana-parameters | hdfs-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | netezza-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters}
"importOptions":{"fileType":"ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters}
required parameter "name":"table-name"
"vars"
:[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>]
"where":"where-expression"
}
},
checkPointOut
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
criticModel
=[{
"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH",
"n":integer,
"size":integer,
"stride":integer,
"type":"CONV" | "FC"
}<, {...}>],
criticOptimizer={"method":"ADAM" | "SGD", method-specific-parameters},
required parameter environment={"type":"BUILTIN" | "REMOTE", type-specific-parameters},
gamma=double,
gpu
={
"deterministic":True | False,
"device":integer,
"enable":True | False
},
required parameter modelOut
={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"lifetime":64-bit-integer,
"maxMemSize":64-bit-integer,
"memoryFormat":"DVR" | "INHERIT" | "STANDARD",
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"threadBlockSize":64-bit-integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
nThreads=integer,
numEpisodes=integer,
required parameter pgAlgorithm={"type":"ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters},
seed=double,
testInterval=integer
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

 checkPointIn

—

Specifies a table that includes a saved model for warm starting.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 checkPointBestOut

—

Specifies a table that to save the best observed model.

 checkPointOut

—

Specifies a table that is used to save the model.

required parametermodelOut

—

Specifies a table that is used to save model structure and weights.

Parameter Descriptions

* actorModel=[{nnParams-1} <, {nnParams-2}, ...>]

Specifies the properties of the layers of the neural net used to build the actor network.

The nnParams value can be one or more of the following:

"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
"n":integer

Number of neurons.

Range1–100000
"size":integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
"stride":integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
"type":"CONV" | "FC"

Layer type.

DefaultFC

actorOptimizer={"method":"ADAM" | "SGD", method-specific-parameters}

The optimization method for the actor model: SGD or ADAM.

AliasactorOptimization

The value that you specify for the method parameter determines the other parameters that apply.

checkPointBestOut={casouttable}

Specifies a table that to save the best observed model.

For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.

checkPointFreq=integer

Specifies how often to save training parameters.

Range1–100000000

checkPointIn={castable}

Specifies a table that includes a saved model for warm starting.

For more information about specifying the checkPointIn parameter, see the common castable parameter.

checkPointOut={casouttable}

Specifies a table that is used to save the model.

For more information about specifying the checkPointOut parameter, see the common casouttable parameter.

criticModel=[{nnParams-1} <, {nnParams-2}, ...>]

Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.

The nnParams value can be one or more of the following:

"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
"n":integer

Number of neurons.

Range1–100000
"size":integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
"stride":integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
"type":"CONV" | "FC"

Layer type.

DefaultFC

criticOptimizer={"method":"ADAM" | "SGD", method-specific-parameters}

The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.

AliascriticOptimization

The value that you specify for the method parameter determines the other parameters that apply.

* environment={"type":"BUILTIN" | "REMOTE", type-specific-parameters}

The reinforcement learning environment.

Aliasenv

The value that you specify for the type parameter determines the other parameters that apply.

gamma=double

Discount factor used to compute long term value.

Default0.99
Range0–1

gpu={gpuOptions}

Specifies GPU related options.

The gpuOptions value can be one or more of the following:

"deterministic":True | False

Specifies that the action will use deterministic GPU functions (may be slower).

DefaultFalse
"device":integer

Specifies a specific GPU device id.

Minimum value0
"enable":True | False

Specifies that action will use a GPU to perform model calculations.

DefaultFalse

maxNormOfGradient=double

Adjust the magnitude of the gradient so that it is less than this value.

Range1–100000000

* modelOut={casouttable}

Specifies a table that is used to save model structure and weights.

For more information about specifying the modelOut parameter, see the common casouttable parameter.

nThreads=integer

Number of CPU threads to use.

Range0–1024

numEpisodes=integer

Number of episodes to train before terminating.

Default10
Range1–100000000

numTestEpisodes=integer

Specifies number of episodes to run when running test.

Default100
Range1–100000000

* pgAlgorithm={"type":"ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}

The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C

AliaspgAlg

The value that you specify for the type parameter determines the other parameters that apply.

seed=double

Seed for random number generator.

Minimum value0

testInterval=integer

Specifies how often to run test.

Default100
Range1–100000000

Parameters for method="ADAM"

"beta1":double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
"beta2":double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
"epsilon":double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
"learningRate":double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
"miniBatchSize":integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
"regL2":double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
"useAMSGrad":True | False

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFalse

Parameters for method="SGD"

"damping":double

Damping for momentum.

Default0
Range0–1
"learningRate":double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
"miniBatchSize":integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
"momentum":double

Specifies the value for momentum for SGD.

Default0
Range0–1
"regL2":double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
"useNesterov":True | False

Enables Nesterov momentum.

DefaultFalse

Parameters for method="ADAM"

"beta1":double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
"beta2":double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
"epsilon":double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
"learningRate":double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
"miniBatchSize":integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
"regL2":double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
"useAMSGrad":True | False

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFalse

Parameters for method="SGD"

"damping":double

Damping for momentum.

Default0
Range0–1
"learningRate":double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
"miniBatchSize":integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
"momentum":double

Specifies the value for momentum for SGD.

Default0
Range0–1
"regL2":double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
"useNesterov":True | False

Enables Nesterov momentum.

DefaultFalse

Parameters for type="BUILTIN"

* "name":"string"

Parameters for type="REMOTE"

"imageStack":integer

Specifies number of consecutive observations to be stacked as an input observation.

Default1
Range1–10000
* "name":"string"

The environment name used in RL.

"render":True | False

Whether render the environment display or not.

DefaultFalse
"renderFreq":integer

Specifies how frequently the environment is rendered.

AliasrenderFrequency
Default1
Range1–10000000
"renderSleep":double

Time to sleep (in seconds) between rendering frames.

Default0
"seed":integer

Seed of the environment.

Default0
* "url":"string"

The url:port used for in RL training.

Default"127.0.0.1:8080"
"wrapper":["string-1" <, "string-2", ...>]

A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.

Parameters for type="ACTORCRITIC"

"entropyCoef":double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
"GAELambda":double

Specifies the value of the lambda in generalized advantage estimation (GAE).

Default0.9
Range0–1
"nRollouts":integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024
"useGAELoss":True | False

Use GAE Loss.

DefaultFalse

Parameters for type="REINFORCE"

"entropyCoef":double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
"nRollouts":integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

Parameters for type="REINFORCERTG"

"entropyCoef":double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
"nRollouts":integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

rlTrainPolicyGradient Action

Do reinforcement learning training using Policy Gradient algorithm.

AliasrlTrainPG
results <– cas.reinforcementLearn.rlTrainPolicyGradient(s,
required parameter actorModel
=list( list(
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
) <, list(...)>),
actorOptimizer=list(method="ADAM" | "SGD", method-specific-parameters),
checkPointBestOut
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
checkPointIn
=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
computedVarsProgram="string",
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
groupByMode="NOSORT" | "REDISTRIBUTE",
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression",
whereTable
=list(
casLib="string"
dataSourceOptions=list(adls_noreq-parameters | bigquery-parameters | cas_noreq-parameters | clouddex-parameters | db2-parameters | dnfs-parameters | esp-parameters | fedsvr-parameters | gcs_noreq-parameters | greenplum-parameters | hadoop-parameters | hana-parameters | hdfs-parameters | impala-parameters | jdbc-parameters | mongodb-parameters | mysql-parameters | netezza-parameters | odbc-parameters | oracle-parameters | path-parameters | postgres-parameters | redshift-parameters | s3-parameters | sapiq-parameters | sforce-parameters | singlestore_standard-parameters | snowflake-parameters | spark-parameters | spde-parameters | sqlserver-parameters | ss_noreq-parameters | teradata-parameters | vertica-parameters | yellowbrick-parameters)
importOptions=list(fileType="ANY" | "AUDIO" | "AUTO" | "BASESAS" | "CSV" | "DELIMITED" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "IMAGE" | "JMP" | "LASR" | "PARQUET" | "SOUND" | "SPSS" | "VIDEO" | "XLS", fileType-specific-parameters)
required parameter name="table-name"
vars
=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>)
where="where-expression"
)
),
checkPointOut
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
criticModel
=list( list(
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH",
n=integer,
size=integer,
stride=integer,
type="CONV" | "FC"
) <, list(...)>),
criticOptimizer=list(method="ADAM" | "SGD", method-specific-parameters),
required parameter environment=list(type="BUILTIN" | "REMOTE", type-specific-parameters),
gamma=double,
gpu
=list(
deterministic=TRUE | FALSE,
device=integer,
enable=TRUE | FALSE
),
required parameter modelOut
=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
lifetime=64-bit-integer,
maxMemSize=64-bit-integer,
memoryFormat="DVR" | "INHERIT" | "STANDARD",
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
threadBlockSize=64-bit-integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
nThreads=integer,
numEpisodes=integer,
required parameter pgAlgorithm=list(type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters),
seed=double,
testInterval=integer
)
indicates a required parameter

Summary: Input and Output Tables

If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.

Parameters for Reading Input Tables

Parameter

Subparameter

Description

 checkPointIn

—

Specifies a table that includes a saved model for warm starting.

Parameters for Creating Output Tables

Parameter

Subparameter

Description

 checkPointBestOut

—

Specifies a table that to save the best observed model.

 checkPointOut

—

Specifies a table that is used to save the model.

required parametermodelOut

—

Specifies a table that is used to save model structure and weights.

Parameter Descriptions

* actorModel=list( list(nnParams-1) <, list(nnParams-2), ...>)

Specifies the properties of the layers of the neural net used to build the actor network.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

actorOptimizer=list(method="ADAM" | "SGD", method-specific-parameters)

The optimization method for the actor model: SGD or ADAM.

AliasactorOptimization

The value that you specify for the method parameter determines the other parameters that apply.

checkPointBestOut=list(casouttable)

Specifies a table that to save the best observed model.

For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.

checkPointFreq=integer

Specifies how often to save training parameters.

Range1–100000000

checkPointIn=list(castable)

Specifies a table that includes a saved model for warm starting.

For more information about specifying the checkPointIn parameter, see the common castable parameter.

checkPointOut=list(casouttable)

Specifies a table that is used to save the model.

For more information about specifying the checkPointOut parameter, see the common casouttable parameter.

criticModel=list( list(nnParams-1) <, list(nnParams-2), ...>)

Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.

The nnParams value can be one or more of the following:

act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"

Activation function.

DefaultRELU
n=integer

Number of neurons.

Range1–100000
size=integer

Size (width and height) of kernel (for convolution layer) or window (for pooling layer).

Range1–1000
stride=integer

Stride of kernel (for convolution layer) or window (for pooling layer).

Default1
Range1–10000
type="CONV" | "FC"

Layer type.

DefaultFC

criticOptimizer=list(method="ADAM" | "SGD", method-specific-parameters)

The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.

AliascriticOptimization

The value that you specify for the method parameter determines the other parameters that apply.

* environment=list(type="BUILTIN" | "REMOTE", type-specific-parameters)

The reinforcement learning environment.

Aliasenv

The value that you specify for the type parameter determines the other parameters that apply.

gamma=double

Discount factor used to compute long term value.

Default0.99
Range0–1

gpu=list(gpuOptions)

Specifies GPU related options.

The gpuOptions value can be one or more of the following:

deterministic=TRUE | FALSE

Specifies that the action will use deterministic GPU functions (may be slower).

DefaultFALSE
device=integer

Specifies a specific GPU device id.

Minimum value0
enable=TRUE | FALSE

Specifies that action will use a GPU to perform model calculations.

DefaultFALSE

maxNormOfGradient=double

Adjust the magnitude of the gradient so that it is less than this value.

Range1–100000000

* modelOut=list(casouttable)

Specifies a table that is used to save model structure and weights.

For more information about specifying the modelOut parameter, see the common casouttable parameter.

nThreads=integer

Number of CPU threads to use.

Range0–1024

numEpisodes=integer

Number of episodes to train before terminating.

Default10
Range1–100000000

numTestEpisodes=integer

Specifies number of episodes to run when running test.

Default100
Range1–100000000

* pgAlgorithm=list(type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters)

The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C

AliaspgAlg

The value that you specify for the type parameter determines the other parameters that apply.

seed=double

Seed for random number generator.

Minimum value0

testInterval=integer

Specifies how often to run test.

Default100
Range1–100000000

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=TRUE | FALSE

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFALSE

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=TRUE | FALSE

Enables Nesterov momentum.

DefaultFALSE

Parameters for method="ADAM"

beta1=double

The exponential decay rate for the first moment estimates.

Default0.9
Range0–1
beta2=double

The exponential decay rate for the second moment estimates.

Default0.999
Range0–1
epsilon=double

The value to add into the second moment of the gradient in the denominator to improve numerical stability.

Aliaseps
Default1E-08
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useAMSGrad=TRUE | FALSE

Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.

DefaultFALSE

Parameters for method="SGD"

damping=double

Damping for momentum.

Default0
Range0–1
learningRate=double

Specifies the learning rate parameter for SGD.

Aliaslr
Default0.001
Range0–1
miniBatchSize=integer

Specifies minibatch size to be used when training the Q model.

Default64
Range1–100000
momentum=double

Specifies the value for momentum for SGD.

Default0
Range0–1
regL2=double

Specifies the L2 regularization parameter, the value must be nonnegative.

AliasweightDecay
Default0
Range0–1
useNesterov=TRUE | FALSE

Enables Nesterov momentum.

DefaultFALSE

Parameters for type="BUILTIN"

* name="string"

Parameters for type="REMOTE"

imageStack=integer

Specifies number of consecutive observations to be stacked as an input observation.

Default1
Range1–10000
* name="string"

The environment name used in RL.

render=TRUE | FALSE

Whether render the environment display or not.

DefaultFALSE
renderFreq=integer

Specifies how frequently the environment is rendered.

AliasrenderFrequency
Default1
Range1–10000000
renderSleep=double

Time to sleep (in seconds) between rendering frames.

Default0
seed=integer

Seed of the environment.

Default0
* url="string"

The url:port used for in RL training.

Default"127.0.0.1:8080"
wrapper=list("string-1" <, "string-2", ...>)

A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.

Parameters for type="ACTORCRITIC"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
GAELambda=double

Specifies the value of the lambda in generalized advantage estimation (GAE).

Default0.9
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024
useGAELoss=TRUE | FALSE

Use GAE Loss.

DefaultFALSE

Parameters for type="REINFORCE"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024

Parameters for type="REINFORCERTG"

entropyCoef=double

Specifies the value for the entropy coefficient in policy gradient algorithms.

Default0.01
Range0–1
nRollouts=integer

Specifies the number of rollouts to generate data.

Default1
Range1–1024
Last updated: February 8, 2023