Reinforcement Learning Action Set: Syntax
Provides actions for building, training, and scoring reinforcement learning models
rlTrainPolicyGradient Action
Do reinforcement learning training using Policy Gradient algorithm.
| Alias | rlTrainPG |
|---|
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that includes a saved model for warm starting. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that to save the best observed model. | |
|
— |
Specifies a table that is used to save the model. | |
|
required parametermodelOut |
— |
Specifies a table that is used to save model structure and weights. |
Parameter Descriptions
* actorModel={{nnParams-1} <, {nnParams-2}, ...>}
Specifies the properties of the layers of the neural net used to build the actor network.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
actorOptimizer={method="ADAM" | "SGD", method-specific-parameters}
The optimization method for the actor model: SGD or ADAM.
| Alias | actorOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
checkPointBestOut={casouttable}
Specifies a table that to save the best observed model.
For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.
checkPointFreq=integer
Specifies how often to save training parameters.
| Range | 1–100000000 |
|---|
checkPointIn={castable}
Specifies a table that includes a saved model for warm starting.
For more information about specifying the checkPointIn parameter, see the common castable parameter.
checkPointOut={casouttable}
Specifies a table that is used to save the model.
For more information about specifying the checkPointOut parameter, see the common casouttable parameter.
criticModel={{nnParams-1} <, {nnParams-2}, ...>}
Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
criticOptimizer={method="ADAM" | "SGD", method-specific-parameters}
The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.
| Alias | criticOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
* environment={type="BUILTIN" | "REMOTE", type-specific-parameters}
The reinforcement learning environment.
| Alias | env |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
gamma=double
Discount factor used to compute long term value.
| Default | 0.99 |
|---|---|
| Range | 0–1 |
gpu={gpuOptions}
Specifies GPU related options.
The gpuOptions value can be one or more of the following:
deterministic=TRUE | FALSE
Specifies that the action will use deterministic GPU functions (may be slower).
| Default | FALSE |
|---|
device=integer
Specifies a specific GPU device id.
| Minimum value | 0 |
|---|
enable=TRUE | FALSE
Specifies that action will use a GPU to perform model calculations.
| Default | FALSE |
|---|
maxNormOfGradient=double
Adjust the magnitude of the gradient so that it is less than this value.
| Range | 1–100000000 |
|---|
* modelOut={casouttable}
Specifies a table that is used to save model structure and weights.
For more information about specifying the modelOut parameter, see the common casouttable parameter.
nThreads=integer
Number of CPU threads to use.
| Range | 0–1024 |
|---|
numEpisodes=integer
Number of episodes to train before terminating.
| Default | 10 |
|---|---|
| Range | 1–100000000 |
numTestEpisodes=integer
Specifies number of episodes to run when running test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
* pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}
The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C
| Alias | pgAlg |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
seed=double
Seed for random number generator.
| Minimum value | 0 |
|---|
testInterval=integer
Specifies how often to run test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=TRUE | FALSE
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | FALSE |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=TRUE | FALSE
Enables Nesterov momentum.
| Default | FALSE |
|---|
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=TRUE | FALSE
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | FALSE |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=TRUE | FALSE
Enables Nesterov momentum.
| Default | FALSE |
|---|
Parameters for type="BUILTIN"
* name="string"
Parameters for type="REMOTE"
imageStack=integer
Specifies number of consecutive observations to be stacked as an input observation.
| Default | 1 |
|---|---|
| Range | 1–10000 |
* name="string"
The environment name used in RL.
render=TRUE | FALSE
Whether render the environment display or not.
| Default | FALSE |
|---|
renderFreq=integer
Specifies how frequently the environment is rendered.
| Alias | renderFrequency |
|---|---|
| Default | 1 |
| Range | 1–10000000 |
renderSleep=double
Time to sleep (in seconds) between rendering frames.
| Default | 0 |
|---|
seed=integer
Seed of the environment.
| Default | 0 |
|---|
* url="string"
The url:port used for in RL training.
| Default | "127.0.0.1:8080" |
|---|
wrapper={"string-1" <, "string-2", ...>}
A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.
Parameters for type="ACTORCRITIC"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
GAELambda=double
Specifies the value of the lambda in generalized advantage estimation (GAE).
| Default | 0.9 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
useGAELoss=TRUE | FALSE
Use GAE Loss.
| Default | FALSE |
|---|
Parameters for type="REINFORCE"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
Parameters for type="REINFORCERTG"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
rlTrainPolicyGradient Action
Do reinforcement learning training using Policy Gradient algorithm.
| Alias | rlTrainPG |
|---|
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that includes a saved model for warm starting. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that to save the best observed model. | |
|
— |
Specifies a table that is used to save the model. | |
|
required parametermodelOut |
— |
Specifies a table that is used to save model structure and weights. |
Parameter Descriptions
* actorModel={{nnParams-1} <, {nnParams-2}, ...>}
Specifies the properties of the layers of the neural net used to build the actor network.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
actorOptimizer={method="ADAM" | "SGD", method-specific-parameters}
The optimization method for the actor model: SGD or ADAM.
| Alias | actorOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
checkPointBestOut={casouttable}
Specifies a table that to save the best observed model.
For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.
checkPointFreq=integer
Specifies how often to save training parameters.
| Range | 1–100000000 |
|---|
checkPointIn={castable}
Specifies a table that includes a saved model for warm starting.
For more information about specifying the checkPointIn parameter, see the common castable parameter.
checkPointOut={casouttable}
Specifies a table that is used to save the model.
For more information about specifying the checkPointOut parameter, see the common casouttable parameter.
criticModel={{nnParams-1} <, {nnParams-2}, ...>}
Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
criticOptimizer={method="ADAM" | "SGD", method-specific-parameters}
The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.
| Alias | criticOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
* environment={type="BUILTIN" | "REMOTE", type-specific-parameters}
The reinforcement learning environment.
| Alias | env |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
gamma=double
Discount factor used to compute long term value.
| Default | 0.99 |
|---|---|
| Range | 0–1 |
gpu={gpuOptions}
Specifies GPU related options.
The gpuOptions value can be one or more of the following:
deterministic=true | false
Specifies that the action will use deterministic GPU functions (may be slower).
| Default | false |
|---|
device=integer
Specifies a specific GPU device id.
| Minimum value | 0 |
|---|
enable=true | false
Specifies that action will use a GPU to perform model calculations.
| Default | false |
|---|
maxNormOfGradient=double
Adjust the magnitude of the gradient so that it is less than this value.
| Range | 1–100000000 |
|---|
* modelOut={casouttable}
Specifies a table that is used to save model structure and weights.
For more information about specifying the modelOut parameter, see the common casouttable parameter.
nThreads=integer
Number of CPU threads to use.
| Range | 0–1024 |
|---|
numEpisodes=integer
Number of episodes to train before terminating.
| Default | 10 |
|---|---|
| Range | 1–100000000 |
numTestEpisodes=integer
Specifies number of episodes to run when running test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
* pgAlgorithm={type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}
The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C
| Alias | pgAlg |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
seed=double
Seed for random number generator.
| Minimum value | 0 |
|---|
testInterval=integer
Specifies how often to run test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=true | false
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | false |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=true | false
Enables Nesterov momentum.
| Default | false |
|---|
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=true | false
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | false |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=true | false
Enables Nesterov momentum.
| Default | false |
|---|
Parameters for type="BUILTIN"
* name="string"
Parameters for type="REMOTE"
imageStack=integer
Specifies number of consecutive observations to be stacked as an input observation.
| Default | 1 |
|---|---|
| Range | 1–10000 |
* name="string"
The environment name used in RL.
render=true | false
Whether render the environment display or not.
| Default | false |
|---|
renderFreq=integer
Specifies how frequently the environment is rendered.
| Alias | renderFrequency |
|---|---|
| Default | 1 |
| Range | 1–10000000 |
renderSleep=double
Time to sleep (in seconds) between rendering frames.
| Default | 0 |
|---|
seed=integer
Seed of the environment.
| Default | 0 |
|---|
* url="string"
The url:port used for in RL training.
| Default | "127.0.0.1:8080" |
|---|
wrapper={"string-1" <, "string-2", ...>}
A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.
Parameters for type="ACTORCRITIC"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
GAELambda=double
Specifies the value of the lambda in generalized advantage estimation (GAE).
| Default | 0.9 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
useGAELoss=true | false
Use GAE Loss.
| Default | false |
|---|
Parameters for type="REINFORCE"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
Parameters for type="REINFORCERTG"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
rlTrainPolicyGradient Action
Do reinforcement learning training using Policy Gradient algorithm.
| Alias | rlTrainPG |
|---|
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that includes a saved model for warm starting. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that to save the best observed model. | |
|
— |
Specifies a table that is used to save the model. | |
|
required parametermodelOut |
— |
Specifies a table that is used to save model structure and weights. |
Parameter Descriptions
* actorModel=[{nnParams-1} <, {nnParams-2}, ...>]
Specifies the properties of the layers of the neural net used to build the actor network.
The nnParams value can be one or more of the following:
"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
"n":integer
Number of neurons.
| Range | 1–100000 |
|---|
"size":integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
"stride":integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
"type":"CONV" | "FC"
Layer type.
| Default | FC |
|---|
actorOptimizer={"method":"ADAM" | "SGD", method-specific-parameters}
The optimization method for the actor model: SGD or ADAM.
| Alias | actorOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
checkPointBestOut={casouttable}
Specifies a table that to save the best observed model.
For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.
checkPointFreq=integer
Specifies how often to save training parameters.
| Range | 1–100000000 |
|---|
checkPointIn={castable}
Specifies a table that includes a saved model for warm starting.
For more information about specifying the checkPointIn parameter, see the common castable parameter.
checkPointOut={casouttable}
Specifies a table that is used to save the model.
For more information about specifying the checkPointOut parameter, see the common casouttable parameter.
criticModel=[{nnParams-1} <, {nnParams-2}, ...>]
Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.
The nnParams value can be one or more of the following:
"act":"IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
"n":integer
Number of neurons.
| Range | 1–100000 |
|---|
"size":integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
"stride":integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
"type":"CONV" | "FC"
Layer type.
| Default | FC |
|---|
criticOptimizer={"method":"ADAM" | "SGD", method-specific-parameters}
The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.
| Alias | criticOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
* environment={"type":"BUILTIN" | "REMOTE", type-specific-parameters}
The reinforcement learning environment.
| Alias | env |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
gamma=double
Discount factor used to compute long term value.
| Default | 0.99 |
|---|---|
| Range | 0–1 |
gpu={gpuOptions}
Specifies GPU related options.
The gpuOptions value can be one or more of the following:
"deterministic":True | False
Specifies that the action will use deterministic GPU functions (may be slower).
| Default | False |
|---|
"device":integer
Specifies a specific GPU device id.
| Minimum value | 0 |
|---|
"enable":True | False
Specifies that action will use a GPU to perform model calculations.
| Default | False |
|---|
maxNormOfGradient=double
Adjust the magnitude of the gradient so that it is less than this value.
| Range | 1–100000000 |
|---|
* modelOut={casouttable}
Specifies a table that is used to save model structure and weights.
For more information about specifying the modelOut parameter, see the common casouttable parameter.
nThreads=integer
Number of CPU threads to use.
| Range | 0–1024 |
|---|
numEpisodes=integer
Number of episodes to train before terminating.
| Default | 10 |
|---|---|
| Range | 1–100000000 |
numTestEpisodes=integer
Specifies number of episodes to run when running test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
* pgAlgorithm={"type":"ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters}
The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C
| Alias | pgAlg |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
seed=double
Seed for random number generator.
| Minimum value | 0 |
|---|
testInterval=integer
Specifies how often to run test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
Parameters for method="ADAM"
"beta1":double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
"beta2":double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
"epsilon":double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
"learningRate":double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
"miniBatchSize":integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
"regL2":double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
"useAMSGrad":True | False
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | False |
|---|
Parameters for method="SGD"
"damping":double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
"learningRate":double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
"miniBatchSize":integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
"momentum":double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
"regL2":double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
"useNesterov":True | False
Enables Nesterov momentum.
| Default | False |
|---|
Parameters for method="ADAM"
"beta1":double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
"beta2":double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
"epsilon":double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
"learningRate":double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
"miniBatchSize":integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
"regL2":double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
"useAMSGrad":True | False
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | False |
|---|
Parameters for method="SGD"
"damping":double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
"learningRate":double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
"miniBatchSize":integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
"momentum":double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
"regL2":double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
"useNesterov":True | False
Enables Nesterov momentum.
| Default | False |
|---|
Parameters for type="BUILTIN"
* "name":"string"
Parameters for type="REMOTE"
"imageStack":integer
Specifies number of consecutive observations to be stacked as an input observation.
| Default | 1 |
|---|---|
| Range | 1–10000 |
* "name":"string"
The environment name used in RL.
"render":True | False
Whether render the environment display or not.
| Default | False |
|---|
"renderFreq":integer
Specifies how frequently the environment is rendered.
| Alias | renderFrequency |
|---|---|
| Default | 1 |
| Range | 1–10000000 |
"renderSleep":double
Time to sleep (in seconds) between rendering frames.
| Default | 0 |
|---|
"seed":integer
Seed of the environment.
| Default | 0 |
|---|
* "url":"string"
The url:port used for in RL training.
| Default | "127.0.0.1:8080" |
|---|
"wrapper":["string-1" <, "string-2", ...>]
A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.
Parameters for type="ACTORCRITIC"
"entropyCoef":double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
"GAELambda":double
Specifies the value of the lambda in generalized advantage estimation (GAE).
| Default | 0.9 |
|---|---|
| Range | 0–1 |
"nRollouts":integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
"useGAELoss":True | False
Use GAE Loss.
| Default | False |
|---|
Parameters for type="REINFORCE"
"entropyCoef":double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
"nRollouts":integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
Parameters for type="REINFORCERTG"
"entropyCoef":double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
"nRollouts":integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
rlTrainPolicyGradient Action
Do reinforcement learning training using Policy Gradient algorithm.
| Alias | rlTrainPG |
|---|
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that includes a saved model for warm starting. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
Specifies a table that to save the best observed model. | |
|
— |
Specifies a table that is used to save the model. | |
|
required parametermodelOut |
— |
Specifies a table that is used to save model structure and weights. |
Parameter Descriptions
* actorModel=list( list(nnParams-1) <, list(nnParams-2), ...>)
Specifies the properties of the layers of the neural net used to build the actor network.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
actorOptimizer=list(method="ADAM" | "SGD", method-specific-parameters)
The optimization method for the actor model: SGD or ADAM.
| Alias | actorOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
checkPointBestOut=list(casouttable)
Specifies a table that to save the best observed model.
For more information about specifying the checkPointBestOut parameter, see the common casouttable parameter.
checkPointFreq=integer
Specifies how often to save training parameters.
| Range | 1–100000000 |
|---|
checkPointIn=list(castable)
Specifies a table that includes a saved model for warm starting.
For more information about specifying the checkPointIn parameter, see the common castable parameter.
checkPointOut=list(casouttable)
Specifies a table that is used to save the model.
For more information about specifying the checkPointOut parameter, see the common casouttable parameter.
criticModel=list( list(nnParams-1) <, list(nnParams-2), ...>)
Specifies the properties of the layers of the neural net used to build the critic network. If this value is left unspecified, the value from actorModel is copied.
The nnParams value can be one or more of the following:
act="IDENTITY" | "RELU" | "SIGMOID" | "TANH"
Activation function.
| Default | RELU |
|---|
n=integer
Number of neurons.
| Range | 1–100000 |
|---|
size=integer
Size (width and height) of kernel (for convolution layer) or window (for pooling layer).
| Range | 1–1000 |
|---|
stride=integer
Stride of kernel (for convolution layer) or window (for pooling layer).
| Default | 1 |
|---|---|
| Range | 1–10000 |
type="CONV" | "FC"
Layer type.
| Default | FC |
|---|
criticOptimizer=list(method="ADAM" | "SGD", method-specific-parameters)
The optimization method for the critic model: SGD or ADAM. If this value is left unspecified, the value from actorOptimizer is copied.
| Alias | criticOptimization |
|---|
The value that you specify for the method parameter determines the other parameters that apply.
* environment=list(type="BUILTIN" | "REMOTE", type-specific-parameters)
The reinforcement learning environment.
| Alias | env |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
gamma=double
Discount factor used to compute long term value.
| Default | 0.99 |
|---|---|
| Range | 0–1 |
gpu=list(gpuOptions)
Specifies GPU related options.
The gpuOptions value can be one or more of the following:
deterministic=TRUE | FALSE
Specifies that the action will use deterministic GPU functions (may be slower).
| Default | FALSE |
|---|
device=integer
Specifies a specific GPU device id.
| Minimum value | 0 |
|---|
enable=TRUE | FALSE
Specifies that action will use a GPU to perform model calculations.
| Default | FALSE |
|---|
maxNormOfGradient=double
Adjust the magnitude of the gradient so that it is less than this value.
| Range | 1–100000000 |
|---|
* modelOut=list(casouttable)
Specifies a table that is used to save model structure and weights.
For more information about specifying the modelOut parameter, see the common casouttable parameter.
nThreads=integer
Number of CPU threads to use.
| Range | 0–1024 |
|---|
numEpisodes=integer
Number of episodes to train before terminating.
| Default | 10 |
|---|---|
| Range | 1–100000000 |
numTestEpisodes=integer
Specifies number of episodes to run when running test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
* pgAlgorithm=list(type="ACTORCRITIC" | "REINFORCE" | "REINFORCERTG", type-specific-parameters)
The policy gradient algorithm type, REINFORCE, REINFORCERTG, ActorCritic, A2C
| Alias | pgAlg |
|---|
The value that you specify for the type parameter determines the other parameters that apply.
seed=double
Seed for random number generator.
| Minimum value | 0 |
|---|
testInterval=integer
Specifies how often to run test.
| Default | 100 |
|---|---|
| Range | 1–100000000 |
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=TRUE | FALSE
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | FALSE |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=TRUE | FALSE
Enables Nesterov momentum.
| Default | FALSE |
|---|
Parameters for method="ADAM"
beta1=double
The exponential decay rate for the first moment estimates.
| Default | 0.9 |
|---|---|
| Range | 0–1 |
beta2=double
The exponential decay rate for the second moment estimates.
| Default | 0.999 |
|---|---|
| Range | 0–1 |
epsilon=double
The value to add into the second moment of the gradient in the denominator to improve numerical stability.
| Alias | eps |
|---|---|
| Default | 1E-08 |
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useAMSGrad=TRUE | FALSE
Use AMSGrad variant of ADAM. See paper: On the Convergence of Adam and Beyond.
| Default | FALSE |
|---|
Parameters for method="SGD"
damping=double
Damping for momentum.
| Default | 0 |
|---|---|
| Range | 0–1 |
learningRate=double
Specifies the learning rate parameter for SGD.
| Alias | lr |
|---|---|
| Default | 0.001 |
| Range | 0–1 |
miniBatchSize=integer
Specifies minibatch size to be used when training the Q model.
| Default | 64 |
|---|---|
| Range | 1–100000 |
momentum=double
Specifies the value for momentum for SGD.
| Default | 0 |
|---|---|
| Range | 0–1 |
regL2=double
Specifies the L2 regularization parameter, the value must be nonnegative.
| Alias | weightDecay |
|---|---|
| Default | 0 |
| Range | 0–1 |
useNesterov=TRUE | FALSE
Enables Nesterov momentum.
| Default | FALSE |
|---|
Parameters for type="BUILTIN"
* name="string"
Parameters for type="REMOTE"
imageStack=integer
Specifies number of consecutive observations to be stacked as an input observation.
| Default | 1 |
|---|---|
| Range | 1–10000 |
* name="string"
The environment name used in RL.
render=TRUE | FALSE
Whether render the environment display or not.
| Default | FALSE |
|---|
renderFreq=integer
Specifies how frequently the environment is rendered.
| Alias | renderFrequency |
|---|---|
| Default | 1 |
| Range | 1–10000000 |
renderSleep=double
Time to sleep (in seconds) between rendering frames.
| Default | 0 |
|---|
seed=integer
Seed of the environment.
| Default | 0 |
|---|
* url="string"
The url:port used for in RL training.
| Default | "127.0.0.1:8080" |
|---|
wrapper=list("string-1" <, "string-2", ...>)
A list of strings that specifies the wrappers that will wrap the environment. The wrappers are used to preprocess the environment observations and rewards.
Parameters for type="ACTORCRITIC"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
GAELambda=double
Specifies the value of the lambda in generalized advantage estimation (GAE).
| Default | 0.9 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
useGAELoss=TRUE | FALSE
Use GAE Loss.
| Default | FALSE |
|---|
Parameters for type="REINFORCE"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |
Parameters for type="REINFORCERTG"
entropyCoef=double
Specifies the value for the entropy coefficient in policy gradient algorithms.
| Default | 0.01 |
|---|---|
| Range | 0–1 |
nRollouts=integer
Specifies the number of rollouts to generate data.
| Default | 1 |
|---|---|
| Range | 1–1024 |