The LMIXED Procedure
OPTIMIZATION Statement
OPTIMIZATION <options>;
The OPTIMIZATION statement specifies the technique and relevant specifications that are used in nonlinear optimization of REML and ML functions.
You can specify the following options. The effect of some options depends on the technique that is specified in the TECHNIQUE= option.
- ABSCONV=r
specifies an absolute function convergence criterion. For minimization, termination requires , where is the vector of parameters in the optimization and is the objective function. The default value of r is the negative square root of the largest double-precision value, which serves only as a protection against overflows.
- ABSFCONV=r
specifies an absolute function difference convergence criterion. For all techniques except Nelder-Mead simplex (TECHNIQUE=NMSIMP), termination requires a small change of the function value in successive iterations:
Here, denotes the vector of parameters that participate in the optimization and is the objective function. The same formula is used for the Nelder-Mead simplex technique, but is defined as the vertex that has the lowest function value and is defined as the vertex that has the highest function value in the simplex. By default, ABSFCONV=0.
- ABSGCONV=r
specifies an absolute gradient convergence criterion. Termination requires the maximum absolute gradient element to be small:
Here, denotes the vector of parameters that participate in the optimization and is the gradient of the objective function with respect to the j parameter. This criterion is not used by the Nelder-Mead simplex technique. By default, ABSGCONV=1E–5.
- FCONV=r
specifies a relative function convergence criterion. For all techniques except the Nelder-Mead simplex technique, termination requires a small relative change of the function value in successive iterations:
Here, denotes the vector of parameters that participate in the optimization and is the objective function. The same formula is used for the Nelder-Mead simplex technique, but is defined as the vertex that has the lowest function value and is defined as the vertex that has the highest function value in the simplex.
The default is , where FDIGITS is and is the machine precision.
- FCONV2=r
specifies a second function convergence criterion. For all techniques except the Nelder-Mead simplex technique, termination requires a small predicted reduction of the objective function:
The predicted reduction
is computed by approximating the objective function f by the first two terms of the Taylor series and substituting the Newton step,
For the Nelder-Mead simplex technique, termination requires a small standard deviation of the function values of the simplex vertices , ,
where . If there are boundary constraints active at , the mean and standard deviation are computed only for the unconstrained vertices.
The default value is r = 1E–6 for the Nelder-Mead simplex technique and r= 0 otherwise.
- GCONV=r
specifies a relative gradient convergence criterion. For all techniques except the conjugate-gradient technique and the Nelder-Mead simplex technique, termination requires that the normalized predicted function reduction be small,
Here, denotes the vector of parameters that participate in the optimization, is the objective function, and is the gradient. For the conjugate-gradient technique (where a reliable Hessian estimate is not available), the following criterion is used:
This criterion is not used by the Nelder-Mead simplex technique. The default value is r=1E–8.
- GCONV2=r
specifies another relative gradient convergence criterion. For Newton-Raphson with ridging and Newton-Raphson techniques, the following criterion of Browne (1982) is used:
This criterion is not used by the other techniques. By default, GCONV2=0.
- HESSTYPE=BFGS | FDFULL | FULL
specifies the type of Hessian to use in the optimization. You can specify the following values:
- BFGS
uses the original Broyden, Fletcher, Goldfarb, and Shanno (BFGS) approximation of the inverse Hessian matrix.
- FDFULL
uses the finite-difference Hessian.
- FULL
uses the exact Hessian or average information Hessian, assuming that it is available.
The default value is HESSTYPE=BFGS when there is no Hessian function available in the optimization.
This option is supported only when TECHNIQUE=ACTIVESET or IPDIRECT.
- MAXFUNC=n
specifies the maximum number of function calls in the optimization process. The optimization can terminate only after completing a full iteration. Therefore, the number of function calls that are performed can exceed n. The default values are as follows, depending on the optimization technique:
When TECHNIQUE=TRUREG, NRRIDG, or NEWRAP, MAXFUNC=123 by default.
When TECHNIQUE=QUANEW or DBLDOG, MAXFUNC=500 by default.
When TECHNIQUE=CONGRA, MAXFUNC=1,000 by default.
When TECHNIQUE=NMSIMP, MAXFUNC=3,000 by default.
- MAXITER=n
specifies the maximum number of iterations in the optimization process. These default values also apply when is specified as a missing value. The default values are as follows, depending on the optimization technique:
When TECHNIQUE=TRUREG, NRRIDG, or NEWRAP, MAXITER=50 by default.
When TECHNIQUE=QUANEW or DBLDOG, MAXITER=200 by default.
When TECHNIQUE=CONGRA, MAXITER=400 by default.
When TECHNIQUE=NMSIMP, MAXITER=1,000 by default.
- MAXTIME=r
specifies an upper limit of seconds of CPU time for the optimization process. The default value is the largest floating-point double representation of your computer. The time specified by the MAXTIME= option is checked only once at the end of each iteration. Therefore, the actual run time can be longer than r.
- MINITER=n
specifies the minimum number of iterations. The default value is 0. If you request more iterations than are needed for convergence to a stationary point, the optimization algorithms can behave strangely. For example, the effect of rounding errors can prevent the algorithm from continuing for the required number of iterations.
- OPTTOL=r
specifies the criterion value by which the current iterate is determined to be an acceptable approximation of a local optimum, where r is a positive real number. Specifically, the optimizer determines that the current iterate is a local optimum when the norm of the scaled vector of the optimality conditions is less than r. By default, OPTTOL=1E–5.
This option is relevant only when you specify TECHNIQUE=ACTIVESET or IPDIRECT.
-
TECHNIQUE=keyword
TECH=keyword specifies the optimization technique for obtaining restricted maximum likelihood estimation (REML) or maximum likelihood estimation (ML). There is no algorithm for optimizing general nonlinear functions that always finds the global optimum for a general nonlinear optimization problem in a reasonable amount of time. Because no single optimization technique is always superior to others, PROC LMIXED provides a variety of optimization techniques that work well in various circumstances.
For more information about traditional techniques such as Newton-Raphson optimization, see the section Optimization Options in the chapter “Shared Concepts” of SAS Visual Statistics: Procedures.
For more information about modern techniques such as active-set and interior-point direct optimization, see the chapter “The Nonlinear Programming Solver” in SAS/OR User's Guide: Mathematical Programming.
You can specify the following keywords:
- ACTIVESET | AS
performs an active-set optimization.
- CONGRA
performs a conjugate-gradient optimization.
- DBLDOG
performs a version of double-dogleg optimization.
- IPDIRECT | IP
performs an interior-point direct optimization.
- NEWRAP
performs a Newton-Raphson optimization that combines a line-search algorithm with ridging.
- NMSIMP
performs a Nelder-Mead simplex optimization.
- NONE
performs no optimization.
- NRRIDG
performs a Newton-Raphson optimization with ridging.
- QUANEW
performs a dual quasi-Newton optimization.
- TRUREG
performs a trust-region optimization.
By default, TECHNIQUE=NRRIDG when DMMETHOD=DENSE. Otherwise, TECHNIQUE=QUANEW by default.
- XCONV=r
specifies the relative parameter convergence criterion, which depends on the technique that is specified in the TECHNIQUE= option as follows:
For all techniques except the Nelder-Mead simplex technique, termination requires a small relative parameter change in subsequent iterations:
For the Nelder-Mead simplex technique, the same formula is used, but is defined as the vertex that has the lowest function value and is defined as the vertex that has the highest function value in the simplex.
The default value is r = 1E–8 for the NMSIMP technique and r = 0 otherwise.