Shared Concepts

Optimization Options

This section applies to the SEVSELECT procedure.

This section describes options that are typically available in the PROC statement or the NLOPTIONS statement of the procedures in this book that perform optimizations. The following notation is used to describe the options. bold-italic beta denotes the p times 1 vector of parameters for the optimization and beta Subscript i is its ith element. The objective function being minimized, its p times 1 gradient vector, and its p times p Hessian matrix are denoted as f left-parenthesis bold-italic beta right-parenthesis, bold g left-parenthesis bold-italic beta right-parenthesis, and bold upper H left-parenthesis bold-italic beta right-parenthesis, respectively. The gradient with respect to the ith parameter is denoted as g Subscript i Baseline left-parenthesis bold-italic beta right-parenthesis. Superscripts in parentheses denote the iteration count; for example, f left-parenthesis bold-italic beta right-parenthesis Superscript left-parenthesis k right-parenthesis is the value of the objective function at iteration k.

ABSCONV=r
ABSTOL=r

specifies an absolute function convergence criterion. For minimization, termination requires f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis less-than-or-equal-to r, where bold-italic beta is the vector of parameters in the optimization and f left-parenthesis dot right-parenthesis is the objective function. The default value of r is the negative square root of the largest double-precision value, which serves only as a protection against overflows.

ABSFCONV=r
ABSFTOL=r

specifies an absolute function difference convergence criterion. For all techniques except NMSIMP, termination requires a small change of the function value in successive iterations:

StartAbsoluteValue f left-parenthesis bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis Baseline right-parenthesis minus f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis EndAbsoluteValue less-than-or-equal-to sans-serif-italic r

Here, bold-italic beta is the vector of parameters in the optimization and f left-parenthesis dot right-parenthesis is the objective function. The same formula is used for the NMSIMP technique, but bold-italic beta Superscript left-parenthesis k right-parenthesis is defined as the vertex that has the lowest function value and bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis is defined as the vertex that has the highest function value in the simplex. By default, ABSFCONV=0.

ABSGCONV=r
ABSGTOL=r

specifies an absolute gradient convergence criterion. Termination requires the maximum absolute gradient element to be small:

max Underscript j Endscripts StartAbsoluteValue g Subscript j Baseline left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis EndAbsoluteValue less-than-or-equal-to sans-serif-italic r

Here, bold-italic beta is the vector of parameters in the optimization and g Subscript j Baseline left-parenthesis dot right-parenthesis is the gradient of the objective function with respect to the jth parameter. This criterion is not used by the NMSIMP technique. By default, ABSGCONV=1E–5.

ABSXCONV=r
ABSXTOL=r

specifies an absolute parameter convergence criterion: For all techniques except NMSIMP, termination requires a small Euclidean distance between successive parameter vectors,

parallel-to bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline minus bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis Baseline parallel-to less-than-or-equal-to sans-serif-italic r

For the NMSIMP technique, termination requires either a small length alpha Superscript left-parenthesis k right-parenthesis of the vertices of a restart simplex,

alpha Superscript left-parenthesis k right-parenthesis Baseline less-than-or-equal-to sans-serif-italic r

or a small simplex size,

delta Superscript left-parenthesis k right-parenthesis Baseline less-than-or-equal-to sans-serif-italic r

where the simplex size delta Superscript left-parenthesis k right-parenthesis is defined as the L1 distance from the simplex vertex bold-italic xi Superscript left-parenthesis k right-parenthesis that has the smallest function value to the other p simplex points bold-italic beta Subscript l Superscript left-parenthesis k right-parenthesis Baseline not-equals bold-italic xi Superscript left-parenthesis k right-parenthesis:

delta Superscript left-parenthesis k right-parenthesis Baseline equals sigma-summation Underscript bold-italic beta Subscript l Baseline not-equals y Endscripts parallel-to bold-italic beta Subscript l Superscript left-parenthesis k right-parenthesis Baseline minus bold-italic xi Superscript left-parenthesis k right-parenthesis parallel-to

The default is r = 1E–8 for the NMSIMP technique and r = 0 otherwise.

FCONV=r
FTOL=r

specifies a relative function difference convergence criterion. For all techniques except NMSIMP, termination requires a small relative change of the function value in successive iterations,

StartFraction StartAbsoluteValue f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis minus f left-parenthesis bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis Baseline right-parenthesis EndAbsoluteValue Over max left-parenthesis StartAbsoluteValue f left-parenthesis bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis Baseline right-parenthesis EndAbsoluteValue comma FSIZE right-parenthesis EndFraction less-than-or-equal-to sans-serif-italic r

where FSIZE is defined by the FSIZE= option. Here, bold-italic beta denotes the vector of parameters that participate in the optimization, and f left-parenthesis dot right-parenthesis is the objective function. The same formula is used for the NMSIMP technique, but bold-italic beta Superscript left-parenthesis k right-parenthesis is defined as the vertex that has the lowest function value and bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis is defined as the vertex that has the highest function value in the simplex.

The default value is r=2 times epsilon where epsilon is the machine precision, which is the smallest double-precision floating-point number such that 1 plus epsilon greater-than 1.

FCONV2=r
FTOL2=r

specifies a second function convergence criterion. For all techniques except NMSIMP, termination requires a small predicted reduction of the objective function:

d f Superscript left-parenthesis k right-parenthesis Baseline almost-equals f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis minus f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline plus bold s Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis

The predicted reduction

StartLayout 1st Row 1st Column d f Superscript left-parenthesis k right-parenthesis 2nd Column equals minus bold g Superscript left-parenthesis k right-parenthesis prime Baseline bold s Superscript left-parenthesis k right-parenthesis Baseline minus one-half bold s Superscript left-parenthesis k right-parenthesis prime Baseline bold upper H Superscript left-parenthesis k right-parenthesis Baseline bold s Superscript left-parenthesis k right-parenthesis Baseline 2nd Row 1st Column Blank 2nd Column equals minus one-half bold s Superscript left-parenthesis k right-parenthesis Super Superscript prime Superscript Baseline bold g Superscript left-parenthesis k right-parenthesis Baseline less-than-or-equal-to sans-serif-italic r EndLayout

is computed by approximating the objective function f by the first two terms of the Taylor series and substituting the Newton step,

bold s Superscript left-parenthesis k right-parenthesis Baseline equals minus left-bracket bold upper H Superscript left-parenthesis k right-parenthesis Baseline right-bracket Superscript negative 1 Baseline bold g Superscript left-parenthesis k right-parenthesis

For the NMSIMP technique, termination requires a small standard deviation of the function values of the p plus 1 simplex vertices bold-italic beta Subscript l Superscript left-parenthesis k right-parenthesis, l equals 0 comma ellipsis comma p,

StartRoot StartFraction 1 Over n plus 1 EndFraction sigma-summation Underscript l Endscripts left-bracket f left-parenthesis bold-italic beta Subscript l Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis minus ModifyingAbove f With bar left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis right-bracket squared EndRoot less-than-or-equal-to sans-serif-italic r

where ModifyingAbove f With bar left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis equals StartFraction 1 Over p plus 1 EndFraction sigma-summation Underscript l Endscripts f left-parenthesis bold-italic beta Subscript l Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis. If there are p Subscript a c t boundary constraints active at bold-italic beta Superscript left-parenthesis k right-parenthesis, the mean and standard deviation are computed only for the n plus 1 minus p Subscript a c t unconstrained vertices.

The default value is r = 1E–6 for the NMSIMP technique and r= 0 otherwise.

FSIZE=r

specifies the FSIZE parameter of the relative function and relative gradient termination criteria. The default value is r = 0. For more information, see the FCONV= and GCONV= options.

GCONV=r
GTOL=r

specifies a relative gradient convergence criterion. For all techniques except CONGRA and NMSIMP, termination requires that the normalized predicted function reduction be small:

StartFraction bold g left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis prime left-bracket bold upper H Superscript left-parenthesis k right-parenthesis Baseline right-bracket Superscript negative 1 Baseline bold g left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis Over max left-parenthesis StartAbsoluteValue f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis EndAbsoluteValue comma FSIZE right-parenthesis EndFraction less-than-or-equal-to sans-serif-italic r

where FSIZE is defined by the FSIZE= option. Here, bold-italic beta denotes the vector of parameters that participate in the optimization, f left-parenthesis dot right-parenthesis is the objective function, and bold g left-parenthesis dot right-parenthesis is the gradient. For the CONGRA technique (where a reliable Hessian estimate bold upper H is not available), the following criterion is used:

StartFraction parallel-to bold g left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis parallel-to Subscript 2 Superscript 2 Baseline parallel-to bold s left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis parallel-to Over parallel-to bold g left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis minus bold g left-parenthesis bold-italic beta Superscript left-parenthesis k minus 1 right-parenthesis Baseline right-parenthesis parallel-to Subscript 2 Baseline max left-parenthesis StartAbsoluteValue f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis EndAbsoluteValue comma FSIZE right-parenthesis EndFraction less-than-or-equal-to sans-serif-italic r

This criterion is not used by the NMSIMP technique. By default, GCONV=1E–8.

GCONV2=r
GTOL2=r

specifies another relative gradient convergence criterion. For the TRUREG, NRRIDG, and NEWRAP techniques, the following criterion of Browne (1982) is used:

max Underscript j Endscripts StartFraction StartAbsoluteValue bold g Subscript j Baseline left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis EndAbsoluteValue Over StartRoot f left-parenthesis bold-italic beta Superscript left-parenthesis k right-parenthesis Baseline right-parenthesis bold upper H Subscript j comma j Superscript left-parenthesis k right-parenthesis Baseline EndRoot EndFraction less-than-or-equal-to sans-serif-italic r

This criterion is not used by the other techniques.

By default, GCONV2=0.

MAXFUNC=n
MAXFU=n

specifies the maximum number n of function calls in the optimization process. The default values are as follows, depending on the optimization technique:

  • TRUREG, NRRIDG, and NEWRAP: 125

  • QUANEW and DBLDOG: 500

  • CONGRA: 1,000

  • NMSIMP: 3,000

The optimization can terminate only after completing a full iteration. Therefore, the number of function calls that are actually performed can exceed the number that is specified by this option. You can specify the optimization technique in the TECHNIQUE= option.

MAXITER=n
MAXIT=n

specifies the maximum number n of iterations in the optimization process. The default values are as follows, depending on the optimization technique:

  • TRUREG, NRRIDG, and NEWRAP: 50

  • QUANEW and DBLDOG: 200

  • CONGRA: 400

  • NMSIMP: 1,000

These default values also apply when n is specified as a missing value. You can specify the optimization technique in the TECHNIQUE= option.

MAXTIME=r

specifies an upper limit of r seconds of CPU time for the optimization process. The time specified by r is checked only once at the end of each iteration. Therefore, the actual running time can be longer than r. The default value is the largest floating-point double representation of your computer.

MINITER=n
MINIT=n

specifies the minimum number of iterations. If you request more iterations than are actually needed for convergence to a stationary point, the optimization algorithms can behave strangely. For example, the effect of rounding errors can prevent the algorithm from continuing for the required number of iterations. By default, MINITER=0.

TECHNIQUE=technique
TECH=technique

specifies the optimization technique for obtaining maximum likelihood estimates. You can specify one of the following techniques:

CONGRA

performs a conjugate-gradient optimization.

DBLDOG

performs a version of double-dogleg optimization.

NEWRAP

performs a Newton-Raphson optimization with line search.

NMSIMP

performs a Nelder-Mead simplex optimization.

NONE

performs no optimization.

NRRIDG

performs a Newton-Raphson optimization with ridging.

QUANEW

performs a dual quasi-Newton optimization.

TRUREG

performs a trust-region optimization

The default method varies by the procedure. For the SEVSELECT procedure, the default is TECHNIQUE=TRUREG.

For more information, see the section Choosing an Optimization Algorithm.

XCONV=r
XTOL=r

specifies the relative parameter convergence criterion. Convergence requires a small relative parameter change in subsequent iterations,

StartFraction max Underscript j Endscripts StartAbsoluteValue beta Subscript j Superscript left-parenthesis k right-parenthesis Baseline minus beta Subscript j Superscript left-parenthesis k minus 1 right-parenthesis Baseline EndAbsoluteValue Over max left-parenthesis StartAbsoluteValue beta Subscript j Superscript left-parenthesis k right-parenthesis Baseline EndAbsoluteValue comma StartAbsoluteValue beta Subscript j Superscript left-parenthesis k minus 1 right-parenthesis Baseline EndAbsoluteValue comma XSIZE right-parenthesis EndFraction less-than-or-equal-to sans-serif-italic r

where XSIZE is defined by the XSIZE= option. and beta Subscript j Superscript left-parenthesis i right-parenthesis is the estimate of the jth parameter at iteration i. For the NMSIMP technique, the same formula is used, but beta Subscript j Superscript left-parenthesis k right-parenthesis is defined as the vertex that has the lowest function value and beta Subscript j Superscript left-parenthesis k minus 1 right-parenthesis is defined as the vertex that has the highest function value in the simplex. The default value is r = 1E–8 for the NMSIMP technique and r = 0 otherwise.

XSIZE=r

specifies the XSIZE parameter of the relative parameter termination criterion. The value of r must be greater than or equal to 0; the default is sans-serif-italic r equals 0. For more information, see the XCONV= option.

Last updated: January 27, 2023