Syntax Supported by the IML Procedure and the iml Action
FDDSOLVE Call
CALL FDDSOLVE (f, g, h, "fun", x0, <, opt> <, "grd"> ) ;
CALL FDDSOLVE (f, g, h, "fun", x0) <OPT=opt> <GRD="grd"> ;
This subroutine is supported by the IML procedure and the iml action.
The FDDSOLVE subroutine uses finite differences to approximate derivatives at a specified point in the domain of a differentiable function. Derivatives are important for many tasks in computational statistics, including optimization, solving nonlinear equations, and fitting statistical models to data.
The FDDSOLVE subroutine approximates derivatives for a function for
. When
, the function is a scalar-valued function. When
, the function is a vector-valued function. The derivatives are evaluated at an n-dimensional point,
, in the domain of f.
The FDDSOLVE subroutine calls the user-defined module "fun" to evaluate the function f and its derivatives at the point x0.
-
If the module "fun" returns a scalar value, the FDDSOLVE subroutine computes the following quantities:
-
If the module "fun" returns a column vector of m function values, the subroutine computes the following quantities:
The subroutine returns a missing value for any result that cannot be computed.
The FDDSOLVE subroutine requires the following input arguments:
The "fun" argument specifies a module that returns the value of the function, f. For vector-valued functions, the module should return a column vector.
The x0 argument is a row vector of length n that defines the point at which the functions and derivatives are computed.
In addition, you can specify the following input arguments:
-
The opt argument is a three-element vector. It is optional for scalar-valued functions, but it is required for vector-valued functions. You can use the OPT keyword to specify the vector opt.
opt[1]indicates how to approximate the first derivatives in the gradient vector or Jacobian matrix. Ifopt[1]is missing or equal to 0, the subroutine uses forward-difference derivatives. Otherwise, it uses central-difference derivatives. Forward differences are faster but less accurate than central differences. The default is to use forward-difference derivatives.opt[2]indicates how to approximate the Hessian matrix for a scalar-valued function. Ifopt[2]is missing or equal to 0, the subroutine uses forward-difference derivatives. Otherwise, it uses central-difference derivatives. The default is to use forward-difference derivatives. This option is ignored for vector-valued functions.opt[3]is relevant only to vector-valued functions. It specifies the number of elements that are returned by the module "fun". Ifopt[3]is missing or less than 1, it is set to 1.
The optional "grd" argument specifies a user-defined function module that returns the analytical gradient of the function at the point x0. You can use the GRD keyword to specify the module "grd". If specified, this module is called to compute the gradient, and the Hessian is approximated by using first-order derivatives of the gradient function. For vector-valued functions, the "grd" argument is ignored.
The FDDSOLVE subroutine is useful for checking first-order and second-order analytical derivatives. It is also useful for approximating derivatives when a formula is not available.
Derivatives for a Scalar-Valued Function
The following example demonstrates how to call the FDDSOLVE subroutine. In the unconstrained Rosenbrock problem, the objective function is
The gradient and the Hessian are
At the point , these expressions evaluate to
The following statements define the Rosenbrock function and use the FDDSOLVE subroutine to compute the gradient and the Hessian. The results are shown in Figure 117.
start F_ROSEN(x);
y1 = 10 * (x[2] - x[1] * x[1]);
y2 = 1 - x[1];
f = 0.5 * (y1 * y1 + y2 * y2);
return(f);
finish F_ROSEN;
x = {2 7};
call fddsolve(f, grad, Hess, "F_ROSEN", x);
print f, grad, Hess;
Figure 117: Gradient and Hessian at a Point
| f |
|---|
| 450.5 |
| grad | |
|---|---|
| -1199 | 300.00001 |
| Hess | |
|---|---|
| 1001.0438 | -400.0018 |
| -400.0018 | 99.99992 |
Figure 117 shows forward-difference approximations to the derivatives. You can compute central-difference derivatives by using the OPT keyword: OPT={1 1}.
If you define a function that evaluates the gradient of the function, you can use the GRD keyword to use the analytic gradient. For example, the following call to the FDDSOLVE subroutine uses the G_ROSEN function to evaluate the analytical gradient. The Hessian is computed by using first-order differences of the gradient function.
start G_ROSEN(x);
g = j(1,2,0.);
g[1] = -200*x[1]*(x[2]-x[1]*x[1]) - (1-x[1]);
g[2] = 100*(x[2]-x[1]*x[1]);
return(g);
finish;
call fddsolve(f, grad, Hess, "F_ROSEN", x) grd="G_ROSEN";
Derivatives for a Vector-Valued Function
If the Rosenbrock problem is considered from a least squares perspective, the two component functions are
The Jacobian and the crossproduct of the Jacobian are
At the point , these matrices evaluate to
The following statements define the Rosenbrock problem in a least squares framework and use the FDDSOLVE subroutine to compute the Jacobian and the crossproduct matrix. The OPT keyword is used to specify the vector opt. Because the last element of opt is 2, the FDDSOLVE subroutine allocates memory for a least squares problem with two functions, and
.
start F_ROSEN(x);
y = j(2, 1, 0);
y[1] = 10 * (x[2] - x[1] * x[1]);
y[2] = 1 - x[1];
return(y);
finish F_ROSEN;
x = {2 7};
parms = {. . 2};
call fddsolve(f, jac, crpj, "F_ROSEN", x) opt=parms;
print f, jac, crpj;
Figure 118: Jacobian and Crossproduct Matrix at a Point
| f |
|---|
| 30 |
| -1 |
| jac | |
|---|---|
| -40 | 10 |
| -1 | 0 |
| crpj | |
|---|---|
| 1601 | -400 |
| -400 | 100 |
For this example, the domain and the range of the objective function are both two-dimensional. However, in general, you can analyze the derivatives of a function for any values of n and m. The dimension of the domain is determined by the number of elements in the x0 vector. The dimension of the range is determined by the last element of the vector opt.
Formulas for Finite-Difference Gradients
In the following formulas for gradients, Jacobians, and Hessians, the vector is the ith basis row vector
, where the 1 is in the ith coordinate position. The symbol
denotes machine precision, which is approximately 2.22E–16.
The FDDSOLVE subroutine approximates derivatives by using finite-difference formulas. The formulas depend on whether you request a forward-difference or central-difference derivative.
If you do not provide an analytic gradient function, then the following formulas are used:
- Forward-difference gradient:
-
The ith component of the gradient vector at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the function f at
locations.
- Central-difference gradient:
-
The ith component of the gradient vector at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the function f at
locations.
Formulas for Finite-Difference Jacobians
The finite-difference formulas for a Jacobian are similar to the formulas for the gradient. The th position of the Jacobian matrix is the partial derivative of the ith component function with respect to the jth variable:
.
- Forward-difference Jacobian:
-
The
th element of the Jacobian matrix at the point x is approximated by using the formula
The step size is defined by
. The formula evaluates each component function
at
locations.
- Central-difference Jacobian:
-
The step size is defined by
. The formula evaluates each component function
at
locations.
Formulas for Finite-Difference Hessians
If you do not provide an analytical gradient, then the Hessian matrix is approximated by using a second-order finite-difference formula.
- Forward-difference Hessian:
-
The
th component of the Hessian matrix at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the function f at
locations.
- Central-difference Hessian:
-
The
th component of the Hessian matrix at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the function f at
locations.
If you provide an analytic gradient function, the Hessian is the result of first-order calls to the gradient function, as follows:
- Forward-difference Hessian:
-
The
th component of the Hessian matrix at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the gradient at
locations.
- Central-difference Hessian:
-
The
th component of the Hessian matrix at the point x is computed by using the formula
The step size is defined by
. The formula evaluates the gradient at
locations.