The FRONTIER Procedure
Getting Started: FRONTIER Procedure
This section illustrates how to estimate stochastic frontier production models. The following statements fit a half-normal model:
proc frontier data=mycas.mydata;
model y = x1 x2 / type=half;
run;
The output variable y and the input variables x1 and x2 are contained in the data table mycas.mydata. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref. The relationship between input and output is specified in the MODEL statement. The distribution of the inefficiency term is specified using the TYPE= option in the MODEL statement. In this example, it is the half-normal distribution. TYPE=EXPONENTIAL specifies an exponential distribution, and TYPE=TRUNCATED specifies a truncated-normal distribution. By default, PROC FRONTIER fits a production model. You can specify the COST option in the MODEL statement to fit a stochastic frontier cost model.
The following example fits the stochastic frontier production model to 1957 state data for the US transportation equipment manufacturing industry. The data were originally published in Zellner and Revankar (1969). The data include observations on aggregate value added (valueadded), aggregate capital service flow (capital), aggregate hours worked (labor), and number of establishments (nfirm) for 25 US states. A Cobb-Douglas production function is estimated for the output variable, log of value added (LNV), using the input variables log of capital (LNK) and log of labor (LNL).
The following statements create the data set:
title1 'Getting Started Example: Stochastic Frontier Production Model';
data transpEqp;
input state$ valueadded capital labor nfirm;
datalines;
Alabama 126.14800262 3.8039999008 31.551000595 68
California 3201.486084 185.44599915 452.84399414 1372
Connecticut 690.66998291 39.712001801 124.0739975 154
Florida 56.296001434 6.5469999313 19.180999756 292
Georgia 304.53100586 11.529999733 45.534000397 71
Illinois 723.02801514 58.986999512 88.39099884 275
Indiana 992.16900635 112.88400269 148.52999878 260
Iowa 35.796001434 2.6979999542 8.0170001984 75
Kansas 494.51501465 10.359999657 86.189002991 76
Kentucky 124.94799805 5.2129998207 12 31
Louisiana 73.32800293 3.7630000114 15.899999619 115
Maine 29.466999054 1.9670000076 6.4699997902 81
Maryland 415.26199341 17.545999527 69.342002869 129
Massachusetts 241.52999878 15.347000122 39.416000366 172
Michigan 4079.5539551 435.10501099 490.38400269 568
Missouri 652.08502197 32.840000153 84.831001282 125
NewJersey 667.11297607 33.291999817 83.032997131 247
NewYork 940.42999268 72.973999023 190.09399414 461
Ohio 1611.8990479 157.97799683 259.91598511 363
Pennsylvania 617.57897949 34.324001312 98.152000427 233
Texas 527.4130249 22.736000061 109.72799683 308
Virginia 174.39399719 7.1729998589 31.301000595 85
Washington 636.94799805 30.806999207 87.962997437 179
WestVirginia 22.700000763 1.5429999828 4.0630002022 15
Wisconsin 349.71099854 22.000999451 52.818000793 142
;
data transpEqp;
set transpEqp;
LNV = log(valueadded);
LNK = log(capital);
LNL = log(labor);
data mycas.transpEqp;
set transpEqp;
run;
The following statements estimate a stochastic frontier half-normal production model that uses the transportation equipment manufacturing industry data:
/*-- Stochastic Frontier Production Model: Half-Normal --*/
proc frontier data=mycas.transpEqp;
model LNV = LNK LNL / type=half;
run;
The "Observation Information" and "Summary Statistics of Dependent Variable" tables shown in Figure 1 contain some details about the observations in the data set and the dependent variable.
Figure 1: Information about Observations and the Dependent Variable
| Getting Started Example: Stochastic Frontier Production Model |
| Observation Information | |
|---|---|
| Number of Observations | 25 |
| Number of Missing Observations | 0 |
| Summary Statistics of Dependent Variable | ||||
|---|---|---|---|---|
| Variable | Mean | Standard Error | Minimum | Maximum |
| LNV | 5.812092 | 1.375304 | 3.122365 | 8.313743 |
The "Model Fit Summary" table shown in Figure 2 lists several details about the model. By default, PROC FRONTIER uses the Newton-Raphson optimization technique, and the estimated covariance matrix is calculated as the inverse of the negative Hessian. For the available optimization and estimation techniques, see the section PROC FRONTIER Statement. The maximum log-likelihood value is shown, in addition to two information measures—Akaike’s information criterion (AIC) and Schwarz’s Bayesian information criterion (SBC)—that can be used to compare competing production models. Smaller values of these criteria indicate better models.
Figure 2: Estimation Summary Table for Half-Normal Production Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | LNV |
| Data Set | TRANSPEQP |
| Model | Production |
| Inefficiency Term Distribution | Half-normal |
| Log Likelihood | 2.469522 |
| Maximum Absolute Gradient | 3.942E-6 |
| Number of Iterations | 11 |
| Optimization Method | Newton-Raphson |
| AIC | 5.060956 |
| SBC | 11.15533 |
| Covariance Estimation | Hessian |
Figure 3 shows the parameter estimates of the model and their standard errors. Both input variables are significant predictors of the log of value added. _Sigma_v is the estimate of the standard deviation of the idiosyncratic component of the error term that is specific to the firm and that could enter the model with either sign. _Sigma_u is the estimate of the standard deviation of the inefficiency component. When the data are in log terms, the inefficiency component becomes a measure of the percentage by which a particular observation fails to achieve an ideal production rate.
Figure 3: Parameter Estimates of Half-Normal Production Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 2.081134 | 0.281648 | 7.39 | <.0001 |
| LNK | 1 | 0.258548 | 0.098766 | 2.62 | 0.0089 |
| LNL | 1 | 0.780245 | 0.119945 | 6.51 | <.0001 |
| _Sigma_v | 1 | 0.175169 | 0.054265 | 3.23 | 0.0012 |
| _Sigma_u | 1 | 0.221507 | 0.123706 | 1.79 | 0.0734 |
Figure 4 shows the estimates of two important variance statistics that are helpful for model diagnosis. Sigma2 is the estimate of the total error variance. It is calculated as the sum of the idiosyncratic error variance and the variance of the technical inefficiency term. Gamma is the estimate of how much the technical inefficiency contributes to the total variation in the model. It is calculated as the ratio of the variance of the technical inefficiency term to the total error variance. In this example, the technical inefficiency explains most of the variation in the model.
Figure 4: Estimates of Important Variance Statistics for Half-Normal Production Model
| Variance Statistics | ||
|---|---|---|
| Parameter | Estimate | Standard Error |
| Sigma2 | 0.079750 | 0.042699 |
| Gamma | 0.615244 | 0.385748 |
The following statements estimate a stochastic frontier exponential production model that uses the same data:
/*-- Stochastic Frontier Production Model: Exponential --*/
proc frontier data=mycas.transpEqp;
model LNV = LNK LNL;
output out=mycas.out_te_exp te1=TE1_exp te2=TE2_exp copyvars=( state );
run;
The default model is the exponential model if you do not specify the TYPE= option in the MODEL statement.
Figure 5 shows some of the results from this production model.
Figure 5: Estimation Summary and Parameter Estimates for Exponential Production Model
| Getting Started Example: Stochastic Frontier Production Model |
| Model Fit Summary | |
|---|---|
| Dependent Variable | LNV |
| Data Set | TRANSPEQP |
| Model | Production |
| Inefficiency Term Distribution | Exponential |
| Log Likelihood | 2.86049 |
| Maximum Absolute Gradient | 4.215E-6 |
| Number of Iterations | 11 |
| Optimization Method | Newton-Raphson |
| AIC | 4.279021 |
| SBC | 10.3734 |
| Covariance Estimation | Hessian |
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 2.069242 | 0.235632 | 8.78 | <.0001 |
| LNK | 1 | 0.262486 | 0.092001 | 2.85 | 0.0043 |
| LNL | 1 | 0.770380 | 0.110964 | 6.94 | <.0001 |
| _Sigma_v | 1 | 0.171392 | 0.038449 | 4.46 | <.0001 |
| _Sigma_u | 1 | 0.135169 | 0.062685 | 2.16 | 0.0311 |
| Variance Statistics | ||
|---|---|---|
| Parameter | Estimate | Standard Error |
| Sigma2 | 0.047646 | 0.015793 |
| Gamma | 0.383467 | 0.285232 |
You can obtain estimates of the efficiency terms by using the OUTPUT statement. PROC FRONTIER offers two types of technical efficiency terms that are proposed by Battese and Coelli (1988) and Jondrow et al. (1982). The OUTPUT statement in the previous and following statements specify the two technical efficiency terms TE1 and TE2. These terms for each observation are written in the data set that you specify in the OUT= option. In these examples, the technical efficiency terms are calculated for each state and included in mycas.out_te_exp for the exponential model and mycas.out_te_half for the half-normal model.
/*-- Obtaining Technical Efficiency Terms for Half-Normal Model --*/
proc frontier data=mycas.transpEqp;
model LNV = LNK LNL / type=half;
output out=mycas.out_te_half te1=TE1_half te2=TE2_half copyvars=( state );
run;
The following statements merge the two output data sets and sort the new data set with respect to the state variable:
/*-- Merging the Two Output Data Sets --*/
data mycas.out_te;
merge mycas.out_te_half mycas.out_te_exp;
by state;
run;
proc sort data=mycas.out_te out=out_te_sorted;
by state;
run;
proc print data = out_te_sorted;
var State TE1_half TE1_exp TE2_half TE2_exp;
title 'TEs from Half-Normal and Exponential Models';
run;
These statistics for each state are shown in Figure 6.
Figure 6: Technical Efficiency for Each State
| TEs from Half-Normal and Exponential Models |
| Obs | state | TE1_half | TE1_exp | TE2_half | TE2_exp |
|---|---|---|---|---|---|
| 1 | Alabama | 0.8231755124 | 0.8690845704 | 0.8178030737 | 0.8642193711 |
| 2 | Californ | 0.8692654721 | 0.9102507468 | 0.8651869818 | 0.9073595422 |
| 3 | Connecti | 0.8318406668 | 0.8783346976 | 0.82667106 | 0.8739011963 |
| 4 | Florida | 0.6016341141 | 0.5623070763 | 0.5959902277 | 0.5541443131 |
| 5 | Georgia | 0.904050962 | 0.9329042531 | 0.901244131 | 0.9310801261 |
| 6 | Illinois | 0.8891712693 | 0.9226121891 | 0.8857974959 | 0.9203129631 |
| 7 | Indiana | 0.8150898568 | 0.8620239532 | 0.8095457608 | 0.8568378457 |
| 8 | Iowa | 0.7853521623 | 0.824843113 | 0.7793217245 | 0.8181327187 |
| 9 | Kansas | 0.9066430761 | 0.9354333085 | 0.9039373908 | 0.9337222102 |
| 10 | Kentucky | 0.9464481304 | 0.960059195 | 0.9452846459 | 0.9593259036 |
| 11 | Louisian | 0.8214223188 | 0.8651803151 | 0.816011214 | 0.8601366506 |
| 12 | Maine | 0.8061243279 | 0.8474149487 | 0.8004097788 | 0.8415937609 |
| 13 | Maryland | 0.8772239419 | 0.914381889 | 0.8734183876 | 0.9116902323 |
| 14 | Massachu | 0.8596337404 | 0.899833078 | 0.855242772 | 0.8964352583 |
| 15 | Michigan | 0.8582037513 | 0.9013389234 | 0.85376816 | 0.8980144395 |
| 16 | Missouri | 0.9050045164 | 0.9337916042 | 0.9022348315 | 0.9320073289 |
| 17 | NewJerse | 0.9111356273 | 0.9378481876 | 0.908606331 | 0.9362433382 |
| 18 | NewYork | 0.7636376949 | 0.8078160421 | 0.7573893975 | 0.8005197149 |
| 19 | Ohio | 0.8010178144 | 0.849549022 | 0.795215466 | 0.8438179328 |
| 20 | Pennsylv | 0.8648740981 | 0.9052355469 | 0.8606506142 | 0.9021007787 |
| 21 | Texas | 0.8217006752 | 0.869420456 | 0.816295656 | 0.8645707216 |
| 22 | Virginia | 0.8732988787 | 0.9106452426 | 0.8693571298 | 0.9077731451 |
| 23 | Washingt | 0.8984320007 | 0.9294176476 | 0.8954080069 | 0.9274350587 |
| 24 | WestVirg | 0.8602609125 | 0.8972449432 | 0.8558896629 | 0.8937211392 |
| 25 | Wisconsi | 0.8727364593 | 0.9103636198 | 0.8687754441 | 0.9074778831 |
The values of each technical efficiency term from the half-normal model are close to the corresponding values from the exponential model.