The FRONTIER Procedure

Getting Started: FRONTIER Procedure

This section illustrates how to estimate stochastic frontier production models. The following statements fit a half-normal model:

proc frontier data=mycas.mydata;
   model y = x1 x2 / type=half;
run;

The output variable y and the input variables x1 and x2 are contained in the data table mycas.mydata. These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref. The relationship between input and output is specified in the MODEL statement. The distribution of the inefficiency term is specified using the TYPE= option in the MODEL statement. In this example, it is the half-normal distribution. TYPE=EXPONENTIAL specifies an exponential distribution, and TYPE=TRUNCATED specifies a truncated-normal distribution. By default, PROC FRONTIER fits a production model. You can specify the COST option in the MODEL statement to fit a stochastic frontier cost model.

The following example fits the stochastic frontier production model to 1957 state data for the US transportation equipment manufacturing industry. The data were originally published in Zellner and Revankar (1969). The data include observations on aggregate value added (valueadded), aggregate capital service flow (capital), aggregate hours worked (labor), and number of establishments (nfirm) for 25 US states. A Cobb-Douglas production function is estimated for the output variable, log of value added (LNV), using the input variables log of capital (LNK) and log of labor (LNL).

The following statements create the data set:

    title1 'Getting Started Example: Stochastic Frontier Production Model';

    data transpEqp;
       input state$ valueadded capital labor nfirm;
    datalines;
    Alabama       126.14800262 3.8039999008 31.551000595 68
    California    3201.486084  185.44599915 452.84399414 1372
    Connecticut   690.66998291 39.712001801 124.0739975  154
    Florida       56.296001434 6.5469999313 19.180999756 292
    Georgia       304.53100586 11.529999733 45.534000397 71
    Illinois      723.02801514 58.986999512 88.39099884  275
    Indiana       992.16900635 112.88400269 148.52999878 260
    Iowa          35.796001434 2.6979999542 8.0170001984 75
    Kansas        494.51501465 10.359999657 86.189002991 76
    Kentucky      124.94799805 5.2129998207 12           31
    Louisiana     73.32800293  3.7630000114 15.899999619 115
    Maine         29.466999054 1.9670000076 6.4699997902 81
    Maryland      415.26199341 17.545999527 69.342002869 129
    Massachusetts 241.52999878 15.347000122 39.416000366 172
    Michigan      4079.5539551 435.10501099 490.38400269 568
    Missouri      652.08502197 32.840000153 84.831001282 125
    NewJersey     667.11297607 33.291999817 83.032997131 247
    NewYork       940.42999268 72.973999023 190.09399414 461
    Ohio          1611.8990479 157.97799683 259.91598511 363
    Pennsylvania  617.57897949 34.324001312 98.152000427 233
    Texas         527.4130249  22.736000061 109.72799683 308
    Virginia      174.39399719 7.1729998589 31.301000595 85
    Washington    636.94799805 30.806999207 87.962997437 179
    WestVirginia  22.700000763 1.5429999828 4.0630002022 15
    Wisconsin     349.71099854 22.000999451 52.818000793 142
    ;
 data transpEqp;
    set transpEqp;
       LNV = log(valueadded);
       LNK = log(capital);
       LNL = log(labor);

 data mycas.transpEqp;
    set transpEqp;
 run;

The following statements estimate a stochastic frontier half-normal production model that uses the transportation equipment manufacturing industry data:

 /*-- Stochastic Frontier Production Model: Half-Normal --*/
 proc frontier data=mycas.transpEqp;
    model LNV = LNK LNL / type=half;
 run;

The "Observation Information" and "Summary Statistics of Dependent Variable" tables shown in Figure 1 contain some details about the observations in the data set and the dependent variable.

Figure 1: Information about Observations and the Dependent Variable

Getting Started Example: Stochastic Frontier Production Model

The FRONTIER Procedure

Observation Information
Number of Observations25
Number of Missing Observations0

Summary Statistics of Dependent Variable
VariableMeanStandard
Error
MinimumMaximum
LNV5.8120921.3753043.1223658.313743


The "Model Fit Summary" table shown in Figure 2 lists several details about the model. By default, PROC FRONTIER uses the Newton-Raphson optimization technique, and the estimated covariance matrix is calculated as the inverse of the negative Hessian. For the available optimization and estimation techniques, see the section PROC FRONTIER Statement. The maximum log-likelihood value is shown, in addition to two information measures—Akaike’s information criterion (AIC) and Schwarz’s Bayesian information criterion (SBC)—that can be used to compare competing production models. Smaller values of these criteria indicate better models.

Figure 2: Estimation Summary Table for Half-Normal Production Model

Model Fit Summary
Dependent VariableLNV
Data SetTRANSPEQP
ModelProduction
Inefficiency Term DistributionHalf-normal
Log Likelihood2.469522
Maximum Absolute Gradient3.942E-6
Number of Iterations11
Optimization MethodNewton-Raphson
AIC5.060956
SBC11.15533
Covariance EstimationHessian


Figure 3 shows the parameter estimates of the model and their standard errors. Both input variables are significant predictors of the log of value added. _Sigma_v is the estimate of the standard deviation of the idiosyncratic component of the error term that is specific to the firm and that could enter the model with either sign. _Sigma_u is the estimate of the standard deviation of the inefficiency component. When the data are in log terms, the inefficiency component becomes a measure of the percentage by which a particular observation fails to achieve an ideal production rate.

Figure 3: Parameter Estimates of Half-Normal Production Model

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Intercept12.0811340.2816487.39<.0001
LNK10.2585480.0987662.620.0089
LNL10.7802450.1199456.51<.0001
_Sigma_v10.1751690.0542653.230.0012
_Sigma_u10.2215070.1237061.790.0734


Figure 4 shows the estimates of two important variance statistics that are helpful for model diagnosis. Sigma2 is the estimate of the total error variance. It is calculated as the sum of the idiosyncratic error variance and the variance of the technical inefficiency term. Gamma is the estimate of how much the technical inefficiency contributes to the total variation in the model. It is calculated as the ratio of the variance of the technical inefficiency term to the total error variance. In this example, the technical inefficiency explains most of the variation in the model.

Figure 4: Estimates of Important Variance Statistics for Half-Normal Production Model

Variance Statistics
ParameterEstimateStandard
Error
Sigma20.0797500.042699
Gamma0.6152440.385748


The following statements estimate a stochastic frontier exponential production model that uses the same data:

 /*-- Stochastic Frontier Production Model: Exponential --*/
 proc frontier data=mycas.transpEqp;
    model LNV = LNK LNL;
    output out=mycas.out_te_exp  te1=TE1_exp te2=TE2_exp copyvars=( state );
 run;

The default model is the exponential model if you do not specify the TYPE= option in the MODEL statement.

Figure 5 shows some of the results from this production model.

Figure 5: Estimation Summary and Parameter Estimates for Exponential Production Model

Getting Started Example: Stochastic Frontier Production Model

The FRONTIER Procedure

Model Fit Summary
Dependent VariableLNV
Data SetTRANSPEQP
ModelProduction
Inefficiency Term DistributionExponential
Log Likelihood2.86049
Maximum Absolute Gradient4.215E-6
Number of Iterations11
Optimization MethodNewton-Raphson
AIC4.279021
SBC10.3734
Covariance EstimationHessian

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValueApprox
Pr > |t|
Intercept12.0692420.2356328.78<.0001
LNK10.2624860.0920012.850.0043
LNL10.7703800.1109646.94<.0001
_Sigma_v10.1713920.0384494.46<.0001
_Sigma_u10.1351690.0626852.160.0311

Variance Statistics
ParameterEstimateStandard
Error
Sigma20.0476460.015793
Gamma0.3834670.285232


You can obtain estimates of the efficiency terms by using the OUTPUT statement. PROC FRONTIER offers two types of technical efficiency terms that are proposed by Battese and Coelli (1988) and Jondrow et al. (1982). The OUTPUT statement in the previous and following statements specify the two technical efficiency terms TE1 and TE2. These terms for each observation are written in the data set that you specify in the OUT= option. In these examples, the technical efficiency terms are calculated for each state and included in mycas.out_te_exp for the exponential model and mycas.out_te_half for the half-normal model.

 /*-- Obtaining Technical Efficiency Terms for Half-Normal Model --*/
 proc frontier data=mycas.transpEqp;
    model LNV = LNK LNL / type=half;
    output out=mycas.out_te_half  te1=TE1_half te2=TE2_half copyvars=( state );
 run;

The following statements merge the two output data sets and sort the new data set with respect to the state variable:

 /*-- Merging the Two Output Data Sets --*/
 data mycas.out_te;
    merge mycas.out_te_half mycas.out_te_exp;
    by state;
 run;

 proc sort data=mycas.out_te  out=out_te_sorted;
    by state;
 run;
 proc print data = out_te_sorted;
    var State TE1_half TE1_exp TE2_half TE2_exp;
    title 'TEs from Half-Normal and Exponential Models';
 run;

These statistics for each state are shown in Figure 6.

Figure 6: Technical Efficiency for Each State

TEs from Half-Normal and Exponential Models

ObsstateTE1_halfTE1_expTE2_halfTE2_exp
1Alabama0.82317551240.86908457040.81780307370.8642193711
2Californ0.86926547210.91025074680.86518698180.9073595422
3Connecti0.83184066680.87833469760.826671060.8739011963
4Florida0.60163411410.56230707630.59599022770.5541443131
5Georgia0.9040509620.93290425310.9012441310.9310801261
6Illinois0.88917126930.92261218910.88579749590.9203129631
7Indiana0.81508985680.86202395320.80954576080.8568378457
8Iowa0.78535216230.8248431130.77932172450.8181327187
9Kansas0.90664307610.93543330850.90393739080.9337222102
10Kentucky0.94644813040.9600591950.94528464590.9593259036
11Louisian0.82142231880.86518031510.8160112140.8601366506
12Maine0.80612432790.84741494870.80040977880.8415937609
13Maryland0.87722394190.9143818890.87341838760.9116902323
14Massachu0.85963374040.8998330780.8552427720.8964352583
15Michigan0.85820375130.90133892340.853768160.8980144395
16Missouri0.90500451640.93379160420.90223483150.9320073289
17NewJerse0.91113562730.93784818760.9086063310.9362433382
18NewYork0.76363769490.80781604210.75738939750.8005197149
19Ohio0.80101781440.8495490220.7952154660.8438179328
20Pennsylv0.86487409810.90523554690.86065061420.9021007787
21Texas0.82170067520.8694204560.8162956560.8645707216
22Virginia0.87329887870.91064524260.86935712980.9077731451
23Washingt0.89843200070.92941764760.89540800690.9274350587
24WestVirg0.86026091250.89724494320.85588966290.8937211392
25Wisconsi0.87273645930.91036361980.86877544410.9074778831


The values of each technical efficiency term from the half-normal model are close to the corresponding values from the exponential model.

Last updated: January 27, 2023