The SPATIALREG Procedure
Getting Started: SPATIALREG Procedure
(View the complete code for this example.)
The SPATIALREG procedure is similar to other SAS regression model procedures, except that you usually need to provide a secondary data set (in the WMAT= option). The spatial weights matrix defines all pairwise spatial relationships and is the most vital component of a spatial regression model. For more information about how to create spatial weights matrix, see the section Specifying the Spatial Weights Matrix.
The following statements fit a SAR model:
proc spatialreg data=one Wmat=W;
model y = x1 x2 / type=SAR;
run;
The response variable y is continuous, and the data set W, which you specify in the WMAT= option, contains the spatial relationships among all spatial units in the data. In this case, W is either contiguity or weights. You specify the TYPE=SAR option to request a SAR model.
The following example illustrates PROC SPATIALREG by using a real-world data set. The data set CRIMEOH is taken from Anselin (1988) and can be found in the SAS/ETS Sample Library. This data set contains variables such as INCOME (household income, measured in $1000), HVALUE (housing value by $1000), and CRIME (number of crimes, including residential burglaries and vehicle thefts, measured per 1,000 households) in 49 neighborhoods in Columbus, Ohio. You want to examine how household income and housing value affect the number of crimes in the 49 neighborhoods of interest.
The first 10 observations in the CRIMEOH data set are shown in Figure 1.
Figure 1: Columbus Crime Data
| Obs | crime | income | hvalue | lat | lon |
|---|---|---|---|---|---|
| 1 | 18.802 | 21.232 | 44.567 | 35.62 | 42.38 |
| 2 | 32.388 | 4.477 | 33.200 | 36.50 | 40.52 |
| 3 | 38.426 | 11.337 | 37.125 | 36.71 | 38.71 |
| 4 | 0.178 | 8.438 | 75.000 | 33.36 | 38.41 |
| 5 | 15.726 | 19.531 | 80.467 | 38.80 | 44.07 |
| 6 | 30.627 | 15.956 | 26.350 | 39.82 | 41.18 |
| 7 | 50.732 | 11.252 | 23.225 | 40.01 | 38.00 |
| 8 | 26.067 | 16.029 | 28.750 | 43.75 | 39.28 |
| 9 | 48.585 | 9.873 | 18.000 | 39.61 | 34.91 |
| 10 | 34.001 | 13.598 | 96.400 | 47.61 | 36.42 |
The following SAS statements fit a linear regression model to the CRIMEOH data set:
proc spatialreg data=crimeoh;
model crime = income hvalue / type=LINEAR;
run;
The "Model Fit Summary" table, shown in Figure 2, lists several fit summary statistics about the model. By default, the SPATIALREG procedure uses the Newton-Raphson optimization technique. The maximum log-likelihood value is shown, in addition to two information measures, Akaike’s information criterion (AIC) and Schwarz’s Bayesian information criterion (SBC). AIC or SBC can be used for model selection. For a set of candidate models, the model with the smallest AIC or SBC is often preferred.
Figure 2: Fit Summary Statistics for a Linear Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | crime |
| Number of Observations | 49 |
| Data Set | WORK.CRIMEOH |
| Model | Linear |
| Log Likelihood | -187.37709 |
| Maximum Absolute Gradient | 7.59852E-7 |
| Number of Iterations | 16 |
| Optimization Method | Newton-Raphson |
| AIC | 382.75418 |
| SBC | 390.32146 |
The parameter estimates of the model and their standard errors are shown in Figure 3. Based on the p-values, both INCOME and HVALUE are significant at the 0.05 level.
Figure 3: Parameter Estimates of the Linear Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 68.618863 | 4.588210 | 14.96 | <.0001 |
| income | 1 | -1.597304 | 0.323739 | -4.93 | <.0001 |
| hvalue | 1 | -0.273931 | 0.099989 | -2.74 | 0.0062 |
| _sigma2 | 1 | 122.751696 | 24.799493 | 4.95 | <.0001 |
The following statements fit a SAR model to the CRIMEOH data set:
proc spatialreg data=crimeoh Wmat=crimeWmat NONORMALIZE;
model crime = income hvalue / type=SAR;
run;
The NONORMALIZE option requests that the spatial weights matrix that is specified in the CRIMEWMAT data set be used "as is" rather than be row-standardized. The "Model Fit Summary" table, shown in Figure 4, lists several fit summary statistics about the SAR model. For this model, the value of AIC is about 374.78—smaller than 382.75, which is the AIC value for the preceding linear model. Based on AIC, the SAR model is preferred.
Figure 4: Fit Summary Statistics for a SAR Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | crime |
| Number of Observations | 49 |
| Data Set | WORK.CRIMEOH |
| Spatial Weights | WORK.CRIMEWMAT |
| Model | SAR |
| Log Likelihood | -182.38860 |
| Maximum Absolute Gradient | 2.7871E-7 |
| Number of Iterations | 16 |
| Optimization Method | Newton-Raphson |
| AIC | 374.77720 |
| SBC | 384.23630 |
The parameter estimates of the SAR model and their standard errors are shown in Figure 5. According to the p-values, both INCOME and HVALUE are significant at the 0.05 level. In addition, the spatial autoregressive coefficient is estimated to be about 0.431, with a p-value of 0.0005.
Figure 5: Parameter Estimates of the SAR Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 45.077070 | 7.870590 | 5.73 | <.0001 |
| income | 1 | -1.031531 | 0.328403 | -3.14 | 0.0017 |
| hvalue | 1 | -0.265924 | 0.088218 | -3.01 | 0.0026 |
| _rho | 1 | 0.431020 | 0.123594 | 3.49 | 0.0005 |
| _sigma2 | 1 | 95.487066 | 19.506312 | 4.90 | <.0001 |
The following statements fit an SDM model. Unlike the previous SAR model, SDM accounts for exogenous interaction effects by introducing spatial lags of two explanatory variables—INCOME and HVALUE.
proc spatialreg data=crimeoh Wmat=crimeWmat NONORMALIZE;
model crime = income hvalue / type=SAR;
spatialeffects income hvalue;
run;
The fit summary statistics for the SDM model are shown in Figure 6. Parameter estimates are provided in Figure 7.
Figure 6: Fit Summary Statistics for the SDM Model
| Model Fit Summary | |
|---|---|
| Dependent Variable | crime |
| Number of Observations | 49 |
| Data Set | WORK.CRIMEOH |
| Spatial Weights | WORK.CRIMEWMAT |
| Model | SDM |
| Log Likelihood | -181.39141 |
| Maximum Absolute Gradient | 5.44802E-8 |
| Number of Iterations | 16 |
| Optimization Method | Newton-Raphson |
| AIC | 376.78282 |
| SBC | 390.02556 |
Figure 7: Parameter Estimates for the SDM Model
| Parameter Estimates | |||||
|---|---|---|---|---|---|
| Parameter | DF | Estimate | Standard Error | t Value | Approx Pr > |t| |
| Intercept | 1 | 42.803457 | 13.924487 | 3.07 | 0.0021 |
| income | 1 | -0.914206 | 0.336439 | -2.72 | 0.0066 |
| hvalue | 1 | -0.293745 | 0.088857 | -3.31 | 0.0009 |
| W_income | 1 | -0.519640 | 0.594772 | -0.87 | 0.3823 |
| W_hvalue | 1 | 0.245716 | 0.176854 | 1.39 | 0.1647 |
| _rho | 1 | 0.426492 | 0.167492 | 2.55 | 0.0109 |
| _sigma2 | 1 | 91.779519 | 18.909222 | 4.85 | <.0001 |
The spatial autoregressive coefficient is estimated to be 0.426 with a p-value of 0.0109 based on an asymptotic t test. This result seems to suggest that there is a significantly positive spatial dependence in the number of crimes.
In the SPATIALREG procedure, the null hypothesis can also be tested against the alternative
by using the likelihood ratio (LR) test, Lagrange multiplier (LM) test, and Wald test. For the LR test, the test statistic is equal to
, where
and
are the log likelihoods for the linear regression model and SAR model, respectively. The likelihood ratio test is significant at the 0.05 level, providing strong evidence of spatial dependence in the data.