HPCOUNTREG Procedure

Getting Started: HPCOUNTREG Procedure

(View the complete code for this example.)

Except for its ability to operate in the multithreaded environment, the HPCOUNTREG procedure is similar to other regression model procedures in the SAS System. For example, the following statements are used to estimate a Poisson regression model:

proc hpcountreg data=one ;
   model y = x / dist=poisson ;
run;

The response variable y is numeric and has nonnegative integer values.

This section illustrates two simple examples that use PROC HPCOUNTREG. The data are taken from Long (1997). This study examines how factors such as gender (fem), marital status (mar), number of young children (kid5), prestige of the graduate program (phd), and number of articles published by a scientist’s mentor (ment) affect the number of articles (art) published by the scientist.

The first 10 observations are shown in Figure 1.

Figure 1: Article Count Data

Obsartfemmarkid5phdment
130121.380008.0000
200004.290007.0000
340003.8500047.0000
410113.5900019.0000
510101.810000.0000
610113.590006.0000
700112.1200010.0000
800104.290002.0000
930122.580002.0000
1030111.800004.0000


The following SAS statements estimate the Poisson regression model. The model is executed in single-machine mode with two threads.



/*-- Poisson Regression --*/
proc hpcountreg data=long97data;
   model art = fem mar kid5 phd ment / dist=poisson method=quanew;
   performance nthreads=2  details;
run;

The "Model Fit Summary" table that is shown in Figure 2 lists several details about the model. By default, the HPCOUNTREG procedure uses the Newton-Raphson optimization technique. The maximum log-likelihood value is shown, in addition to two information measures—Akaike’s information criterion (AIC) and Schwarz’s Bayesian information criterion (SBC)—which can be used to compare competing Poisson models. Smaller values of these criteria indicate better models.

Figure 2: Estimation Summary Table for a Poisson Regression

The HPCOUNTREG Procedure

Model Fit Summary
Dependent Variableart
Number of Observations915
Data SetWORK.LONG97DATA
ModelPoisson
Log Likelihood-1651
Maximum Absolute Gradient0.0002080
Number of Iterations13
Optimization MethodQuasi-Newton
AIC3314
SBC3343


Figure 3 shows the parameter estimates of the model and their standard errors. All covariates are significant predictors of the number of articles, except for the prestige of the program (phd), which has a p-value of 0.6271.

Figure 3: Parameter Estimates of Poisson Regression

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValuePr > |t|
Intercept10.30460.10302.960.0031
fem1-0.22460.05461-4.11<.0001
mar10.15520.061372.530.0114
kid51-0.18490.04013-4.61<.0001
phd10.012820.026400.490.6271
ment10.025540.00200612.73<.0001


To allow for variance greater than the mean, you can fit the negative binomial model instead of the Poisson model by specifying the DIST=NEGBIN option, as shown in the following statements. Whereas the Poisson model requires that the conditional mean and conditional variance be equal, the negative binomial model allows for overdispersion, in which the conditional variance can exceed the conditional mean.

/*-- Negative Binomial Regression --*/
proc hpcountreg data=long97data;
   model art = fem mar kid5 phd ment / dist=negbin(p=2) method=quanew;
   performance nthreads=2 details;
run;

Figure 4 shows the fit summary and Figure 5 shows the parameter estimates.

Figure 4: Estimation Summary Table for a Negative Binomial Regression

The HPCOUNTREG Procedure

Model Fit Summary
Dependent Variableart
Number of Observations915
Data SetWORK.LONG97DATA
ModelNegBin
Log Likelihood-1561
Maximum Absolute Gradient0.0000666
Number of Iterations16
Optimization MethodQuasi-Newton
AIC3136
SBC3170


Figure 5: Parameter Estimates of Negative Binomial Regression

Parameter Estimates
ParameterDFEstimateStandard
Error
t ValuePr > |t|
Intercept10.25610.13861.850.0645
fem1-0.21640.07267-2.980.0029
mar10.15050.082111.830.0668
kid51-0.17640.05306-3.320.0009
phd10.015270.036040.420.6718
ment10.029080.0034708.38<.0001
_Alpha10.44160.052978.34<.0001


The parameter estimate for _Alpha of 0.4416 is an estimate of the dispersion parameter in the negative binomial distribution. A t test for the hypothesis upper H 0 colon alpha equals 0 is provided. It is highly significant, indicating overdispersion (p less-than 0.0001).

The null hypothesis upper H 0 colon alpha equals 0 can be also tested against the alternative alpha greater-than 0 by using the likelihood ratio test, as described by Cameron and Trivedi (1998, pp. 45, 77–78). The likelihood ratio test statistic is equal to minus 2 left-parenthesis script upper L Subscript upper P Baseline minus script upper L Subscript upper N upper B Baseline right-parenthesis equals minus 2 left-parenthesis negative 1651 plus 1561 right-parenthesis equals 180, which is highly significant, providing strong evidence of overdispersion.

Last updated: June 19, 2025