The HPFOREST Procedure

References

  • Archer, K. J., and Kimes, R. V. (2008). “Empirical Characterization of Random Forest Variable Importance Measures.” Computational Statistics and Data Analysis 52:2249–2260.

  • Berk, R. A. (2008). Statistical Learning from a Regression Perspective. New York: Springer.

  • Breiman, L. (1996). “Bagging Predictors.” Machine Learning 24:123–140.

  • Breiman, L. (2001). “Random Forests.” Machine Learning 45:5–32.

  • Breiman, L., and Cutler, A. (2003). “Manual—Setting Up, Using, and Understanding Random Forests V4.0.”

  • Breiman, L., Friedman, J., Olshen, R. A., and Stone, C. J. (1984). Classification and Regression Trees. Belmont, CA: Wadsworth.

  • De Ville, B., and Neville, P. G. (2013). Decision Trees for Analytics Using SAS Enterprise Miner. Cary, NC: SAS Institute Inc.

  • Fisher, W. D. (1958). “On Grouping for Maximum Homogeneity.” Journal of the American Statistical Association 53:789–798.

  • Freedman, J. H., and Popescu, B. E. (2003). Importance Sampled Learning Ensembles. Technical report, Department of Statistics, Stanford University.

  • Friedman, J. H. (1977). “A Recursive Partitioning Decision Rule for Nonparametric Classification.” IEEE Transactions on Computers 26:404–408.

  • Friedman, J. H. (1991). “Multivariate Adaptive Regression Splines.” Annals of Statistics 19:1–67.

  • Friedman, J. H. (2001). “Greedy Function Approximation: A Gradient Boosting Machine.” Annals of Statistics 29:1189–1232.

  • Grömping, U. (2009). “Variable Importance Assessment in Regression: Linear Regression versus Random Forest.” American Statistician 63:308–319.

  • Hothorn, T., Hornik, K., and Zeileis, A. (2006). “Unbiased Recursive Partitioning: A Conditional Inference Framework.” Journal of Computational and Graphical Statistics 15:651–674.

  • Kass, G. V. (1980). “An Exploratory Technique for Investigating Large Quantities of Categorical Data.” Journal of the Royal Statistical Society, Series C 29:119–127.

  • King, G., and Zeng, L. (2001). “Logistic Regression in Rare Events Data.” Political Analysis 9:137–163.

  • Lichman, M. (2013). “UCI Machine Learning Repository.” School of Information and Computer Sciences, University of California, Irvine. http://archive.ics.uci.edu/ml.

  • Loh, W.-Y. (2002). “Regression Trees with Unbiased Variable Selection and Interaction Detection.” Statistica Sinica 12:361–386.

  • Loh, W.-Y. (2009). “Improving the Precision of Classification Trees.” Annals of Applied Statistics 3:1710–1737.

  • Loh, W.-Y., and Shih, Y.-S. (1997). “Split Selection Methods for Classification Trees.” Statistica Sinica 7:815–840.

  • Neville, P. G., and Tan, P.-Y. (2014). “A Forest Measure of Variable Importance Resistant to Correlations.” In Proceedings of the 2014 Joint Statistical Meetings. Alexandria, VA: American Statistical Association.

  • Nicodemus, K. K., and Malley, J. D. (2009). “Predictor Correlation Impacts Machine Learning Algorithms: Implications for Genomic Studies.” Bioinformatics 25:1884–1890.

  • Quinlan, J. R. (1993). C4.5: Programs for Machine Learning. San Francisco: Morgan Kaufmann.

  • Radcliffe, N., and Surry, P. (2011). Real-World Uplift Modelling with Significance-Based Uplift Trees. Portrait Technical Report TR-2011-1, Stochastic Solutions. http://www.stochasticsolutions.com/pdf/sig-based-up-trees.pdf.

  • Schapire, R. E., and Freund, Y. (2012). Boosting: Foundations and Algorithms. Cambridge, MA: MIT Press.

  • Smith, J. W., Everhart, J. E., Dickson, W. C., Knowler, W. C., and Johannes, R. S. (1988). “Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus.” In Proceedings of the Symposium on Computer Applications and Medical Care, 261–265. Los Alamitos, CA: IEEE Computer Society Press.

  • Strobl, C., Boulesteix, A.-L., Kneib, T., Augustin, T., and Zeileis, A. (2008). “Conditional Variable Importance for Random Forests.” BMC Bioinformatics 9:307.

  • Su, X., Tsai, C.-L., Wang, H., Nickerson, D. M., and Li, B. (2009). “Subgroup Analysis via Recursive Partitioning.” Journal of Machine Learning Research 10:141–158.

  • Van der Laan, M. J. (2006). “Statistical Inference for Variable Importance.” International Journal of Biostatistics 2:1–31. Article 2.

Last updated: July 02, 2020