The HPF Procedure
ID Statement
ID variable INTERVAL=interval < options >;
The ID statement names a numeric variable that identifies observations in the input and output data sets. The ID variable’s values are assumed to be SAS date, time, or datetime values. In addition, the ID statement specifies the (desired) frequency to be associated with the actual time series. The options in the ID statement also specify how to accumulate the observations and how to align the time ID values to form the actual time series. Consequently, the specified information affects all variables that are specified in subsequent FORECAST statements. If an ID statement is not specified, the observation number, with respect to the BY group, is used as the time ID.
You must specify the following option:
- INTERVAL=interval
-
specifies the frequency of the input time series. For example, if the input data set consists of quarterly observations, then INTERVAL=QTR should be used. If the SEASONALITY= option is not specified, the length of the seasonal cycle is implied by the INTERVAL= option. For example, INTERVAL=QTR implies a seasonal cycle of length 4. If the ACCUMULATE= option is also specified, the INTERVAL= option determines the time periods for the accumulation of observations.
The basic intervals are YEAR, SEMIYEAR, QTR, MONTH, SEMIMONTH, TENDAY, WEEK, WEEKDAY, DAY, HOUR, MINUTE, SECOND. For information about the intervals that can be specified, see Chapter 4, Date Intervals, Formats, and Functions (SAS/ETS User's Guide).
You can also specify the following options:
- ACCUMULATE=option
-
specifies how to accumulate the data set observations within each time period. The frequency (width of each time interval) is specified by the INTERVAL= option. The ID variable contains the time ID values, each of which corresponds to a specific time period. The accumulated values form the actual time series, which is used in subsequent model fitting and forecasting.
The ACCUMULATE= option is particularly useful when there are zero or more than one input observations that coincide with a particular time period (for example, transactional data). The EXPAND procedure offers additional frequency conversions and transformations that can also be useful in creating a time series.
The following options determine how the observations are to be accumulated within each time period based on the ID variable and the frequency specified by the INTERVAL= option. By default, ACCUMULATE=NONE.
- NONE
does not accumulate observations; the ID variable values must be equally spaced with respect to the frequency.
- TOTAL
accumulates observations based on the total sum of their values.
- AVERAGE | AVG
accumulates observations based on the average of their values.
- MINIMUM | MIN
accumulates observations based on the minimum of their values.
- MEDIAN | MED
accumulates observations based on the median of their values.
- MAXIMUM | MAX
accumulates observations based on the maximum of their values.
- N
accumulates observations based on the number of nonmissing observations.
- NMISS
accumulates observations based on the number of missing observations.
- NOBS
accumulates observations based on the number of observations.
- FIRST
accumulates observations based on the first of their values.
- LAST
accumulates observations based on the last of their values.
- STDDEV | STD
accumulates observations based on the standard deviation of their values.
- CSS
accumulates observations based on the corrected sum of squares of their values.
- USS
accumulates observations based on the uncorrected sum of squares of their values.
If the ACCUMULATE= option is specified, the SETMISSING= option is useful for specifying how accumulated missing values are treated. If missing values should be interpreted as zero, then SETMISSING=0 should be used. For more information about accumulation, see the section Details: HPF Procedure.
- ALIGN=option
-
controls the alignment of SAS dates used to identify output observations.
You can specify the following options:
- BEGINNING | BEG | B
represents each time period by using the beginning SAS date or datetime value of the time period.
- ENDING | END | E
represents each time period by using the ending SAS date or datetime value of the time period.
- MIDDLE | MID | M
represents each time period by using the middle SAS date or datetime value of the time period. The middle is calculated as the average of the beginning and ending values.
By default, ALIGN=BEGINNING.
- END=value
specifies a SAS date, datetime, or time value that represents the end of the data. If the last time ID variable’s value is less than value, the series is extended with missing values. If the last time ID variable’s value is greater than value, the series is truncated. For example, END="&sysdate"D uses the automatic macro variable SYSDATE to extend or truncate the series to the current date. This option and the START= option can be used to ensure that data that are associated with each BY group contain the same number of observations.
- FORMAT=format
specifies the SAS format for the time ID values. If this option is not specified, the default format is implied by the INTERVAL= option.
- NOTSORTED
requests that the time ID values not be in sorted order. The HPF procedure sorts the data with respect to the time ID prior to analysis.
- SETMISSING=option | number
-
specifies how to assign missing values (either actual or accumulated) in the accumulated time series. If a number is specified, missing values are set to number. If a missing value indicates an unknown value, do not specify this option. If a missing value indicates no value, specify SETMISSING=0. You would typically use SETMISSING=0 for transactional data because a lack of recorded data usually implies no activity. You can specify the following options.
- MISSING
sets missing values to missing. If the MODEL= option specifies a smoothing model, the missing observations are smoothed over. If MODEL=IDM is specified, missing values are assumed to be periods of no demand—that is, SETMISSING=MISSING is equivalent to SETMISSING=0.
- AVERAGE | AVG
sets missing values to the accumulated average value.
- MINIMUM | MIN
sets missing values to the accumulated minimum value.
- MEDIAN | MED
sets missing values to the accumulated median value.
- MAXIMUM | MAX
sets missing values to the accumulated maximum value.
- FIRST
sets missing values to the accumulated first nonmissing value.
- LAST
sets missing values to the accumulated last nonmissing value.
- PREVIOUS | PREV
sets missing values to the previous accumulated nonmissing value. Missing values at the beginning of the accumulated series remain missing.
- NEXT
sets missing values to the next accumulated nonmissing value. Missing values at the end of the accumulated series remain missing.
By default, SETMISSING=MISSING.
- START=value
specifies a SAS date, datetime, or time value that represents the beginning of the data. If the first time ID variable value is greater than value, missing values are added at the beginning of the series. If the first time ID variable value is less than value, the series is truncated. This option and the END= option can be used to ensure that data that are associated with each BY group contain the same number of observations.
- ZEROMISS=option
-
specifies how to interpret beginning and ending zero values (either actual or accumulated) in the accumulated time series. The following options can also be used to determine how beginning and/or ending zero values are assigned:
- NONE
does not change beginning and ending zeros.
- LEFT
sets beginning zeros to missing.
- RIGHT
sets ending zeros to missing.
- BOTH
sets both beginning and ending zeros to missing.
If the only values in the accumulated series are missing or zero, the series is not changed. By default, ZEROMISS=NONE