Scoring (computing predicted values) for new observations or a validation data set


Contents: Scoring Methods and Examples

1. Use the STORE Statement and ...
PROC PLM (SAS® 9)
RESTORE= in modeling procedure or PROC ASTORE (SAS® Viya®)
2. Use Built-In Scoring Capabilities
PROC SCORE
SCORE Statement
CODE Statement
3. Augment the Training Data Set
Example 1: Logistic Model Validation Using PROC GENMOD
4. Use the Saved Parameter Estimates to Score Generalized Linear Models
Example 2: A Poisson Model with Offset
Example 3: A Probit Model
Example 4: Scoring a model containing spline effects

Four ways to score (compute predicted values for) new observations using a previously fitted model are discussed below. Note that several conditions can make it impossible to score a new observation, resulting in a missing predicted value. These conditions are described in SAS KB0057141.

1. Use the STORE Statement

Many modeling procedures provide a STORE statement to save the fitted model. In SAS®9, you can then use the SCORE statement in PROC PLM to score a data set using the saved model. In SAS® Viya®, you can score additional data using the SCORE statement in PROC ASTORE. In some procedures, such as LOGSELECT and GENSELECT, you can also use the RESTORE= option along with an OUTPUT statement. For more on the STORE statement, see "STORE statement" in the Shared Concepts and Topics chapter in the SAS/STAT® documentation or in the Shared Concepts chapter in the SAS® Visual Statistics documentation.

PROC PLM (SAS9)

See the example titled "Scoring with PROC PLM" in the Examples section of the PROC PLM documentation and the example below.

The following uses the remission data in the example titled "Stepwise Logistic Regression and Predicted Values" in the LOGISTIC procedure documentation. The PROC GENMOD step below fits a binary logistic model to the data and the STORE statement saves the fitted model in an item store named LOGMOD. Two data sets, New1 and New2, are scored using SCORE statements in PROC PLM. The RESTORE= option reads the model saved by the STORE statement in GENMOD. The PRED=, LCLM=, and UCLM= options in the SCORE statements request the predicted values and their confidence limits. The ILINK option applies the inverse of the link function (logit, in this case) to obtain estimates on the mean (probability) scale. The scored data are saved in the Preds1 and Preds2 data sets specified in the OUT= options.

proc genmod data=remiss descending;
   model remiss = smear blast / dist=binomial;
   store out=logmod;
   run;
proc plm restore=logmod;
   score data=New1 out=Preds1 pred=pred lclm=lower uclm=upper / ilink;
   run;
   score data=New2 out=Preds2 pred=pred lclm=lower uclm=upper / ilink;
   run;

RESTORE= in modeling procedure or PROC ASTORE (SAS Viya)

In SAS Viya, many procedures offer a STORE statement to save the fitted model in an analytic store. You can score new observations using either the SCORE statement in PROC ASTORE or, in some procedures, using the RESTORE= option. Both are illustrated below.

After starting a CAS session and creating a CAS libref called SASCAS1, the following statements fit a logistic model to the remission data using PROC LOGSELECT, save the predicted values, and save the fitted model in an analytic store named LogStor. PROC LOGSELECT is run a second time using the RESTORE= option to read the analytic store. A subset of the remission data, SUB, is scored by specifying it in the DATA= option and including an OUTPUT statement with the PRED= option requesting predicted values. The subset is also scored using the SCORE statement in PROC ASTORE. You can verify that the predicted values for the subset (OUT2 and OUT3) match those produced when fitting the model (OUT1). Note that the SCORE statement in PROC ASTORE provides the predicted probabilities of both response levels (P_remiss1 and P_remiss0) and the predicted response level (I_remiss) according to the level with the larger predicted probability.

data sascas1.remiss;
   set remiss;
   run;
data sascas1.sub;
   set remiss(obs=10);
   run;
proc logselect data=sascas1.remiss;
   model remiss(event='1') = smear blast;
   output out=sascas1.out1 p=p;
   store sascas1.LogStor;
   run;
proc logselect restore=sascas1.LogStor data=sascas1.sub;
   output out=sascas1.out2 pred=p;
   run;
proc astore;
   score data=sascas1.sub out=sascas1.out3
         rstore=sascas1.LogStor;
   run;

2. Use Built-In Scoring Capabilities

Some procedures include features that make scoring new observations easier:

PROC SCORE

For ordinary regression models fit using PROC REG, you can use PROC SCORE to compute predicted values for new observations. See the example titled "Regression Parameter Estimates" in the SCORE documentation. It is not necessary to refit the model. However, PROC SCORE does not directly provide scoring for other types of models such as logistic or other generalized linear models. It also does not provide standard error estimates or confidence limits.

SCORE Statement

For a logistic or probit model, the scoring process is greatly simplified in PROC LOGISTIC. Its SCORE statement enables you to score a data set of new observations. The FITSTAT and OUTROC= options in the SCORE statement enable you to evaluate the model applied to the new data set. The FITSTAT option provides fit statistics such as the area under the ROC curve (AUC) and R-square. The OUTROC= option produces a data set for plotting the ROC curve. An ROC plot and analysis for validation data can be obtained as described in this note. As with ordinary regression models, refitting the model is not necessary if the model is saved using the OUTMODEL= option and then retrieved during scoring by the INMODEL= option. The example titled "Scoring Data Sets with the SCORE Statement" in the LOGISTIC documentation illustrates the use of the SCORE statement with a nominal logistic model.

A SCORE statement is also available in several other modeling procedures such as GLMSELECT, GAM, LOESS, TPSPLINE, and ADAPTIVEREG. See the procedure documentation for discussion and examples.

PROC DISCRIM provides a TESTDATA= option that enables you to specify a data set to be scored, and a TESTOUT= option that includes posterior probabilities and predicted classifications. See the example titled "Linear Discriminant Analysis of Remote-Sensing Data on Crops" in the DISCRIM documentation.

CODE Statement

The CODE statement is available in several modeling procedures. The CODE statement generates SAS code that can be used in a DATA step to score a data set. For more on the CODE statement, see "CODE statement" in the Shared Concepts and Topics chapter in the SAS/STAT documentation or in the Shared Concepts chapter in the SAS® Visual Statistics documentation.

3. Augment the Training Data Set

You can get predicted values for one or more settings of your model predictors by adding observations to the input data that you use to fit (train) the model. The predictors in these new observations should be set to the values for which you want predicted values. For the added observations, either the response variable should be set to missing, or if the new observations have observed values then a WEIGHT variable should be created with value 1 for the training observations and value 0 (or missing) for the new observations.

With these new observations appended to your training data set, the fitted model should be identical to the model fit using only the training data. This is because any observation that has a missing response value or zero (or missing) weight is ignored when fitting the model. (The exception to this is when the model includes spline effects defined in the EFFECT statement. See the Extrapolation section of this note for details.) The procedure can compute predicted values for such observations as long as they have nonmissing values for all of the model predictors and have values for CLASS predictors that existed in the training data set. This is further explained and illustrated in this note. In many procedures, you can request predicted values by specifying the P= option in the OUTPUT statement, but some procedures use other syntax. See the procedure's documentation.

Example 1: Logistic Model Validation Using PROC GENMOD

Model validation often involves getting predictions for a potentially large number of observations that were held out from the original data. That is, the original data set is split into a data set to train the model and a data set to validate the model. Validation is done by comparing the values predicted under the model to the observed values in the validation data set (often called a hold-out data set). One way this can be done is by concatenating the training and validation data sets and using the combined data set as the input data set to the modeling procedure. It is often convenient for the output data set to contain only the validation observations, excluding the observations used to train the model. To do this, add a variable to the combined data set that indicates which observations are the training data set and which observations are the validation data set. You can use this indicator variable in a WHERE= data set option in the OUTPUT statement to select only the validation observations for output.

The following uses the remission data in the example titled "Stepwise Logistic Regression and Predicted Values" in the LOGISTIC procedure documentation. This DATA step creates a validation data set, NEW. For purposes of illustration, the first eight observations of the training data set are used.

data new;
   set remiss(obs=8);
   run;

The following DATA step concatenates the training and validation data sets into a single data set, BOTH, for input to PROC GENMOD. The IN= option in the SET statement creates a temporary variable, InNew, which equals 1 when the observation comes from the validation data set (NEW) and equals 0 when it comes from the training data set (REMISS). The inverse of this variable, W, is created for use as a weight variable. W equals 1 for the training observations, 0 for the validation observations. Data set BOTH is then displayed by PROC PRINT (not shown).

data both;
   set remiss new (in=InNew);
   w=not(InNew);
   run;
proc print noobs; 
   var remiss w smear blast;
   run;

These statements fit the model using the combined data set, BOTH. The training indicator variable, W, is specified in the WEIGHT statement. The results are identical to a GENMOD analysis on just the training data set because observations in the validation data set have zero weight and are ignored in the model fitting process. The OUTPUT statement produces a data set, PREDS, containing predicted values. The WHERE clause after the OUT= data set name causes only those observations from the validation data set to be written to the data set. The L= and U= options request that 95% confidence limits be computed and output in addition to the predicted values requested by the P= option. Data set PREDS is then displayed by PROC PRINT (not shown).

proc genmod data=both descending;
   weight w;
   model remiss = smear blast / dist=binomial;
   output out=preds(where=(w=0)) p=pred l=lower u=upper;
   run;
proc print data=preds noobs; 
   var remiss pred lower upper smear blast;
   run;

4. Use the Saved Parameter Estimates to Score Generalized Linear Models

Some important issues must be remembered in order to correctly and accurately compute predicted values:

Note that if the value of a CLASS variable in an observation to be scored does not appear in the training data set, then that observation cannot be scored. This is because, unlike with a continuous predictor, there is no parameter corresponding to that value in the trained model. For more, see SAS KB0057141 and Example 3 below.

Predicted values can be obtained by this method, but the computations for the standard errors of the predicted values are generally more complex and cannot be computed. As a result, confidence limits for the predicted values also cannot be computed.

Example 2: A Poisson Model with Offset

The following Poisson model is based on the data in the "Getting Started" section of the GENMOD documentation. Note that the model includes a continuous variable (age), a CLASS variable (car), and their interaction. The CLASS variable uses GENMOD's default coding method (PARAM=GLM). The model also includes an offset variable (ln). The following statements create the training data set and fit the desired model. The XVARS and P options in the MODEL statement display the predictor values and the predicted counts (Pred) for the observations in the training data set, shown below. In this example, the specified model happens to be a saturated model, so the predicted values equal the actual values. But this has no influence on the manner of scoring.

data insure;
   input n c car $ age;
   ln = log(n);
   datalines;
   500   42  small  1
   1200  37  medium 1
   100    1  large  1
   400  101  small  2
   500   73  medium 2
   300   14  large  2
    ;
proc genmod data=insure;
   class car;
   model c = car age car*age / dist=poisson link=log offset=ln xvars p;
   ods output parameterestimates=pe;
   run;

Notice that the parameter estimates table was saved via the ODS OUTPUT statement. The variable containing the parameter estimates (Estimate) is displayed with high precision by the FORMAT statement in the following PROC PRINT step. The 12.10 format displays the estimates in a field 12 digits wide and with 10 decimal places. These more precise values are used in the scoring computations below to more closely match what GENMOD does internally with full precision values.

proc print data=pe; 
   id parameter level1;
   format estimate 12.10; 
   var estimate; 
   run;

The following step does the scoring. In this example, the training data set is scored, so it is specified in the SET statement. A SELECT group should appear for each predictor in the CLASS statement to create the appropriately coded design variables. Since the PARAM= option was not specified in the CLASS statement, the default GLM coding is used. If a different coding method is requested via the PARAM= option in the CLASS statement, the coding of the design variables (named carlarge, carmedium, and carsmall in this example) would change. This is discussed further below. The parameter estimates from the preceding PROC PRINT step are used in the computation of the linear predictor, xβ. By definition, the parameter associated with an offset variable equals 1. xβ is computed by multiplying parameter estimates by predictor (or design) variables and adding the products. Finally, the inverse link function is applied to get a predicted mean. Since this Poisson model uses the log link, the inverse link function is exponentiation that can be done with the EXP function. For ordinary regression models, such as those fit by PROC REG or PROC GLM, the link is the identity link and xβ is the predicted mean.

data scores;
  set insure;
  select (car);
     when ("large") do;
        carlarge=1; carmedium=0; carsmall=0; end;
     when ("medium") do;
        carlarge=0; carmedium=1; carsmall=0; end;
     when ("small") do;
        carlarge=0; carmedium=0; carsmall=1; end;
     otherwise;
     end;
  xbeta=-3.577532930 +
        -2.568082297*carlarge +
        -1.456636259*carmedium +
        0*carsmall +
        1.1005944499*age +
        0.4398505911*age*carlarge +
        0.4544158160*age*carmedium +
        0*age*carsmall +
        1*ln
        ;
  mu_hat=exp(xbeta);
  run;
proc print noobs;
  var car age c xbeta mu_hat;
  run; 

Notice that the computed scores (mu_hat) match the predicted values computed by the P option (Pred) in PROC GENMOD.

Had effects coding (PARAM=EFFECT) been specified, the following SELECT group would properly code the design variables for use in scoring:

select (car);
   when ("large") do;
     carlarge=1;  carmedium=0;  end;
   when ("medium") do;
     carlarge=0;  carmedium=1;  end;
   when ("small") do;
     carlarge=-1; carmedium=-1; end;
   otherwise;
   end; 

For reference coding (PARAM=REF), this SELECT group would be used:

select (car);
   when ("large") do;
     carlarge=1; carmedium=0; end;
   when ("medium") do;
     carlarge=0; carmedium=1; end;
   when ("small") do;
     carlarge=0; carmedium=0; end;
   otherwise;
   end;

Example 3: A Probit Model

The following uses data from the example titled "Logistic Modeling with Categorical Predictors" in the LOGISTIC procedure documentation. PROC GENMOD is used to fit a probit model to the data to model the probability of no pain as specified by the response variable option EVENT="No". Effects coding is used for the categorical predictor Treatment (A, B, or P) and reference coding is used for the Sex (F or M) predictor with males (M) as the reference category. The P= option in the OUTPUT statement produces predicted values and saves them in data set PREDS. The ODS OUTPUT statement saves the parameter estimates in data set PARMS.

proc genmod data=Neuralgia;
  class Treatment (param=effect) Sex (param=ref ref="M");
  model Pain(event="No") = Treatment Sex Treatment*Sex Age Duration / 
                           dist=binomial link=probit;
  output out=preds p=PrNoPain;
  ods output parameterestimates=parms;
  run;
proc print data=preds(obs=6) noobs;
  run; 
proc print data=parms;
  id parameter level:;
  format estimate 12.10; 
  var estimate; 
  run;

The Class Level Information and Parameter Estimates tables produced by GENMOD are shown below. Note the coding used for the two CLASS predictors. Following those tables are the scores (predicted probabilities of no pain) for the first six observations. Finally, the parameter estimates are displayed using higher precision.

To illustrate scoring, the first six observations of the training data set are used as a validation data set. The scores for these observations should match the predicted values shown above. Two additional observations are included — one with an invalid Treatment code (X) and one with a missing value for Sex. The first observation cannot be scored because there is no parameter for Treatment X in the model. In order to score this observation, the training data would need to include some subjects who were given Treatment X. The second observation cannot be scored since values for all predictors in the model must be nonmissing in order to make a valid computation. See this note for more discussion.

data valid;
  input Treatment $ Sex $ Age Duration Pain $ @@;
  datalines;
P  F  68   1  No   B  M  74  16  No  P  F  67  30  No
P  M  66  26  Yes  B  F  67  28  No  B  F  77  16  No
X  F  50  10  .    B  .  32  15  .
;
proc print noobs;
   run;

In the scoring step below, a SELECT group is included for each of the two categorical predictors, Treatment and Sex, using coding that matches the coding used when training the model — effects coding for Treatment and reference coding for Sex. The "Class Level Information" table (above) produced by PROC GENMOD shows you how the design variables are coded. xβ is computed using the high precision parameter estimates displayed above. Since the inverse of the probit link function is the probability from the standard normal distribution, you can use the PROBNORM function in SAS. Had the logit link been used to produce a logistic model, you would use the inverse logit function, 1/(1+exp(-xβ)), which can be computed using the LOGISTIC function: logistic(xbeta).

data scores;
  set valid;
  select (Treatment);
     when ("A") do;
        TrtA=1; TrtB=0; end;
     when ("B") do;
        TrtA=0; TrtB=1; end;
     when ("P") do;
        TrtA=-1; TrtB=-1; end;
     otherwise;
     end;
  select (Sex);
     when ("F") SexF=1;
     when ("M") SexF=0;
     otherwise;
     end;
  xbeta=9.9221347360 +
        0.5139396491*TrtA +
        0.7200040408*TrtB +
        0.9404279566*SexF +
        -.1621550070*TrtA*SexF +
        0.1580056737*TrtB*SexF +
        -.1439733988*age +
        0.0006380643*duration
        ;
  PrNoPain=probnorm(xbeta);
  run;      
proc print noobs;
  var treatment sex age duration pain xbeta PrNoPain;
  run;

Notice that the predicted probabilities for the first six observations match those computed by PROC GENMOD above, and the predicted probabilities for the last two observations are missing as expected.

PROC LOGISTIC can also fit a probit model and provides effects and reference coding. Since it has built-in scoring capability via its SCORE statement, you can fit the model and score the validation data all in a single step. Any slight differences are due to minor differences in starting values and iteration methods used by GENMOD and LOGISTIC.

proc logistic data=Neuralgia;
  class Treatment (param=effect) Sex (param=ref ref="M");
  model Pain = Treatment Sex Treatment*Sex Age Duration / link=probit;
  score data=valid out=validscore;
  run;
proc print data=validscore noobs;
  var treatment sex age duration pain P_No;
  run; 

Example 4: Scoring a model containing spline effects

See this example that discusses the types of spline transformations available in the EFFECT statement and illustrates reproducing the spline basis functions and scoring data.