Producing odds ratios for logistic GEE (Generalized Estimating Equations) models


In SAS/STAT®, you can fit a logistic GEE (Generalized Estimating Equations) model using the REPEATED statement in the GEE and GENMOD procedure when the response is binary (DIST=BIN option) or multinomial (DIST=MULT option) and the link involves a logit function (LINK=LOGIT or LINK=CUMLOGIT option). In these procedures, you can then estimate odds ratios using the ESTIMATE statement with the EXP option. Odds ratio estimates can also be obtained using the EXP option in the LSMESTIMATE statement, or the DIFF and ODDSRATIO (or EXP) options in the LSMEANS or SLICE statement. To use the LSMEANS, LSMESTIMATE, or SLICE statement, you must use the default parameterization (PARAM=GLM) in the CLASS statement. These statements are described more below and in the GENMOD and GEE documentation.

A binary logistic example using a main effects model is shown below. An example using all of these statements in a binary logistic model with interaction appears in SAS KB0057059. See "Example 3: A Two-Factor Logistic Model with Interaction Using Dummy and Effects Coding." An example in a generalized logit model for a nominal, multinomial response appears in SAS KB0057397.

In SAS® Viya®, you can fit the logistic GEE model using the REPEATED statement in the LOGSELECT or GENSELECT procedures. Estimation of odds ratios is greatly simplified by the ODDSRATIO statement that is available in PROC LOGSELECT. It can provide odds ratio estimates for categorical or continuous predictors even if involved in interactions. It also does not require GLM parameterization of CLASS variables. This statement is not available in PROC GENSELECT.

Available statements for odds ratios in PROC GEE and PROC GENMOD

ESTIMATE Statement: While the ESTIMATE statement is the most flexible way to estimate an odds ratio, it is also the most complex and difficult to use since you must correctly specify the contrast coefficients defining the desired difference of log odds. For this reason, it is advisable to only use the ESTIMATE statement when one of the statements, described below, cannot be used. However, only the ESTIMATE statement can be used to estimate odds ratios for a continuous predictor. To obtain an odds ratio estimate, the ESTIMATE statement should specify coefficients of a linear combination of model parameters that define a difference between two groups. This estimates the difference in log odds (equivalently, the log odds ratio) for the two groups in the logistic model. The process for determining appropriate ESTIMATE statement coefficients is described and illustrated in SAS KB0057059. The same can be done with longitudinal logistic models fit by Generalized Estimating Equations (GEE) using the REPEATED statement.

LSMEANS Statement: With the GEE or GENMOD procedure, the LSMEANS statement is the easiest way to obtain odds ratio estimates if the groups to be compared only involve levels of CLASS variables. The DIFF option provides all pairwise comparisons of the levels of the specified variable or interaction. This removes the need for you to specify a linear combination for the specific comparison desired. In a logistic model, an LS-mean is a log odds. The DIFF option computes differences of the log odds, which are log odds ratios. Adding the ODDSRATIO (or EXP) option exponentiates the log odds ratios resulting in odds ratio estimates. The CL option provides confidence limits. The ILINK option applies the inverse of the logit link to the individual LS-means (log odds) estimates resulting in estimates of the event probabilities.

LSMESTIMATE Statement: The LSMESTIMATE statement can be much easier to use than the ESTIMATE statement since you need only specify a linear combination of the LS-means themselves rather than of the model parameters. This makes specifying the desired comparison much easier particularly when the model involves interactions. Since the LSMEANS statement provides the list of LS-means, you just need to specify a corresponding list of coefficients that defines and compares the two groups of interest. Adding the EXP option exponentiates the log odds differences resulting in odds ratio estimates. The CL option again provides confidence limits.

SLICE Statement: The SLICE statement can be used to estimate odds ratios for CLASS predictors involved in interactions. It essentially provides a specified subset of what the LSMEANS statement can provide. By specifying an interaction effect such as A*B in the SLICE statement along with the SLICEBY=B option, you obtain comparisons among the levels of CLASS predictor A at each level of CLASS predictor B. The DIFF and ODDSRATIO (or EXP) options again provide estimates of differences of log odds and odds ratios, and the CL option provides confidence limits. See SAS KB0057062 for an example of using the SLICE statement to estimate odds ratios in a logistic model with interaction.

Odds ratio estimation in PROC GEE (or GENMOD)

Consider the Generalized Estimating Equations (GEE) example in the "Getting Started" section of the GENMOD documentation. The interpretations below also apply to models that do not use the REPEATED statement. In the MODEL statement, the EVENT="1" response variable option ensures that the probability of wheezing (WHEEZE=1) is modeled. The first ESTIMATE statement below produces the odds ratio estimate and confidence interval for a unit increase in AGE. The ESTIMATE statement must be used since AGE is a continuous predictor. The remaining statements all estimate the CITY odds ratio. The odds ratio comparing the Kingston to Portage cities is provided by the second ESTIMATE statement. The LSMEANS statement provides log odds, odds (EXP option), and probability (ILINK option) estimates for each city and all pairwise comparisons of the cities giving log odds ratio and odds ratio estimates and confidence limits. The LSMESTIMATE statement does the same but only for the specified comparison of cities.

proc gee data=six;
    class case city;
    model wheeze(event="1") = city age smoke / dist=bin;
repeated subject=case / type=exch; estimate "log O.R. Age" age 1 / exp cl; estimate "log O.R. Kingston vs Portage" city 1 -1 / exp cl; lsmeans city / ilink exp diff cl; lsmestimate city 'Kingston vs Portage' 1 -1 / exp cl; run;

The first "Estimate" table contains the results from the first ESTIMATE statement. The estimated change in the odds of wheezing for a one-year increase in AGE is 0.8158 with confidence interval (0.4723, 1.4093).

In the second table from the next ESTIMATE statement, the odds of wheezing for children in Kingston is estimated to be 1.1301 times the odds of wheezing for children in Portage with confidence interval (0.2933, 4.3548).

The first table from the LSMEANS statement produces, for each city, estimates of the log odds of wheezing (Estimate column), the probability of wheezing (Mean column from the ILINK option), and the odds of wheezing (Exponentiated column from the EXP option). The DIFF option produces the table of LS-means differences containing the estimated log odds ratio (Estimate column) comparing the cities. The EXP and CL options add the odds ratio estimate and confidence limits (Exponentiated columns), which match the results from the second ESTIMATE statement above. If you replace the EXP option with the ODDSRATIO option, the same exponentiated differences are provided, but are instead labeled as Odds Ratios in the Differences table.

The LSMESTIMATE statement with the EXP and CL options requests one comparison of cities and produces the following table showing the same result for this comparison as from the DIFF option in the LSMEANS statement.

Odds ratio estimation in PROC LOGSELECT

The logistic GEE model can also be fit in PROC LOGSELECT or PROC GENSELECT in SAS Viya using the REPEATED statement.

After starting a CAS session and creating a CAS libref called SASCAS1, the following statements fit the same logistic GEE model as above. The ODDSRATIO statement then estimates the odds ratio for a 1 unit increase in the continuous Age predictor as well as the odds ratio comparing the cities.

data sascas1.six;
   set six;
   run;
proc logselect data=sascas1.six;
   class case city;
   model wheeze(event="1") = city age smoke;
   repeated subject=case / type=exch;
   oddsratio age city;
   run;

The requested odds ratio estimates are below and agree with the estimates produced by the ESTIMATE and LSMEANS statements above.