The input data for the MDC procedure should have a separate observation for each choice presented to an individual. PROC MDC allows different choice sets for different individuals. Therefore, the total number of observations in the data set must equal the sum (over all individuals in the study) of the numbers of choices presented to all individuals.
If your data set is structured such that choice-specific information is stored in multiple variables and the total number of observations equals to the number of individuals, then you can either use DATA step code or, beginning in SAS 9.2, the MDCDATA statement to transform your data to match the required format. Both are illustrated in the example titled "Conditional Logit and Data Conversion" in the PROC MDC documentation. The example titled "Conditional Logit: Estimation and Prediction" in the Getting Started section of the PROC MDC documentation also illustrates the use of DATA step code to transform the input data.
If individuals have choice sets varying in size, and if the value of the variable corresponding to an alternative that is not in the choice set for a specific individual is missing in the original data set, then after the transformation the observation corresponding to this unavailable alternative will have a missing value and will be ignored during estimation. PROC MDC automatically omits observations missing on any variable in the model. So, the number of observations actually used in estimation is equal to the sum of the numbers of choices that are presented to all individuals. With differing numbers of choices across individuals, you will need a numeric variable in the data set containing the possible choices for each individual. Specify this variable in the CHOICE= option in the MODEL statement since the NCHOICE= option is not allowed when choice sets differ in size.
The following example illustrates transforming the original data with differing numbers of choices for each individual into the required format using both DATA step code and the MDCDATA statement.
In data set TRAVEL, note that only the Transit mode is available in the choice set for individual 2. Consequently, the AUTO variable is missing for this individual.
data travel;
input auto transit mode $;
datalines;
52.9 4.4 Transit
. 28.5 Transit
4.1 86.9 Auto
56.2 31.6 Transit
51.8 20.2 Transit
0.2 91.2 Auto
27.6 79.7 Auto
89.9 2.2 Transit
41.5 24.5 Transit
95.0 43.5 Transit
99.1 8.4 Transit
18.5 84.0 Auto
82.0 38.0 Auto
8.6 1.6 Transit
22.5 74.1 Auto
51.4 83.8 Auto
81.0 19.2 Transit
51.0 85.0 Auto
62.2 90.1 Auto
95.1 22.2 Transit
41.6 91.5 Auto
;
The following DATA step code transforms the data to the required format.
data new;
set travel;
retain id 0;
id+1;
/*-- create auto variable --*/
decision = (upcase(mode) = 'AUTO');
ttime = auto;
autodum = 1;
trandum = 0;
choice_var = 1;
output;
/*-- create transit variable --*/
decision = (upcase(mode) = 'TRANSIT');
ttime = transit;
autodum = 0;
trandum = 1;
choice_var = 2;
output;
run;
These statements display the transformed data set, NEW.
proc print data=new noobs;
var decision autodum trandum ttime choice_var;
id id;
run;

The analysis (not shown) of data set NEW is done by specifying the new variable, CHOICE_VAR, in the CHOICE= option.
proc mdc data=new;
model decision = autodum ttime / type=clogit choice=(choice_var 1 2);
id id;
run;
Alternatively, you can use the MDCDATA statement to transform data as shown in these statements. The VARLIST option specifies a variable name to use in the model for each set of the choice-specific variables specified in parentheses after the equal sign. In this example data set, choice-specific information is contained in variables AUTO and TRANSIT which represents travel time for the two alternatives, auto and transit respectively. The variable is named TTIME in the VARLIST option to be consistent with the data set NEW above. The SELECT= option specifies the variable in the original data set that contains the chosen alternative for each individual. Note that the variable specified in the SELECT= option must be a character variable with values matching the names that appear in the first set of choice-specific variable names in the VARLIST option. In this example data set, SELECT=MODE and the values of MODE in data set TRAVEL are Auto and Transit matching the names for TTIME in the VARLIST option. For details about the SELECT= option, see this problem note. The ID= option creates a variable in the transformed data set which identifies the individual. The ALT= option creates a variable in the transformed data set that contains the alternatives for each individual. This is the variable you will need to specify in the CHOICE= option when different individuals are presented with different choice sets. The DECVAR= option creates a 0,1-coded decision variable in the transformed data set with value 1 indicating the chosen alternative. The decision variable is created according to the SELECT= variable in the original data set. The OUT= option stores the transformed data. The analysis results (not shown) are the same as from the previous step.
proc mdc data = travel;
mdcdata varlist(ttime=(auto transit))
select = mode
id = id
alt = choice_var
decvar = decision / out= new2;
model decision = auto ttime / type=clogit choice=(choice_var 1 2);
id id;
run;
Below are the data transformed and stored by the MDCDATA statement.
proc print data = new2 noobs;
var decision auto transit ttime choice_var;
id id;
run;
