Methods
What UtilityLab computes, exactly as the code does it: the choice model, the D-error, how designs are generated and checked, and how blocks reach respondents. Read it with one caution in mind: a D-efficient design is only as good as the priors it was built on (Walker et al., 2018).
Last updated 7 October 2026
1. The choice model
UtilityLab scores designs under a multinomial logit (MNL) model. The utility of alternative j in choice task s is V_sj = β′x_sj, and the choice probabilities are the usual logit probabilities. The coding of x_sj is:
- Numeric attributes enter linearly with one parameter, in the value entered (for example minutes or euros), without centring or scaling.
- Categorical and boolean attributes are dummy coded with L − 1 parameters. The base is the first level in the attribute’s level order. For a yes/no attribute that is the first level listed (by default “Yes”), so its parameter reads “No vs Yes”.
- Attribute parameters are generic: one parameter (or set of dummy parameters) per attribute, shared by every alternative it applies to, in labeled experiments too. Alternative-specific attribute parameters and interactions are not supported.
- An attribute that does not apply to an alternative contributes zeros for that alternative. Context variables are not part of the model: they are shown to respondents but do not enter the D-error.
- Pivoted attributes enter with the values as entered (the offsets or multipliers), not the values a respondent sees after pivoting on a reference.
Constants depend on the experiment type and on whether there is an opt-out:
| Experiment | Constants | Reference (utility 0) |
|---|---|---|
| Labeled, no opt-out | One per alternative except the last | The last alternative in the list |
| Labeled, with opt-out | One per designed alternative | The opt-out |
| Unlabeled, no opt-out | None | — |
| Unlabeled, with opt-out | One, shared by the designed alternatives | The opt-out |
The opt-out has no attribute levels and a utility of zero. The number of parameters K counts the constants and the attribute parameters.
2. The D-error
For a design with choice tasks s = 1…S, UtilityLab builds the MNL Fisher information matrix (Huber & Zwerina, 1996)
F(β) = Σ_s X_s′ ( diag(P_s) − P_s P_s′ ) X_s
where X_s stacks the coded alternatives of choice task s (opt-out included) and P_s holds their choice probabilities at the priors β. The asymptotic variance–covariance (AVC) matrix is its inverse, and the D-error is
D-error = det( F(β)⁻¹ )^(1/K) = det( F(β) )^(−1/K)
Following Walker et al. (2018) and Rose and Bliemer (2009), the D-error is a Dz-error when every prior is zero and a Dp-error when the priors are fixed, non-zero values. UtilityLab labels it accordingly. It does not compute a Bayesian Db-error.
- The information is summed over all choice tasks of all blocks, as if one respondent answered the whole design. Blocks play no part in it.
- Lower is better, but only like for like: adding choice tasks lowers the D-error, so compare designs with the same number of choice tasks, the same model (attributes, coding, constants) and the same priors.
- The log-determinant is computed by a Cholesky factorisation. If a pivot is below 10⁻¹⁰ of the largest diagonal element the information matrix is treated as singular: the design does not identify every parameter, and the D-error is shown as “not computable” rather than as a misleading number.
- The D-error is computed for any design, generated or uploaded.
3. Priors, and why D-efficiency is not the whole story
Priors are fixed point values, entered per attribute parameter on the utility scale (per unit for numeric attributes, relative to the base level for dummy-coded ones). Constants always have a prior of zero. With every prior at zero, the Dz-optimal design is the one that would be efficient if no attribute mattered: a starting point, not a design tuned to any particular preferences.
A D-efficient design built on non-zero priors concentrates its choice tasks where those priors make the alternatives close in utility. That is efficient if the priors are right and can be poor if they are wrong. Walker et al. (2018) show this for a binary mode choice with travel time and cost, where the parameter of interest is the value of time (VOT):
- The D-efficient design built on a prior VOT of $20/hour was the most efficient design when the true VOT was roughly $10–30/hour, became the least efficient outside about $5–40/hour, and estimation degraded markedly above about $50/hour.
- Orthogonal and random designs were robust over a much wider range of true VOT. Bayesian efficient designs were more robust than D-efficient ones, depending on the prior’s variance, and two-stage designs (priors from a small first sample) were not more robust than uninformative designs.
- Their conclusion: arguably, “the random design (which is the easiest to generate) performs as well as any design”, and any design does better once choice tasks in which one alternative dominates the other are removed. They recommend efficient designs only when the priors are excellent.
What this means in UtilityLab:
- Use non-zero priors only when you are confident in them, for example from a pilot or a closely comparable study. Otherwise prefer zero priors, the balanced search or a random design.
- Check robustness yourself: keep the design and change the priors. Diagnostics recomputes the D-error of the same design at the new values, which is the analytical check Walker et al. use. If it rises sharply for plausible alternative values, the design is fragile.
- In unlabeled experiments, look at the dominance check and avoid choice tasks with a dominant alternative (Bliemer, Rose & Chorus, 2017).
4. Generating a design
All three methods take the number of choice tasks and blocks, record their random seed, and can be reproduced by entering that seed again with the same structure and version.
- D-efficient search. A coordinate-exchange algorithm (Meyer & Nachtsheim, 1995) that minimises the D-error at the current priors. Each start is a random design; each pass visits every choice task, designed alternative and applicable attribute and tries every other level, keeping the change that most lowers the D-error (or, when constraints are set, one that removes a violation). It stops after a pass without improvement, or after 30 passes. Several random starts are run (5 by default) and the best is kept. The result is a locally optimal, D-efficient design, not a proven D-optimal one. A run is refused when the design cannot identify every parameter: fewer choice tasks than ⌈K / (J − 1)⌉, an attribute with fewer than two levels for an alternative it applies to, or fewer than two alternatives.
- Balanced search. Draws many random designs (1,000 by default) and keeps the one with the lowest weighted score of level-balance deviation, largest within-alternative correlation, dominance relations and pairs of identical alternatives. The weights can be changed.
- Random. Draws every level independently and uniformly.
Common to all three:
- Blocks are not optimised. Choice tasks are dealt to blocks in turn, so block sizes differ by at most one, and the blocking is not balanced against the attributes.
- Context variables are drawn at random for each choice task; they are not optimised.
- Constraints forbid combinations of levels within an alternative. Random starts and random designs redraw a violating choice task up to 200 times; the D-efficient search never accepts a change that adds a violation and prefers changes that remove one. Any choice task that still violates a constraint is reported.
An uploaded design (for example from Ngene or R) is used as given: UtilityLab maps it to the structure, computes its D-error and runs the same checks.
5. Checking a design
The diagnostics are heuristics. They flag patterns worth a look; they are not tests of significance.
- Dominance (unlabeled experiments only, since in labeled ones the label itself carries utility). Alternative A dominates B in a choice task if, on every attribute, A is at least as good given the attribute’s preference direction and strictly better on at least one. Pairs that do not show the same attributes are not compared, and attributes without a preference direction count as trade-offs, so no dominance is declared through them.
- Identical alternatives. Two alternatives that show the same level on every attribute they share.
- Attribute correlation. Pearson’s r between each pair of attributes within an alternative, across all choice tasks. Numeric attributes use their values, yes/no attributes 1 and 0, and categorical ones their level order, which treats nominal categories as ordered, so read those values with care. Warning above |r| = 0.3, concern above 0.5 (adjustable).
- Level balance. How far each level’s count is from an equal share, across the whole design (not per block). Flagged above 20% (adjustable).
- Constraints. Choice tasks that contain a forbidden combination, and constraints that can never match.
- Observations per parameter = respondents × choice tasks per respondent ÷ K, rated low below 25, borderline from 25 and enough from 50. This is a rule of thumb, not a power analysis; for sample size requirements of efficient designs see Rose and Bliemer (2013).
The analyst view shows each choice task’s MNL choice probabilities at the current priors. With all priors at zero they are equal by construction.
6. Fielding: blocks and answers
- Qualtrics QSF: the blocks sit under a randomizer that presents one of them, evenly.
- Qualtrics TXT: the format cannot hold survey flow, so the file’s setup notes explain how to add that randomizer after import.
- LimeSurvey (file and push): a hidden question BLK draws one block at random for each respondent and is stored with the answers; only that block’s choice tasks are shown, in a random order. The draw is uniform, not balanced, so block sizes can come out unequal in a small sample.
- Answers: each study chooses whether respondents may skip a choice task (optional), are prompted once before skipping (the default), or must answer.
Pivoted attributes are exported with values resolved at one fixed reference for every respondent. Pivoting on each respondent’s own value needs the wiring described in the pivot wiring guide.
7. Limitations
- MNL only, with generic attribute parameters; no alternative-specific attribute parameters or interactions.
- Fixed point priors only; no Bayesian efficient designs (Db-error), and no priors on constants.
- Blocks and context variables are assigned at random or in turn, not optimised.
- Pivoted attributes are modelled on the offsets or multipliers entered, not on respondents’ resolved values.
- Correlations for categorical attributes use level order.
- The Qualtrics and Sawtooth exports follow those platforms’ published formats but have not been tested on live accounts; the LimeSurvey exports have.
8. Reporting
The Export step generates a methods paragraph for your paper from the study: the experiment, the design and how it was generated (with its seed), the Dz- or Dp-error, the response requirement and the check results. Review it before use, and cite the UtilityLab version you used: each release has its own DOI (see How to cite (opens in a new tab)). The source code of everything described here is open, in lib/ (opens in a new tab).
References
- Bliemer, M. C. J., Rose, J. M., & Chorus, C. G. (2017). Detecting dominance in stated choice data and accounting for dominance-based scale differences in logit models. Transportation Research Part B: Methodological, 102, 83–104. doi:10.1016/j.trb.2017.05.005 (opens in a new tab)
- Huber, J., & Zwerina, K. (1996). The importance of utility balance in efficient choice designs. Journal of Marketing Research, 33(3), 307–317.
- Meyer, R. K., & Nachtsheim, C. J. (1995). The coordinate-exchange algorithm for constructing exact optimal experimental designs. Technometrics, 37(1), 60–69. doi:10.1080/00401706.1995.10485889 (opens in a new tab)
- Rose, J. M., & Bliemer, M. C. J. (2009). Constructing efficient stated choice experimental designs. Transport Reviews, 29(5), 587–617.
- Rose, J. M., & Bliemer, M. C. J. (2013). Sample size requirements for stated choice experiments. Transportation, 40(5), 1021–1041.
- Walker, J. L., Wang, Y., Thorhauge, M., & Ben-Akiva, M. (2018). D-efficient or deficient? A robustness analysis of stated choice experimental designs. Theory and Decision, 84(2), 215–238. doi:10.1007/s11238-017-9647-3 (opens in a new tab) (open access: MIT DSpace (opens in a new tab))
See also the terms and disclaimer.