Building Better Prediction Models for Consumer Choices
Caltech and MIT researchers have addressed a long-standing gap in consumer choice modeling by refining the random utility model (RUM) to work with incomplete datasets, potentially improving demand predictions across industries.
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
Economists have long used the random utility model (RUM) to predict consumer choices, but its effectiveness diminishes when not all options are known. Researchers at Caltech and MIT, led by Professor Kota Saito and MIT graduate student Alec Sandroni, have now solved a key challenge tied to the linear ordering problem, which has persisted since the 1980s. Their findings, published in the *American Economic Review*, demonstrate how to refine predictions even when some alternatives remain unobserved, addressing a critical gap in empirical analysis.
The team’s work focuses on the limitations of the "outside option" approach, where unobserved choices are grouped into a single category. Sandroni, Saito, and collaborator Haruki Kono found that this method can produce biased estimates. By applying network flow theory, they developed a technique to test RUM with fewer data points than previously thought necessary, enabling more accurate demand evaluations in sectors like retail, automotive, and policy analysis.
The research builds on Sandroni’s undergraduate work at Caltech, where he initially pursued environmental science before joining Saito’s lab through the Summer Undergraduate Research Fellowships (SURF) program. Despite limited prior knowledge of economics, Sandroni contributed significantly to the project, including rigorous proofs and manuscript preparation. Saito praised his rapid advancement, noting the demanding nature of the work for an undergraduate student.
Building on these results, the team has expanded their inquiry into how aggregated data—such as broad product categories—can obscure underlying consumer preferences. A recent working paper explores the ambiguity introduced by such aggregation, showing that it weakens the testable implications of RUM. The researchers are now developing a more general model using monotonicity conditions and exploring applications in machine learning, with support from the National Science Foundation and Caltech’s SURF program.