A Taxonomy of AI Experiments

AI is everywhere these days. Including economic experiments. But what exactly are AI experiments? What bar does an experiment have to clear to deserve that honorable label?

That’s the very question Christina Strobel and I were both pondering, and it’s what brought us together on this project. Why were we thinking about it? It’s a question we often heard while we presented our own work. “Why is that an AI experiment?” people at seminars and conferences would ask. Having that question implies you have some benchmark in mind.

As we started thinking about what that benchmark could be, we realized that we were asking the wrong question. Instead of searching for an ideal AI experiment, it is more productive to allow for a broad definition and then classify different types of AI experiments. That’s the approach inspired, in part, by the work of my doctoral advisor Glenn Harrison and John List on the taxonomy of field experiments. Their work introduced a spectrum of field experiments, instead of defining “the” field experiment.

We decided therefore to go with the following broad definition of AI experiments:

These are experiments that study interactions between human subjects and computers, algorithms, artificial agents, machines, robots, automated agents, and artificial intelligence agents with the goal of better understanding how the (actual or hypothetical) interaction with, or the presence of, these agents affects human behavior and outcomes in organizations and markets.

That seemed broad enough, so then how do you classify these experiments? Sophistication of the AI is what people often tend to think about when they think whether an experiment is a “proper” AI experiment or not. That seemed like a reasonable feature to use for classification.

Our broad approach suggests, however, that you might have no AI implemented at all and still have an AI experiment. Think about vignette studies. We call this type of experiment conceptual AI experiments. These are mostly vignette studies, although not exclusively. They are easy to implement and allow you to model scenarios that are hard to simulate for real, for example, moral dilemmas. The obvious downside is that these experiments are not very natural and lack real consequences of decisions.

Let’s say you want to implement some AI or an algorithm but don’t want to go big on LLMs. Perhaps you want to explore a certain feature of an algorithm that would be otherwise hidden inside the black box of a big model. You want to have control over the AI. In that case, it is natural to create an AI specifically for your study. We call these stylized AI experiments. The AI in those can get fairly sophisticated actually, but it would probably not become something other people would use in daily lives or commercially. The benefit of such experiments is that you can have real consequences and a lot of control, although it’s still not very natural.

If the naturalness of the AI implementation is key but you want to keep a tightly controlled environment, you can conduct a lab or online experiment with an actual LLM or robot. Here you have less control over the design. We call these quasi-natural AI experiments.

Finally, perhaps you want to have an actual AI in a natural environment where it’s used. That would often be the case in a firm or on an online platform. In fact, firms run thousands of these experiments, called A/B testing. Conducting these as scientific studies, however, requires that you find a balance between whether the findings would add to generalizable knowledge and the value for a firm. Otherwise, why would they let you do it in the first place?

To conclude, there is no “the” AI experiment. It’s more useful to think of different types of AI experiments, each with their own pros and cons. Each type has its own use, and one should tailor their design to the research question rather than chasing one aspect like naturalness because it seems fashionable when the research question actually calls for more control. When you chase one aspect, you sacrifice another, which could be more relevant for your question.

You can read our open-access paper here, and here you can find other articles in the Special Issue on “AI/ML in Behavioural Experiments” published in the Journal of Behavioral and Experimental Economics.