9 Causation: Randomized Experiments & Quasi-Experiments
9.1 Introduction
The contents of this chapter are being updated for the Fall 2025 semester. Check back soon.
9.2 Experiments, Randomized Experiments, Quasi-Experiments
In Science—including IWO Psychology—the word experiment has a specific meaning that you may not find in a standard English dictionary. An experiment is any process in which the researcher manipulates one or more variables and then observes for any subsequent changes in one or more variables (usually different than the manipulated variables) as a result. A research study that has an experiment-based design may be called experimental research. In contrast, observational research doesn’t involve manipulating any variables.
One of the main reasons why a scientist might want to use an experiment-based design—versus an observational design—is because the experiment tends to be a much more straightforward approach for discovering a causal relationship (e.g., how one variable causes another). In fact, most scientists are taught that the experiment is the only valid way to make an inference about causation—but we will learn about modern advances in scientific reasoning that allow us to use observational studies to make causal inferences in the next chapter.
In everyday English, the word experimental also connotes something that is new or not yet established or finalized (e.g., “the doctors performed an experimental procedure”). However, in most scientific domains, the phrases experimental research or experimental study almost always connote a procedure in which a variable is manipulated.
In this chapter, I focus on two broad types of experiments: randomized experiments and quasi-experiments.
A randomized experiment (also known as a random experiment) is any research process in which the researcher randomly assigns each sample element into being subjected to one of two or more versions of a manipulated variable (that the researcher is manipulating), such that some of the sample elements will receive one version of the manipulated variable whereas other sample elements will receive another version (and there may be several versions across the sample). Usually, we say the researcher is randomly assigning the sample elements into a condition (e.g., condition 1 versus condition 2 versus condition 3, etc.), and each condition is a unique version of the manipulated variable(s). We can also say a study has a randomized experiment design.
Here is a basic example of a randomized experiment: Imagine the sample elements are employees and the manipulated variable is type of training and there are three types of training (e.g., Training A, Training B, and no training). If the researcher wants to do a randomized experiment, the researcher would randomly assign a type of training to one employee, and then randomly assign a type of training to another employee, and the researcher would continue like that until each employee is assigned a type of training. Thus, some of the employees will be assigned into Training A, and others will be assigned into Training B, and others will be assigned into no training. This is a basic setup for a randomized experiment.
A quasi-experiment is any research process in which the researcher causes the sample elements into being subjected to a manipulated version of a variable (that the researcher is manipulating) but either of the following is true:
- the experiment involves only one condition (i.e., one version of the manipulated variable), such that all the sample elements receive the same condition (i.e., there is no assignment to two or more conditions);
- OR, the experiment involves two or more conditions, but the assignment to condition isn’t random (i.e., it is nonrandom assignment).
For example: Imagine the sample elements are employees, and the manipulated variable is type of training. There are at least two types of quasi-experiments that could be done:
- the researcher assigns all employees to receive
Training A(or they all receiveTraining B, or they all receiveno training); - OR, the researcher allows each employee to be nonrandomly assigned to either of the types of training. For example, the researcher might let each employee pick their favorite type. Thus, some employees will end up in
Training A, others will end up inTraining B, and others will end up inno training. However, they weren’t randomly assigned into those conditions.
The distinction between a randomized experiment versus a quasi-experiment boils down to the difference between random assignment to condition versus nonrandom assignment to condition. Next, I explain the importance of this distinction.
9.3 Randomized experiments: the usefulness of random assignment to condition
Random assignment to condition is a very useful method to help us discover the causal effect of the condition(s) we’re studying. The conditions are different versions of a variable we’re manipulating (e.g., receiving Training A versus Training B). A causal effect is a change (also known as a difference) that the cause creates.
Imagine you implement a training system on some employees, and it causes a change in those employees. The change is observable as a difference between what things were like before versus after the training. Of course, some of the difference might be caused by other things unrelated to the training, so we would need to figure out what portions of the difference were caused by the training.
Unfortunately, when we’re studying people, we can’t simply manipulate a variable in one group of people and then observe for any subsequent changes in those people. This is because there are an infinite number of factors that could’ve also affected those people during and/or after our manipulation of the variable.
- Imagine a group of workers received training for their jobs, but then soon after the training several of them started using artificial intelligence to help them do their work (this is called a history effect). If we didn’t know that and we measured their
job performancethree months after the training, we would incorrectly believe the training helped a lot more than it actually did! - Imagine the workers just needed more time to get accustomed to their jobs, and they were already going to achieve higher levels of
job performancewithout the training (this is called a maturation effect). In that case, we wouldn’t know whether (nor how much) the training actually helped. - Imagine the company’s beloved CEO resigned, and consequently all the workers’
job performancewas down that week because they were all in a downmood(this is also called a history effect). Everyone’sjob performancewas measured that week, because the managers wanted to measure everyone’sjob performancebefore an upcomingtraining, to compare it against performance after the training. A couple weeks later, everyone naturally started feeling back to a normalmood, and theirjob performancewent up (but no one measured it). One week later, everyone was trained and their post-trainingjob performancewas measured. In that case, we wouldn’t know whether (nor how much) anyjob performanceincrease was due to them feeling back to a normalmoodafter the bad news, versus being due to thetraining.
The next For example… block provides a clear example of how we can get around this issue and actually measure the true effect of any manipulation, without accidentally including any effects from other factors that might also be influencing our results.
Imagine you magically could teleport back and forth to a duplicate universe where everything was 100% identical to your home universe. In other words, Universe 1 (your home universe) is identical to Universe 2, down to the smallest atoms and grains of sand. The only way you can tell them apart is because you see an exact duplicate of you in Universe 2.
Now, imagine you want to measure the effect of a training system for employees at ABC Company. You teleport to Universe 2 and convince your clone to not implement the training system at ABC Company in their universe. Then, you go back to your universe, and you implement the training system at ABC Company in your universe. At this point, the only difference between Universe 1 and Universe 2 is the fact that the training system was implemented in Universe 1, but not in Universe 2.
Then, you measure everyone’s job performance at ABC Company in your universe, and then you teleport to Universe 2 and convince your clone to do the same measurement of everyone’s job performance at ABC Company in their universe. At this point, if there are any differences in job performance between Universe 1 versus Universe 2, those differences must’ve been caused by the only difference between the two universes: the training system.
In this situation, we wouldn’t have to worry about whether the workers used artificial intelligence to increase their performance, nor whether their performance increased because they naturally got accustomed to their jobs, nor whether they started feeling better a few weeks after their beloved CEO resigned. Although those factors impact everyone’s performance, their impact is identical in Universe 1 and Universe 2. Thus, Universe 1 and Universe 2 would both have experienced the same enhancements from A.I., the same enhancements from naturally getting accustomed to the job, and the same enhancements from feeling better after the CEO resigned.
Therefore, after we implement the training system in only Universe 1, if we later see any difference in job performance between Universe 1 and Universe 2, it must’ve been caused by the only difference between the two universes: the training system.
The lesson in the above For example… block is that the identical universes allow us to compare two groups that are identical to each other (i.e., the ABC Company employees in Universe 1 are identical to those in Universe 2) except for one difference: the manipulated variable (i.e., the training system). Thus, we were able to guarantee that any subsequent changes in job performance were actually caused by the training system.
Of course, in reality we don’t knowingly have access to duplicate universes to do the comparisons to calculate the effect of our manipulated variable. Thus, we must rely on the next best thing: comparing two groups that are nearly equivalent to each other except for the one thing we are manipulating. Next, I explain how to do that.
9.3.1 How to create nearly-equivalent groups: random assignment to condition
A simple way to help create nearly-equivalent groups is by randomly assigning each of our sample elements into a condition (remember: a condition is one version of our manipulated variable), such that several of our sample elements will end up being subjected to condition 1, whereas others will be subjected to condition 2, or condition 3 (if there are more than two conditions), etc. Because each sample element was randomly assigned to a condition, the characteristics of the group of elements in condition 1 will on average be similar to the characteristics of the group of elements in condition 2, and condition 3, etc. In other words, it is unlikely that any group will “stand out” or be systematically different than the other groups, because each of their elements were randomly assigned (i.e., each element from our sample had an equal probability of being assigned to any of the conditions). Thus, we have a high likelihood of ending up with nearly-equivalent groups, such that the only difference between the groups will (on average) be the different conditions.
Remember from Section 7.2 there are several other methods for randomly selecting elements. You can use those for random assignment too. For example, stratified random assignment and/or randomized matches can be excellent techniques if your sample size is small. When the sample size is small, there’s a higher likelihood that simple random assignment won’t produce (nearly) equivalent groups.
Keep in mind, when we use random assignment, the likelihood of ending up with nearly-equivalent groups is increased when we use larger groups (this is because of a mathematical fact known as the Law of Large Numbers, but that is beyond the scope of this course).
If you have multiple versions of your manipulated variable, each one could correspond to a different intervention-group. So you could end up with several intervention-groups and one control-group.
When we create nearly-equivalent groups, we end up with two or more groups that we can compare against each other. In the simplest scenario, we could make one of the groups be subjected to a special condition we’re interested in (this is often called the intervention-group, or the treatment-group, or the experimental-group), and make another group be subjected to a non-special condition (this is often called the control-group). Often, if the manipulated variable is something like type of training or type of intervention, the control-group will receive no training or no intervention (which is still technically a version of the manipulated variable).
A classic type of experiment is called a randomized controlled trial (RCT) because it involves randomized assignment to condition and also one of the conditions is a control-condition (i.e., no intervention, no treatment, etc.).
If we want to measure the effect of a training system on the employees, our manipulated variable would be experiencing the training and each employee would be subjected to one of two possible experiences: yes they experience the training or no they do not experience the training. Thus, we can randomly assign each employee to one of those two possibilities, such that we end up with half of the employees in the yes they experience the training condition (also known as the intervention-condition or treatment-condition), and the other half is in the no they do not experience the training (also known as the control-condition).
We could also have two types of training (i.e., Training A and Training B)—corresponding to condition 1 and condition 2—and the third condition would receive no training (thus they would be the control-group). Thus, we would have three separate conditions.
Either way, by using random assignment, each sample element will have an equal probability of ending up in any of the groups. Thus, if our sample contains some younger workers, some older workers, some highly-motivated workers, and some unmotivated workers, random assignment to condition will help us end up with two or more groups that contains relatively equal proportions of younger workers, older workers, highly-motivated workers, and unmotivated workers. In other words, the group of workers who get Training A will on average be similar to the group of workers who get Training B. Of course, the likelihood of ending up with nearly-equivalent groups is increased if we use larger groups (e.g., each group contains 100 employees, instead of only 15).
Random assignment is a subtly different concept than random sampling. Random sampling helps our sample be representative of our target population. Random assignment to one of two or more conditions will help us create two (or more) groups that are nearly-equivalent to each other on average (assuming we end up with sufficiently large groups). Technically, you can think of random assignment as the process of randomly sampling from your sample so that you end up with two essentially equivalent subsamples from your sample.
9.4 Quasi-experiments: when you can’t or aren’t allowed to randomly assign to conditions
As an IWO Psychologist, sometimes you might be asked to evaluate the effect of a new system that impacts the employees (e.g., a new hiring system, a new training system, a new promotion system), but the company managers already implemented it without creating any comparison-groups (e.g., a control-group) or they want you to implement it without any comparison-groups. Often, this occurs because the company managers want everyone to receive the anticipated benefits of the new system (but the managers don’t understand the necessity of a comparison-group to help measure the effect of the system). Other times, the company might have comparison-groups, but they weren’t randomly assigned. For example, the company might’ve allowed the employees to self-select into whichever condition they prefer, or the company might’ve assigned all of the engineers to one type of condition and the marketers to the other type of condition.
Either way, those are quasi-experiments. Remember, earlier in this chapter I said a quasi-experiment is:
any research process in which the researcher causes the sample elements into being subjected to a manipulated version of a variable (that the researcher is manipulating) but either of the following is true:
- the experiment involves only one condition (i.e., one version of the manipulated variable), such that all the sample elements receive the same condition (i.e., there is no assignment to two or more conditions);
- OR, the experiment involves two or more conditions, but the assignment to condition isn’t random (i.e., it is nonrandom assignment).
Next, I explain how you can enhance the validity of your methods if your only option is to move forward with a quasi-experiment (instead of a randomized experiment).
9.4.1 What to do if all you have is a quasi-experiment
If your only option is to move forward with a quasi-experiment (instead of a randomized experiment), try to collect as much data as possible—especially from before and after the intervention (i.e., after you manipulate the variable). Having all that data can help you figure out whether any post-intervention changes are actually caused by the intervention, versus being caused by other things.
If your quasi-experiment involves only one version of the manipulated variable (i.e., only one condition), then all of your sample elements belong to that one group. Thus, we would call that a one-group quasi-experiment. If you measure the outcome variable (e.g., job performance) only after the intervention, we call that a one-group post-intervention design. A slightly stronger method would be to measure the outcome variable (e.g., job performance) before and after the intervention, and we would call that a one-group pre-and-post-intervention design. If you collect pre-intervention and post-intervention data multiple times, we call that a one-group interrupted time-series design.
Technically, if you have a randomized experiment, at minimum you need to compare the results (of the outcome variable) across each group after the intervention (i.e., after the sample elements are subjected to their assigned condition). However, you can increase the power and precision of your statistical analyses by including pre-intervention data (e.g., pre-training data, pre-condition data). If you’re worried about having a small sample size, you could also use the pre-intervention data to help you use fancier methods for random assignment, such as randomized blocks and/or randomized matches.
If your quasi-experiment involves multiple versions of the manipulated variable, such that some sample elements nonrandomly ended up into condition 1, whereas other sample elements nonrandomly ended up into condition 2, etc. (i.e., you have nonrandom assignment to condition), then your groups are less likely to be nearly-equivalent to each other (i.e., less likely than if we would’ve used random assignment to condition). Thus, we would say those groups are non-equivalent groups (NEG). Still, because the ultimate goal of any experiment is to discover a causal effect of the manipulated variable (e.g., type of training) onto an outcome variable (e.g., job performance), a NEG design is generally a more valid approach for making such causal inferences—although the randomized experiment is generally more valid. In the next section, I describe how we can think about the validity of any experiment. Still, when we have non-equivalent groups in a quasi-experiment, we can have any of these designs:
- nonequivalent groups post-intervention design (i.e., the outcome variable is only measured after the intervention);
- nonequivalent groups pre-and-post-intervention design (i.e., the outcome variable is measured before and after the intervention);
- nonequivalent groups interrupted time-series design (i.e., the outcome variable is measured several times before and after the intervention).
If the intervention is an exam or a test, researchers often use the word test instead of intervention (e.g., a post-test design). Also, sometimes people write non-equivalent (with the hyphen) instead of nonequivalent. You can also write NEG as an acronym for nonequivalent groups.
9.5 The validity of an experiment
Remember from Section 7.1.1.1 I said a good definition of validity in the domain of Psychological Science is:
the extent to which a structure (e.g., a tool or psychometric instrument) or process (e.g., a method or experiment) is well-founded on or in accordance with facts or reasonable principles, and appropriate for a specified context or purpose.
When we talk about the validity of an experiment, we are talking about the extent to which the experiment procedure is logical and reasonably good for discovering a causal relationship between two or more variables.
Recall from Section 7.1.1.1 there are many ways to define what is valid, and often it depends on the specific context you’re operating in. For experiments, scientists usually think about two types of validity: internal validity versus external validity. An experiment’s internal validity is the extent to which the experiment allows the researcher to make a valid inference about a causal relationship between the manipulated variable and an outcome variable in the experiment. An experiment’s external validity is the extent to which the experiment allows a researcher to use a valid inference about a causal relationship between the manipulated variable and an outcome variable (from the experiment) to make another valid inference about another population, or other context, or about other variables.
Here’s an example to help you understand internal validity versus external validity.
Imagine we have an experiment in which we want to discover whether the type of training the employees receive will impact their job performance. If there are two types of training (Training A and Training B) and we let employees choose their favorite type to participate in, our results would likely end up biased because the employees who chose Training A might be very different than the employees who chose Training B (e.g., if Training A is a harder training, perhaps most of the highly motivated employees chose it, and most of the unmotivated employees chose Training B). Thus, if we were to measure job performance three months after training, we wouldn’t be able to know how much of the average difference in job performance from the employees who chose Training A versus Training B was actually caused by the trainings, versus being caused by the differences that already pre-existed among the workers who preferred one training versus the other training. Thus, we would say the internal validity of the experiment is threatened.
Now, imagine we used a different design for the experiment (to correct that threat against its internal validity), such that we randomly assign each employee to either Training A or Training B (so, some employees end up in Training A and others end up in Training B). In that case, assuming we have a large enough sample, there is a much lower likelihood that the employees in Training A will be consistently different than the employees in Training B on average, because they were all randomly assigned. That’s the beauty of random assignment to condition. Thus, we would have greater confidence that any subsequent differences in job performance three months later is likely caused by the type of training (because there are likely little or no pre-existing differences on average between the employees in Training A versus those in Training B). Thus, we would say the experiment appears to be internally valid.
However, even if our experiment is internally valid, a separate issue would be its external validity (also known as generalizability).
- For example, if all of the employees are highly educated, we might not be able to use the results to make an inference about less-educated populations, or other populations that are meaningfully different than our sample.
- Also, although we can probably generalize the experiment’s results to help us make inferences about the effects of
type of trainingon other variables that are similar tojob performance(e.g.,job self-efficacy), we probably can’t use it to make inferences about variables that are very different thanjob performance, such asextroversionoremotional intelligence. - Also, depending on the details of the employees’ work environment and the trainings, we might not be able to generalize the experiment’s results to help us make inferences about the effects of training within elementary school classrooms.
As you can see, the external validity of an experiment depends on what you wish to generalize the results to (e.g., other populations, or other variables, or other contexts). An experiment may be externally valid for some purposes, but not for other purposes.
9.5.1 Threats against the internal validity of an experiment
There are numerous features that can threaten the internal validity of an experiment. A complete list is theoretically impossible, because of the vast variety of perspectives about the meaning of validity. However, here I introduce some of the most important and common threats.
- Bad luck with the randomization process: When our randomized experiment has large groups, we have a reasonably high confidence that the groups are at least nearly-equivalent to each other on average, in terms of any characteristics that might impact our outcome variable(s). However, even if we use a truly random procedure to assign participants into their condition, it’s still possible to have bad luck such that we end up with groups that aren’t at least nearly-equivalent. This effect goes by several different names: baseline imbalance, random imbalance, covariate imbalance. That reduces (or possibly destroys) the internal validity of our randomized experiment.
- Anything in a quasi-experiment that causes our two or more groups to be very non-equivalent to each other on average. A common example is called an allocation bias, which occurs when the sample elements that end up into
condition 1(of the manipulated variable) are consistently different than the sample elements that end up intocondition 2(andcondition 3,condition 4, etc. if there are more conditions). - Measurement reactivity effects: Measurement reactivity is any change in a person or other thing being measured, and the change is caused by the measurement process. For example, if an employee knows their performance is being measured for an experiment, they may perform better at their job because they are excited about being part of an experiment (regardless of whether they are consciously trying to perform better or not). A classic example of measurement reactivity is called stereotype threat, in which a person behaves different than they normally would because they are reminded (during the measurement process) about a stereotype about a group they belong to.
- Social desirability effect (also known as the positive self-presentation effect): This is the tendency of people to present themselves (whether consciously or subconsciously) in a manner that they believe is socially acceptable (i.e., socially desirable), even if it requires hiding their true feelings, beliefs, or habits. This can threaten or destroy your experiment’s internal validity.
- Expectancy effects: An expectancy effect is any effect (i.e., any difference in an outcome variable) that was caused by a person’s expectations about what they believe will happen. A classic example is the placebo effect, in which a person actually experiences an improvement in an outcome variable because they were deceived into believing they received a beneficial treatment.
- Attrition effects: Attrition is any reduction in the number of participants in a study after the study begins. Imagine you randomly assigned half of your sample into
condition 1, and the other half intocondition 2. If participants decide to not show up, or they drop out after the study already began, the people incondition 1might no longer be at least nearly-equivalent to those incondition 2. For example: the people who dropped out ofcondition 1might be all of the unmotivated or disorganized workers, or perhaps they are all wealthy workers who don’t need the small monetary incentive from the experiment. - Non-attrition missing-data effects: If your data-collection procedure allows participants to refuse to answer some questions or refuse to provide some data, you will end up with missing-data. That is not necessarily a problem, as long as there is no pattern in the missing data. However, if you are missing some data on only some types of people (e.g., unmotivated workers, disorganized workers, etc.) then your final data is biased (e.g., your final data contains more responses from highly motivated workers, and highly organized workers).
- The intervention didn’t occur exactly as was intended. For example: the employee training wasn’t what the researcher planned it to be. Imagine you wanted to compare an active-learning training versus a lecture-based training, but then what was supposed to be an active-learning session ended up being nearly identical to a lecture-based training.
- Noncompliance effects: If your sample elements are persons, they might not follow instructions. If they are randomly assigned to receive
condition 1, but they instead force their way intocondition 2, then they weren’t randomly assigned intocondition 2. If enough people do that in your experiment, you now have a quasi-experiment, instead of a randomized experiment. You might also have a crossover effect, in which a participant starts within their randomly assigned condition (e.g.,receiving Training A) for a while, but then they make themselves switch into another condition (e.g.,receiving Training B).
9.6 Suggested Readings
Shadish, Cook, and Campbell’s (2002) book titled Experimental and Quasi-Experimental Designs for Generalized Causal Inference is a classic in the Social and Behavioral Sciences—including IWO Psychology. Reichardt’s (2019) book takes a somewhat more modern approach to the topic. In our shared Group Zotero Library I’ve included a copy of the chapter titled Randomized Experiments from Reichardt’s (2019) book.
Fun fact: Charles S. “Chip” Reichardt was one of Donald T. “Don” Campbell’s Ph.D. students, so it’s not surprising Reichardt wrote a book similar to the classic that Campbell co-authored.