Presenters and teachers can either animate their slides’ elements or not. When they are animated, a presenter’s slides come up a little bit at a time. When they’re not, the text and images are presented all-at-once, in one big blob. You’ve definitely experienced both. Which way do you think is more effective for learning? Make a prediction!

A recent study tested the two styles of presentation, and it provides a great example of a pretest-posttest design. The empirical study in question appears the Journal of Computer-Assisted Learning, and it was accessibly described by a journalist for Study Finds. Here’s some of the Study Finds story:
Researchers at Tokyo University of Science tested two ways of delivering the same recorded biology lesson on interactions between populations, covering predator-prey relationships, symbiosis, and interspecies competition. In one version, students saw a fully loaded slide from the start, with all charts, labels, and illustrations visible before narration began. In the other version, visual elements appeared one by one, timed to match the narration as it unfolded.
We already mentioned this was a pretest-posttest design. Knowing that, read the next passage and identify the Independent and Dependent variables:
Forty Japanese university students took part, all science and engineering majors with normal vision and hearing. Before the lesson began, everyone completed a short test measuring prior knowledge. This baseline confirmed both groups started at a similar level. Participants were randomly split into two groups. One group watched a roughly 20-minute recorded lesson, created by the study’s first author playing a Japanese high school teacher, with all visual content visible the moment each slide appeared. The other group watched the same lesson with identical narration, but slides built up piece by piece, each new element appearing at the moment narration began discussing it. Participants were asked not to take notes and to keep their eyes on the screen.
a) Complete the table below, by identifying the major variables in the study:
| Variable name | Levels of this variable | IV? DV? | If it’s an IV, was it independent-groups or within-groups? |
Students in the cumulative presentation group outperformed those in the standard, full-slide group on the post-lesson test, scoring 17.70 out of 20 on average compared to 16.70. That’s a modest gap, but researchers found it was real and not just due to chance.
The study also did some eyetracking to see where on the slide students were looking during the study. I’m not focusing on that here, but you could read more about it in the paper.
b) Sketch a graph of the test results. You can make a bar graph or a line graph. Here are some hints. First, Study Finds doesn’t mention the exact pretest score for each group–you’ll have to check the empirical article or estimate what you think the pretest scores will be, given other parts of the study. Second, put “pretest” and “posttest” on the x-axis.
Now let’s address four big validities.
c) Reread the study description. Note the sample and think about external validity. Are we able to generalize from this sample to a population of college students in Tokyo? Why or why not?
d) Take a look at the size of the difference (the effect size) and think about statistical validity. Specifically, how much more did students learn in the animated slide presentation compared to the all-at-once presentation?
e) Let’s take construct validity next. What might you need to know about the knowledge measure in order to decide if it is construct valid?
f) There’s a lot to consider with internal validity. For one thing, the students were randomly assigned to the two groups. Why was this important? What internal validity threat does this rule out?
g) The students in both groups also heard the same narration and saw the same content on the slides. What internal validity threat does this step rule out?
h) The students in both groups took a pretest and a posttest. Which internal validity threat does this rule out?
i) Some students might think this study is susceptible to a maturation threat or a testing threat. It’s not. Why not?
Finally, I need to call your attention to what the journalist says about the results: “That’s a modest gap, but researchers found it was real and not just due to chance.”
Despite what this journalist wrote, a statistically significant difference does not mean a result is “real”. And, it doesn’t mean it was “not just due to chance.”
Here’s what the journalist should have said: “the researchers found that it was a statistically significant difference, which means it was unlikely to have been selected, by chance, from a population in which there is no difference between the animated and all-at-once groups.” It’s a lot harder to say, but it’s more accurate!
Selected answers:
f) There’s a lot to consider with internal validity. The students were randomly assigned to the two groups. Why? What internal validity threat does this rule out? Random assignment helps rule out a selection effect. Because random assignment was used, we know that the two groups contain the same types of people. We can’t surmise that “maybe one group was just smarter” or “maybe some students just knew that content already.”
g) The students in both groups also heard the same narration and saw the same content on the slides. What internal validity threat does this rule out? The researcher used these control variables to rule out design confounds.
h) The students in both groups took a pretest and a posttest. Which internal validity threat does this rule out? This rules out a design confound, too. Both groups did the same things, except for the one active ingredient: How the slides were animated.
i) Some students might think this study is susceptible to a maturation threat or a testing threat. It’s not. Why not? It might seem that when people take the test twice (pretest and posttest) they will score better on the second test because of that practice (that’s a testing effect). That’s likely true, but when both groups do it, you have controlled for any testing threat. The animated slide group did even better than the all-at-once group, even though both had the same opportunity to take the test twice.
Pre-test posttest studies take place over time, which makes some students suspicious of a history threat. But it’s the same reasoning: Both groups had the same amount of time pass. The animated slide group did even better than the all-at-once group, even though both had the same amount of time to mature (develop naturally).