Video Modeling: A Practitioner's Guide for BCBAs, RBTs, and School Teams
Based on 51 experimental studies (36 controlled, 15 suggestive); 92% report positive effects; where reported, effects are predominantly large. Updated July 2026.
How we grade →01What the research shows
Across 51 experimental studies (36 controlled, 15 suggestive), 92% of the studies reporting a direction found positive effects. Where effect size was reported, effects were predominantly large.
Populations studied: autism, intellectual disability, neurotypical learners.
Computed across 51 corpus articles (51 experimental, 0 contextual). Regenerated monthly as new studies are ingested.
02The variants, and how they differ
Video modeling teaches a target skill by having the learner watch a recorded example of the behavior, then perform it, rather than watching a live demonstration in the moment. The recording is the active ingredient: it can be scripted, edited, reviewed before a session, and reused across learners in a way a live model cannot. The variants below differ in who appears in the recording, how it's paired with other procedures, and who the intended viewer actually is.
Standard video modeling: another person performs the target skill
This is the reference form: an adult, peer, or sibling is recorded performing the target behavior correctly, and the learner watches the clip before an opportunity to perform it themselves. Home safety skills held up especially well here, four preschoolers with ASD reached 100% accuracy avoiding household chemicals after watching brief videos their own fathers recorded (Mortaş Kum et al., 2025). The same structure extended to a higher-stakes safety target, teaching children with ASD to refuse abduction lures from strangers and known adults using a code word, though here the video was paired with in-vivo rehearsal rather than run as a standalone procedure (Abadir et al., 2021). Standard video modeling also taught symbolic play, three young children with ASD acquired, maintained, and generalized playing with imaginary objects after watching an adult model the pretend actions (Lee et al., 2021).
Video self-modeling and feedforward self-modeling
Here the learner is the model. The clip is edited to show the learner performing the target behavior correctly, sometimes stitched together from brief correct fragments the learner hasn't yet strung into a full performance (feedforward self-modeling). A comparison of teacher-model video and feedforward video self-modeling for reading fluency and comprehension found the self-modeling format worked for two of four participants, but results were inconsistent across the group, this is not a format to default to without a fluency probe confirming it's the better fit for the individual learner (Egarr et al., 2021).
Human video modeling versus animated video modeling
Both use a recorded model, but an animated model swaps the human actor for a cartoon or digitally rendered character. A direct comparison teaching intraverbal responding and motor imitation of facial expression found seven of eight participants acquired the targets with one or both formats, two learners did better with human video, three did better with animated, and three showed little difference, neither format was conclusively superior (Bloh et al., 2025). Treat the choice as an empirical question to probe per learner, not a format with a default answer.
Joint video modeling
The learner watches the recorded model alongside a peer rather than alone, embedding the viewing itself in a social context. Joint video modeling improved unscripted, unprompted pretend-play verbalizations for preschoolers with ASD interacting with typically developing peers in an inclusive classroom, a result specific to the spontaneous layer of play rather than the scripted actions the video directly demonstrated (Dueñas et al., 2019). Video modeling for peer play targets more broadly has also taught scripted commenting during shared leisure activities between dyads of children with ASD, though roughly half the participants needed the video paired with tangible reinforcement and additional prompts to reach mastery, video modeling alone was sufficient for the other half (Ezzeddine et al., 2020).
Video modeling as a component within a larger package
In several studies, the video is not the sole active ingredient but a supplement layered onto another primary procedure. Self-monitoring paired with a daily video-based supplement improved turn-taking and question-asking for adolescents with ASD, the video functioned as a refresher model rather than the mechanism doing the teaching on its own (Ayvazo et al., 2024). A related but mechanically distinct format, live video-chat coaching rather than a pre-recorded clip, taught social conversation skills to seven-year-olds with ASD talking with family members over video, this is video-mediated instruction in real time and shouldn't be conflated with recorded video modeling when deciding what a "video-based" program actually involves (Brodhead et al., 2019).
Video modeling to train the interventionist, not the direct learner
One study in this set used video modeling to train neurotypical adolescent peers, not autistic learners, to deliver a ten-step peer-mediated social interaction procedure to classmates with ASD (MacFarland et al., 2025). Flag this one carefully: the direct recipients of the video model were the neurotypical peer-trainers, and the ASD learners who benefited did so secondhand, through the trained peers' improved delivery. It's a legitimate and increasingly common use of the format, but it answers a different clinical question than a video model aimed straight at the learner with the skill deficit.
03Which one, and when
Before picking a video-modeling format, confirm the learner can actually use one. Watching a screen and extracting a model to imitate requires sustained visual attending to a two-dimensional representation and an imitation repertoire strong enough to translate what's seen into a motor or vocal response. A learner without either prerequisite is a candidate for live, in-person modeling first, not a video workaround. Model-lead-test, a live-modeling procedure with no recorded component, reliably taught an elementary student with ASD and severe problem behavior to calibrate, track, and code a robot, with full generalization and maintenance (Knight et al., 2019). That result is worth keeping in view precisely because it shows live modeling doing the job cleanly: when a skill is procedural, the learner already attends well to an in-person demonstration, and there's no standing need to reuse the model across sessions or sites, the production overhead of a video model buys you little.
Reach for video modeling instead when at least one of three things is true: the model needs to be reused across many sessions or multiple learners, the behavior needs to be demonstrated by someone who can't be physically present for every trial (a parent recording once rather than repeating a live demonstration daily), or the target is a safety skill where practicing the actual error live carries real risk. The home-safety and abduction-prevention evidence here leans on exactly that logic, a father recorded a 30-second clip once and it drove four children to 100% accuracy avoiding household chemicals (Mortaş Kum et al., 2025), and a scripted video plus in-vivo rehearsal taught differentiated, generalized refusal of abduction lures without ever needing a live confederate to model the correct refusal response (Abadir et al., 2021).
Within video modeling, don't default to a human actor or to teaching-others-as-model without checking the individual learner. A head-to-head comparison of human video modeling against animated video modeling found neither format won consistently, results split roughly into thirds across learners who did better with human video, better with animated, or showed no difference (Bloh et al., 2025). The same caution applies to self-modeling: feedforward video self-modeling outperformed a teacher-model video for only half the participants in one reading-fluency comparison, with inconsistent results overall (Egarr et al., 2021). Treat human-versus-animated and self-versus-other as empirical questions you probe per learner across a few sessions, not defaults you commit to before collecting data.
Don't assume video modeling alone will carry a target to mastery. In a play-comments program, video modeling by itself was sufficient for half the participants, the remaining half needed it paired with tangible reinforcement and additional prompts (Ezzeddine et al., 2020), and a conversational-skills program built the video in as a supplement to a primary self-monitoring procedure rather than as the sole intervention (Ayvazo et al., 2024). Build in a decision point after a handful of sessions: if the video alone isn't moving the data, add reinforcement or prompting rather than concluding video modeling failed as a format.
Choose live video-chat coaching over a pre-recorded model specifically when the goal is generalized conversation with a real, current communication partner, not a scripted demonstration. Video-chat mediated coaching taught social conversation to children with ASD talking with actual family members, and skills generalized to unfamiliar adults who had no prior experience with the participants (Brodhead et al., 2019). That's a different tool than a recorded clip: it trades reusability for real-time responsiveness, and it's the better fit when the target partner and context can't be scripted in advance.
04What this means Monday morning
Once you've decided video modeling fits, what determines whether it works is how the clip gets built and how tightly the viewing is tied to the opportunity to respond, not how polished the production is.
Keep production simple and the script tight to the exact target response. The home-safety evidence here didn't come from a studio-quality video, it came from a father recording a 30-second clip of himself saying "no, chemicals are dangerous" and walking away, and that low-production model was enough to drive four children to 100% accuracy (Mortaş Kum et al., 2025). Extra context or dialogue in the clip is something the learner has to filter out, not something that helps. For a safety target with multiple lure types, script separate short clips per variation rather than folding stranger and known-adult lures into one video, that separation is part of what let responding generalize and differentiate correctly by lure type in the abduction-prevention program (Abadir et al., 2021).
If you're building a self-modeling or feedforward self-modeling clip, budget real editing time. The learner has to be filmed attempting the target, then the footage gets cut to only the correct fragments or spliced into a full performance the learner hasn't yet produced end to end. That's meaningfully more work than a standard other-model video, and the format doesn't reliably outperform a teacher-model video, so don't take on the production cost by default; probe both formats across a handful of sessions before investing in a self-modeling library (Egarr et al., 2021).
Run the viewing immediately before the opportunity to respond, not as a separate activity earlier in the day, priming the model right at the point of performance rather than hoping it transfers hours later. For play and social targets, decide up front whether the learner watches alone or with a peer: joint viewing before a play session improved unscripted, spontaneous play talk in an inclusive classroom, a benefit specific to the joint-viewing setup (Dueñas et al., 2019).
Watch the data for a stall after the first few sessions, that's a cue to add support, not switch formats. In one commenting program, video modeling alone got half the participants to mastery; the rest needed the same video paired with tangible reinforcement and extra prompts (Ezzeddine et al., 2020). Build that reinforcement layer into the plan as a pre-written fallback, not an improvised fix.
Program generalization deliberately. Vary the setting, the person present, and, for safety targets, the specific scenario across probes once the learner responds reliably to the trained clip, the abduction-prevention data held up because generalization was assessed across novel confederates and locations, not just re-tested with the same clip in the same room (Abadir et al., 2021). Fade the pre-response viewing itself once a learner holds criterion across two or three consecutive sessions without needing the refresher.
05From the experts
So three RBTs were paired with these trainees. We then wanted to see, could the skills that we learned in the video modeling and voiceover transfer to the therapist? And it did. Long story short, there was 88% integrity, or excuse me, three to six feedback sessions all reached mastery, meaning that the pyramidal sprinkling worked. Video modeling with voiceover was able to teach the supervised and effective skill that they were then able to push on to their behavior technicians. So what does that mean? It means a couple really important things.
I try to incorporate short videos of examples with actual students. So yeah, a lot of people are visual, right? So even the video modeling can be really successful in teaching certain learners as well. Fidelity data, job aids. Yeah, fantastic. I think that these are all great ideas. Another thing that I noticed here, right, is there was a little bit of fear created or contrived in the office scenario. And sometimes without us really recognizing it, we may be using fear or some guilt to shape BT behavior.
And so you'll see, we have a list of the different safety skills here, behavioral skills, trainings up at the top, DTT, prompt and prompt fading, using reinforcers, video modeling, all the way down, right? I want to focus a little bit on that top one on behavior skills training. So for those of you that don't know what it is, it's, it's a really structured teaching strategy that's going to walk a client through four steps in order to help keep them safe.
06Common questions
- My learner won't attend to the screen for more than a few seconds. Should I keep pushing video modeling or switch to a live demonstration?
- Switch to live modeling for now. Video modeling assumes the learner can sustain attending to a two-dimensional representation and has enough imitation repertoire to translate what's seen into a response. A live, in-person model-lead-test procedure with no video component reliably taught a student with severe problem behavior a full procedural skill sequence with generalization and maintenance. Build the screen-attending skill separately, then revisit video.
- Should I use a human actor or an animated character for my video model?
- Probe both instead of picking one. A head-to-head comparison found neither format won consistently: some learners did better with human video, some with animated, several showed no meaningful difference. Run a short alternating comparison before committing production time to either format.
- I filmed my client attempting the task and want to edit out the mistakes to build a 'self-model' video. Is that a legitimate technique or am I fabricating data?
- It's a legitimate, published technique, feedforward video self-modeling, not fabrication, since the goal is teaching a not-yet-fluent response, not documenting one that occurred. Don't assume it outperforms a standard teacher-model video by default: in one comparison it worked better for only half the participants, with inconsistent results overall. Probe it against a standard video model rather than defaulting to it.
- Is it better for my learner to watch the model video alone or with a peer present?
- For play and social targets, joint viewing has a documented benefit solo viewing doesn't automatically produce. Watching the model with a peer before a play session improved unscripted, spontaneous play talk in an inclusive classroom. If the goal is spontaneous social behavior rather than a scripted response, build the peer into the viewing itself.
- A caregiver wants to coach my client over a live video call instead of me sending pre-recorded clips. Is that the same intervention as video modeling?
- No, treat it as a different tool. Video-chat coaching is real-time and responsive to whoever is on the call, which is why it generalized conversation skills to unfamiliar adults during live sessions with actual family members. A pre-recorded model is reusable and scripted but can't adjust in the moment. Choose live coaching when the partner and context can't be scripted in advance, a recorded model when you need a consistent demonstration across many sessions.
07The studies behind this grade
The strongest 12 of 51 constituent studies. Each links to its record in the research database and its source.
- Teaching home safety skills to children with autism spectrum disorders
- Comparing Human Video Modeling to Animated Video Modeling for Learners with Autism
- Using Video Modeling to Teach Neurotypical Adolescents to Interact Socially with Peers with ASD.
- Supporting the Conversational Behavior of Adolescents with Autism Spectrum Disorders with Self-Monitoring and a Video-Based Supplement
- Making Deception Fun: Teaching Autistic Individuals How to Play Friendly Tricks
- Effects of video modeling on abduction-prevention skills by individuals with autism spectrum disorder
- Effects of Video Modeling on the Acquisition, Maintenance, and Generalization of Playing with Imaginary Objects in Children with Autism Spectrum Disorder.
- Model Teachers or Model Students? A Comparison of Video Modelling Interventions for Improving Reading Fluency and Comprehension in Children with Autism
- Using video modeling to teach play comments to dyads with ASD
- Effects of Joint Video Modeling on Unscripted Play Behavior of Children with Autism Spectrum Disorder.
- Teaching Robotics Coding to a Student with ASD and Severe Problem Behavior.
- A Pilot Evaluation of a Treatment Package to Teach Social Conversation via Video-Chat.