Assessment & Research · Sub-Pillar

MSWO Preference Assessment: Procedure, Variants, and Decision Logic for BCBAs and RBTs

By Matt Harrington, BCBA · BBC Editorial Team · Search target: MSWO
BBC Evidence Grade: MODERATE

Based on 29 experimental studies (7 controlled, 22 suggestive); 73% report positive effects; where reported, effects are predominantly large. Updated July 2026.

Experimental base 29 studies
Controlled (T1) 7
Suggestive (T2) 22
Convergence 73% positive
How we grade →

01What the research shows

Across 29 experimental studies (7 controlled, 22 suggestive), 73% of the studies reporting a direction found positive effects. Where effect size was reported, effects were predominantly large.

Populations studied: autism, developmental delay, neurotypical learners, intellectual disability.

Computed across 36 corpus articles (29 experimental, 7 contextual). Regenerated monthly as new studies are ingested.

02The variants, and how they differ

Standard MSWO (DeLeon and Iwata array)

The standard procedure arrays a set of items (commonly five to seven) in a line, the learner selects one, the selected item is removed, the remaining items are re-spaced to close the gap, and the sequence repeats until every item is chosen or the learner stops responding. Because a selected item never returns, one pass produces a complete rank order rather than a frequency score, the format's core advantage over multiple-stimulus-with-replacement (MSW) and the reason it earns its own mechanics treatment here rather than a comparison paragraph. Sessions are typically repeated two to three times and ranks averaged, since the last item or two in any single pass was never really chosen so much as left over.

Brief and abbreviated MSWO

A brief-format MSWO compresses the array-and-remove cycle without changing the underlying logic. A free web-based tool built for this compressed version, the MSWO Preference Assessment Tool (MSWO PAT), paired with a short concurrent-operants reinforcer check, identified video preference hierarchies for young children and carried predictive value for most participants (Curiel et al., 2024). A separate evaluation of the same tool with three children diagnosed with autism spectrum disorder ranked high- and low-preferred videos and confirmed the top item functioned as a reinforcer for two of the three (Curiel et al., 2024). Brief administration buys speed, not an exemption from confirming the top rank actually reinforces behavior.

Digital and app-based administration

Moving the array onto a screen or training platform changes who can run the assessment and how fast they learn to, not the selection-and-remove logic itself. A self-instructional online package, manual plus video modeling with no in-person instructor, brought university students from 35% to 94% correct and staff from 23% to 87% (Arnal Wishnowski et al., 2018). An AI-guided coaching tool produced comparable efficiency gains and fewer scoring errors for preservice speech-language pathologists learning the procedure with clients who have intellectual and developmental disabilities (Griffen et al., 2024). Both results are about how reliably a new administrator executes the array, not about client-side outcomes.

Picture-based and verbal/topic MSWO

For learners who cannot physically sample every item, or stimulus classes that aren't discrete objects, the array substitutes photographs, icons, or spoken labels while keeping the selection-and-remove structure. A picture-based MSWO produced valid hierarchies for social interactions most often among learners who could match, identify, and tact the pictures presented, while a format sampling the interaction directly (SIPA) worked better for learners who could not (Morris et al., 2020). The same logic extends to conversation topics for individuals with autism spectrum disorder and advanced vocal-verbal skills, where an MSWO-style ranking of topic labels sat on a continuum with a faster self-report format and a slower, more certain response-restriction assessment (Kronfli et al., 2024). Stimulus modality is a repertoire match, not a difficulty setting.

Staff-training pathway as its own variant

Because scoring depends on the administrator removing the correct item and re-presenting the array correctly every trial, how staff are trained to run MSWO functions almost like a format variable in its own right. Behavioral skills training brought three special education teachers to accurate, generalized implementation across ten children from a baseline of inconsistent scoring (Yang, 2022). A video-modeling package with no live coaching taught staff at a summer camp and school in mainland China to run MSWO alongside a token economy and error correction, generalizing to sessions with children on the autism spectrum and holding at follow-up (Lim et al., 2020). The accuracy ceiling looks similar across BST, video modeling, and online packages; what changes is coaching time.

03Which one, and when

Reach for MSWO when the referral question needs a complete rank order and one trial series is worth more than PSPA's pairwise granularity. Because a selected item is withdrawn rather than returned, one pass yields a full hierarchy without the combinatorial trial count a full paired-stimulus array requires as the item pool grows, a full ranking on a timeline closer to a single-stimulus sweep than a PSPA session.

The prerequisite is a reliable, discriminated selection response across a multi-item array, not just a two-item choice. A learner who reaches, points, or indicates a clear selection among several visible options can run MSWO; a learner whose response only holds up across two items, or who doesn't have one yet, needs single-stimulus or a non-choice format instead. Confirm that repertoire before building the array, not mid-session.

Weigh the without-replacement mechanic itself as a clinical risk factor, not just a scoring convenience. Both paired-stimulus and MSWO produced higher rates of problem behavior than a free-operant assessment for individuals with intellectual and developmental disabilities whose problem behavior is maintained by access to tangibles (Tung et al., 2017). MSWO's removal step is not incidental to that finding, it is the mechanic the finding is about: every selected item disappears from the array, a predictable evocative event every trial for a tangible-maintained profile. If intake data already flags that profile, route to free operant before the first MSWO session rather than running one cautiously.

Match the array's stimulus modality to the learner's tacting and identification repertoire whenever the target isn't a plain physical object. A picture-based MSWO produced valid hierarchies most reliably for learners who could match, identify, and tact the pictures in the array; learners who couldn't do that generated more valid data from a format sampling the interaction directly (Morris et al., 2020). The same logic extends to verbal stimuli: an MSWO-style ranking of conversation topics sits on a continuum with faster self-report and slower, more certain response-restriction formats, so for non-object stimuli the real question is where on that continuum a given case can afford to land, with MSWO as one point on it rather than the only option (Kronfli et al., 2024).

Decide up front whether to array items from a single stimulus class or mix classes in one pass. Combining edible and social-interaction items in one array, rather than running two separate arrays, is how displacement between the two classes gets detected at all; three of six typically developing preschoolers showed one class crowding out selections from the other once both were available together (Lasinski et al., 2024). That result is from a typically developing sample rather than the autism and intellectual-disability caseload most of this literature addresses, so treat it as a reason to check for displacement, not a rate to import directly.

Budget for staff training as part of the format decision, not an afterthought once the array is built, since scoring accuracy is the mechanic MSWO depends on most. A self-instructional online package with no in-person instructor and a live-coached BST package have both moved inconsistent administrators to accurate, generalized implementation (Arnal Wishnowski et al., 2018; Yang, 2022), so pick based on available coaching bandwidth, not an assumption that the unsupervised route is second-best.

Extending MSWO outside the autism and intellectual-disability population it was built around is defensible but should be checked, not assumed. In a general-education elementary sample, MSWO-selected and teacher-ranked rewards produced no measurable difference for any of four participants, and only half the sample showed either reward outperforming no reward (Resetar et al., 2008). Read that as a faster indirect ranking possibly being adequate in a general-ed setting, not as confirmation that MSWO itself underperforms there.

04What this means Monday morning

Set the array before the learner sits down. Space five to seven items in a line, far enough apart that reaching for one doesn't require passing over another, and randomize position across trials so the same edge or corner isn't winning by proximity. Define the selection response in advance, touch, pick up, point to, so a second scorer watching the same trial marks the same item you did.

Decide how you're presenting item magnitude before the first trial and hold that decision constant across the array. A comparison of uniform-mass presentation against a caregiver-report-consistent presentation, where portion matched how the item is normally delivered, found both methods identified the same top three items in three of five cases; disagreement concentrated in the lower-preference tail, not the top ranks (Moore et al., 2017). Don't switch magnitude presentation mid-array chasing a cleaner middle-of-the-pack ranking, and don't over-read a reshuffled fourth or fifth item as a real preference change when it may just reflect how much of that item was on the table.

Run the trial-and-remove cycle to completion, or until the learner stops responding, then re-space and present again. Score every trial, including the last item or two nobody actively chose, a rank at the bottom still carries information, what the learner tolerated least, not just what they liked most. Repeat the full pass two to three times and average ranks rather than trusting a single trial series.

When time is genuinely short, a web-based brief administration is a legitimate substitute for the live tabletop version, not a watered-down one. A tool built around the same without-replacement logic identified video preference hierarchies for young children in a compressed format and still carried predictive value for five of seven participants against a separate reinforcer test (Curiel et al., 2024). That two-of-seven miss rate is the practical lesson: brief administration is fast enough for a first pass, but the ranking still needs a confirmation step before it goes into a formal plan, the same as a live-array ranking.

Confirm the top item functions as a reinforcer before writing it into a program, especially with a young or newly assessed learner. A video-based version of this check found the top-ranked video maintained higher task engagement than the lowest-ranked item for two of three children tested, not all three (Curiel et al., 2024). A rank order is a strong candidate list, not a guarantee; a short contingent-delivery check on the top one or two items closes that gap faster than assuming the highest rank automatically works.

Re-run the array on a regular schedule rather than waiting for a plan to visibly fail. Expect the top two or three items to be the most stable part of the ranking across repeated administrations and the middle-to-lower tier to shift more; that instability is closer to expected format noise than a sign something went wrong with the session, so don't chase a perfectly matching re-test before trusting the top of a fresh ranking.

05From the experts

Are they going to be able to sit at a table and tolerate 150 trials of a PSPA? Or would it better to be a quicker MSWO? What items are they engaging with? If you want to contrive your free operant preference assessment a little bit more, one thing you could do is if you have perhaps a playroom or something like that, take the client in the playroom, let them engage, let them identify certain things they like.
From the talk — Matt Harrington 5 Days of Manding Mastery
It beats MSWO out of the water. And there's been a lot of cool research that I'll point to in a second. If you want to learn more about it, definitely something to check out. Secret word for our CEUs? Dark roast. Right? You see the pattern here? I am indeed obsessed with coffee. So, that's our secret word for this section. Dark roast. Write it down. Make sure you don't forget it. So, let's move forward. Let's look a little bit about what the concurrent chains arrangement is.
From the talk — Matt Harrington Analyzing Assent and Taking Data

06Common questions

Should I run MSWO or MSW (with replacement) for a given case?
Default to MSWO when you want a single-pass hierarchy and don't need repeated-selection frequency data. Because MSWO withdraws each selected item, every trial after the first is a forced choice among fewer options, producing a complete rank order in one series. Weigh that removal step against the case file: paired-stimulus and MSWO have both produced higher rates of problem behavior than free-operant assessment for a tangible-maintained profile, since MSWO removes a chosen item every trial. MSW keeps the array intact and tracks how often an item gets picked, useful when frequency is the question, but won't produce the same clean rank order in one pass.
What do I do with a trial where the learner doesn't select anything from the array?
Score it and move on rather than discarding or repeating the trial. A no-response trial still tells you the learner didn't find any remaining item worth selecting at that point, meaningful especially late in the array where only lower-preference items are left. If no-response trials show up early, that's a different signal worth checking before trusting the rest of the pass: confirm the selection response is still reliable and the session hasn't run into fatigue or satiation.
How many trials or sessions does it take before I can trust the rank order?
Run the full array-and-remove sequence two to three times and average, rather than relying on a single pass. Expect the top ranks to stabilize faster than the middle and bottom; a comparison of two presentation methods found agreement on the top three items in three of five cases, with disagreement concentrated in the lower-preference tail. A brief, tool-based version of MSWO has produced hierarchies with predictive value for most, though not all, participants against a reinforcer test, a reasonable first pass under time pressure that still benefits from a confirmation step.
Are app-based or tablet-administered MSWO tools as reliable as running the physical array by hand?
The evidence so far supports them as a legitimate substitute, particularly for speed and training new administrators, rather than a downgrade. A free web-based MSWO tool identified video preference hierarchies for young children with predictive value confirmed for most participants, and a self-instructional online package with no in-person coaching brought staff accuracy from roughly a quarter correct to the high 80s. Digital delivery changes how fast a hierarchy gets produced; it doesn't remove the requirement to confirm the top item functions as a reinforcer first.
My learner's MSWO ranking changes somewhat every time I re-run it, especially in the middle and lower positions. Is that a problem with my procedure?
Usually not, if the top two or three items stay roughly consistent. Rank instability in this format concentrates in the lower-preference tail: a comparison of two array-presentation methods found agreement on the top three items in three of five cases while disagreement clustered among less-preferred items. Treat a reshuffled middle-of-the-pack ranking as expected noise in an idiographic, repeated-measures format rather than evidence your array construction is off, and confirm the top-ranked item still functions as a reinforcer, that's worth re-checking, not the exact order below it.

07The studies behind this grade

The strongest 12 of 36 constituent studies. Each links to its record in the research database and its source.

  1. The Effects of Behavioral Skills Training on Staff Implementation of Multiple Stimulus without Replacement Preference Assessment
    Yang, 2022 · Journal of Higher Education Research Controlled
  2. A comparison of methods for assessing preference for social interactions
    Morris et al., 2020 · Journal of Applied Behavior Analysis Controlled
  3. The effects of video modeling on staff implementation of behavioral procedures in China
    Lim et al., 2020 · Behavioral Interventions Controlled
  4. Effects of computer-aided instruction on the implementation of the MSWO stimulus preference assessment
    Arnal Wishnowski et al., 2018 · Behavioral Interventions Controlled
  5. The effects of preference assessment type on problem behavior
    Tung et al., 2017 · Journal of Applied Behavior Analysis Controlled
  6. The Impact of Stimulus Presentation and Size on Preference
    Moore et al., 2017 · Behavior Analysis in Practice Controlled
  7. Evaluating preference assessments for use in the general education population.
    Resetar et al., 2008 · Journal of applied behavior analysis Controlled
  8. A Continuum of Methods for Assessing Preference for Conversation Topics
    Kronfli et al., 2024 · Behavior Analysis in Practice Suggestive
  9. Evaluating Artificial Intelligence on the Efficacy of Preference Assessments for Preservice Speech-Language Pathologists.
    Griffen et al., 2024 · Journal of Developmental and Physical Disabilities Suggestive
  10. The use of a preference assessment tool with young children diagnosed with autism
    Curiel et al., 2024 · Behavioral Interventions Suggestive
  11. Evaluating preference displacement of edible stimuli and social interactions for typically developing preschool children
    Lasinski et al., 2024 · Behavioral Interventions Suggestive
  12. The multiple-stimulus-without-replacement preference assessment tool and its predictive validity
    Curiel et al., 2024 · Journal of Applied Behavior Analysis Suggestive
Get the monthly evidence update. When new studies change this grade, we email the diff. Free, for BCBAs and RBTs.