Preference Assessment in ABA: Formats, Decision Logic, and How to Pick Reinforcers That Actually Reinforce
Based on 126 experimental studies (21 controlled, 105 suggestive); 81% report positive effects; where reported, effects are predominantly large. Updated July 2026.
How we grade →01What the research shows
Across 126 experimental studies (21 controlled, 105 suggestive), 80% of the studies reporting a direction found positive effects. Where effect size was reported, effects were predominantly large.
Populations studied: autism, intellectual disability, developmental delay, mixed clinical.
Computed across 153 corpus articles (126 experimental, 27 contextual). Regenerated monthly as new studies are ingested.
02The variants, and how they differ
Stimulus preference vs. reinforcer assessment: what a ranking actually tells you
A stimulus-preference assessment identifies which items or activities a learner approaches, selects, or engages with most, producing a rank order or a hierarchy of high, moderate, and low preference. It does not, by itself, confirm that the top-ranked item will function as a reinforcer once made contingent on behavior. In practice the two correspond well enough, often enough, that most clinicians treat a validated preference assessment as a reasonable stand-in for a formal reinforcer assessment. That shortcut has real exceptions: preference and reinforcer value diverged between two structurally different formats in at least one direct comparison (Kodak et al., 2009), which is why "ran a preference assessment" and "confirmed a reinforcer" stay separate claims even when they usually land on the same item.
Single-stimulus and paired-stimulus (PSPA)
Single-stimulus presentation offers one item at a time and records approach or engagement. It is the simplest format to run and the most tolerant of learners who struggle with choice-making, but it produces the least differentiated ranking because nothing is ever placed in competition. Paired-stimulus presents two items at a time across all combinations and forces a selection, sharpening the rank order considerably, and it remains the closest thing the field has to a reference format. The tradeoff is session length, since a full combinatorial array grows quickly as the item pool grows, and every trial that ends in a selection typically requires removing the unselected item from view, a design choice with consequences (below). A newer iteration that presents each pair once instead of twice, a single-presentation PSPA, has held up against the traditional double-presentation version on accuracy while cutting session time roughly in half (MacNaul et al., 2024).
Multiple-stimulus without replacement (MSWO)
MSWO arrays several items at once, removes each selected item before the next trial, and repeats until every item has been chosen or the learner stops responding, producing a full rank order in a single pass rather than the many discrete pairs a PSPA requires. It sits between single-stimulus and paired-stimulus on time cost while producing a comparably differentiated rank. Its format mechanics, array size, replacement rules, holding item magnitude constant, matter enough to accuracy that a full treatment of running and scoring it belongs on its own page.
Multiple-stimulus with replacement (MSW)
MSW is structurally identical to MSWO except a selected item returns to the array before the next trial rather than being withdrawn, so the same item can be chosen across multiple trials and its selection frequency, not just its rank, becomes part of the picture. That difference matters in practice: MSW and free-operant formats have produced different top-ranked items for the same learner in direct comparison, disagreeing often enough that neither substitutes for the other without a confirmatory reinforcer probe when the format choice is close (Kodak et al., 2009).
Free-operant preference assessment
Free-operant preference assessment gives the learner continuous, concurrent access to every item in an enriched environment for an extended period and measures engagement duration rather than discrete selections, avoiding the removal step entirely. It tends to run calmer for learners whose challenging behavior is triggered by having a preferred item taken away, and it is the format most associated with lower problem behavior during the assessment itself for tangibly maintained presentations specifically (Tung et al., 2017). A response-restriction variant that adds brief, controlled removal periods back into an otherwise free-operant structure has outperformed standard paired-stimulus on reducing challenging behavior during identification without giving up ranking utility (Herbek et al., 2026). Full mechanics live on its own page; here it is enough to know it trades some of PSPA's discrimination sharpness for a meaningfully calmer session.
Concurrent-chains, duration-based, and eye-gaze formats
Concurrent-chains preference assessment adds an initial-link choice, which activity to sample, before a terminal-link engagement period, letting duration of engagement do the ranking work instead of a forced discrete selection. Head-to-head against paired-stimulus, it produced consistent, well-correlated rankings while running measurably faster (Basile et al., 2021). Duration-based scoring is worth reaching for whenever the target is naturally continuous, leisure items, break activities, rather than discrete. When the stimulus under assessment is a social interaction rather than an object, match the instrument to the learner's prerequisite repertoire: an MSWO-style ranking works when the learner can tact pictures of the interaction, a stimulus-interaction format substitutes when they can't, and a vocal-response paired format works for a vocal learner (Morris et al., 2020). Eye-gaze duration extends the same logic to learners without a reliable motor response at all, using fixation time in a paired array to identify reinforcers for individuals with severe multiple disabilities (Cannella-Malone et al., 2015).
Caregiver and indirect report
Indirect formats, caregiver interview, checklist, or a simple ranking exercise, trade direct observation for speed and are the natural starting point when no direct assessment has happened yet or time is genuinely too short for a discrete-trial format. They are not a compromise of last resort: teacher ranking has performed comparably to a full MSWO for selecting reinforcers in a general-education sample (Resetar et al., 2008), though that finding sits in a typically developing school population rather than the autism and intellectual-disability samples most of this literature draws from, and should be weighted accordingly on a clinical caseload. Pairing structured interview with adjustments to pacing and instruction clarity has also improved comfort and reduced confusion during assessment for adults with dementia, without changing which items came out on top (Bigwood et al., 2026).
03Which one, and when
The choice of format starts with what the learner's response repertoire actually supports, not with which format is fastest to run. A learner who reaches, points, or otherwise makes a clear, reliable selection response can run any choice-based format: single-stimulus, paired-stimulus, MSWO, MSW, concurrent-chains. A learner without a reliable motor selection response, whether from significant motor impairment or a very early skill repertoire, needs a format that doesn't require one. Eye-gaze-based paired-choice procedures have substituted successfully where a reach-and-select response wasn't reliably available (Cannella-Malone et al., 2015).
Time cost is the second axis, and it scales with how discriminated a ranking you actually need. If you need a full hierarchy before building a token economy, paired-stimulus or MSWO earns its session length. If you need one or two reliable reinforcers fast, a single-presentation PSPA variant halves the standard session time without losing accuracy (MacNaul et al., 2024), and concurrent-chains has produced comparably reliable rankings on a shorter timeline than traditional paired-stimulus (Basile et al., 2021). When even that is too slow, an indirect caregiver or teacher ranking is a legitimate starting point: it has held up against a full MSWO in at least one direct comparison, though that comparison ran in a typically developing, general-education sample rather than the autism or intellectual-disability caseload this page addresses, so confirm the ranking directly before you build on it (Resetar et al., 2008).
Problem-behavior risk during the assessment itself is the axis practitioners most often skip. Any format that removes a selected item from view, standard PSPA, MSWO, MSW, creates a moment where item removal is the antecedent, and for a learner whose challenging behavior is maintained in part by access to tangibles, that moment is a predictable evocative event inside the assessment itself, not only in treatment later. Free-operant formats sidestep the removal step and have produced fewer instances of problem behavior during the assessment for tangibly maintained presentations specifically (Tung et al., 2017), and a response-restriction hybrid that reintroduces brief, controlled removal periods into a free-operant structure has outperformed standard paired-stimulus on this exact measure (Herbek et al., 2026). If a case file already flags tangible-maintained problem behavior, start there rather than discovering the removal sensitivity mid-session.
A few secondary decisions are worth making deliberately. Decide up front whether social interaction belongs inside the assessment: for leisure and social items, adding brief social interaction, comments, praise, joint attention, during sampling can shift which item comes out on top relative to a solitary-only presentation (Kanaman et al., 2022). Hold stimulus size constant across items unless magnitude is the specific variable being tested; uneven size has shifted rank order among lower-tier items in at least one comparison, an easy confound to eliminate and an easy one to miss (Moore et al., 2017). When the item isn't a physical, handleable stimulus, a break, an activity, an environment, use a pictorial or symbolic representation rather than forcing a live-sample trial; a brief pictorial assessment has reliably identified break options that functioned as reinforcers for work behavior in autistic clients (Castelluccio et al., 2019).
Finally, treat correspondence between a preference ranking and an actual reinforcer as something to verify, not assume, whenever the format choice was close or the stakes are high. MSW and free-operant rankings have disagreed for the same learner often enough that a brief contingent-reinforcement probe on the top-ranked item is worth the extra minutes before committing a plan to the ranking (Kodak et al., 2009).
04What this means Monday morning
Run a stimulus-preference assessment before you build any reinforcement-based program, not after a first session reveals the token economy isn't working. For a new case with no known reinforcers, start with a brief indirect step, a caregiver interview or rapid ranking, to generate a candidate item pool, then confirm it with a direct assessment before finalizing the plan. Teacher and caregiver ranking has held up well enough against direct assessment to make this a legitimate first pass rather than a stopgap, though the direct comparison behind that came from a typically developing, general-education sample rather than the autism and intellectual-disability caseload here, which is exactly why the indirect step generates the candidate pool and the direct assessment confirms it (Resetar et al., 2008).
Pick the direct format based on what you already know about the learner, not a clinic default. If the intake flags tangible-maintained problem behavior, start with a free-operant or response-restriction free-operant format rather than paired-stimulus, since removal-triggered items are a predictable evocative event you can design around from session one (Tung et al., 2017; Herbek et al., 2026). If time is short, run a single-presentation PSPA or a concurrent-chains format rather than the full double-presentation array (MacNaul et al., 2024; Basile et al., 2021).
Build the array deliberately. Hold item size constant across items unless magnitude is the variable you're testing, since uneven size has shifted rankings among lower-tier items in at least one comparison (Moore et al., 2017). If the target reinforcer is a social or leisure activity, decide whether to include brief social interaction during sampling and stay consistent about it, since adding it has shifted which item ranks highest for leisure items specifically (Kanaman et al., 2022). If the item being assessed is a social interaction itself, match the instrument to the learner's tacting and vocal repertoire rather than defaulting to a standard MSWO (Morris et al., 2020). For a learner without a reliable selection response, move to duration-based free-operant scoring, or an eye-gaze-based format for a learner without a reliable motor response at all (Cannella-Malone et al., 2015).
Before you write the top-ranked item into a formal reinforcement plan, run a brief contingent-reinforcement check: deliver the item contingent on a simple, already-fluent response and confirm responding actually increases. MSW and free-operant assessments have identified different top items for the same learner in direct comparison, so a ranking alone isn't a guaranteed reinforcer, and the confirmation probe is cheap compared to building a plan around an item that doesn't function as one (Kodak et al., 2009).
Re-assess on a schedule, not only when a plan stops working. Preferences are not fixed properties of a learner; treat a stale preference list as a live hypothesis to retest at a regular cadence, most clinics land on every few months, or sooner for a young or rapidly changing learner, and immediately whenever responding to a previously reinforcing item flattens out or a new problem behavior appears during item removal. If you're assessing a population outside the usual autism and intellectual-disability caseload, dementia care is the clearest example in this evidence base, adjust pacing and instruction clarity rather than assuming the standard procedure transfers unchanged; slower pacing and clearer instructions have improved comfort and reduced confusion for adults with dementia without changing which items came out on top (Bigwood et al., 2026).
05From the experts
The first thing to do is start with a preference assessment, right? We need to know what the reinforcers are, and we need to know what the, well, I should say potential reinforcers are, because we're not running a reinforcer assessment. But on that note, do we need to run a reinforcer assessment to make sure that our preference assessment is lining up with potential reinforcers? I would argue no. There's been a lot of research that has corresponded preference assessments to reinforcer assessments.
Like sensory preference assessment. Why is it that behavior analysts say sensory is not a thing when everything about stimuli is sensory? Preference assessment should include sensory preferences. Values. Preference assessments should include values and reinforcers. Now, here's where the little bit of the difference is. The difference is in relation to masking. And I don't mean neurodivergent masking. I mean stimulus masking, the behavior analytic concept. Because we know that a reinforcer can be so powerful that it can interfere with access for things the individual needs. We know that.
I just saw. Yes. What motivates people? Private verbal. Yeah. Some companies do preference assessments. Yes. From what the reinforcement style question is. Absolutely. Yeah. Starting at the interview process. That's a really great point. We actually do that as well, even just to having nothing to do with employee recognition for our clinic specific employees. We want to make sure that our break room is stocked with the food and drinks they actually like, like things they actually want to eat. And guess what? People's preferences change.
06Common questions
- If an item ranks highest on a preference assessment, do I still need to run a reinforcer assessment before using it in a program?
- Treat the ranking as a strong candidate, not a confirmed reinforcer, especially when the format choice was close or the stakes are high. MSW and free-operant assessments have produced different top-ranked items for the same learner in direct comparison, so a brief contingent-reinforcement probe on the top item is worth the few extra minutes before building a full plan around it.
- My learner gets upset or engages in problem behavior when I remove an item during a paired-stimulus or MSWO trial. What should I switch to?
- Move to a free-operant or response-restriction free-operant format, especially if the case file already flags tangible-maintained problem behavior. Free-operant formats skip the removal step entirely and have produced fewer instances of problem behavior during the assessment for tangibly maintained presentations, and a response-restriction variant that reintroduces brief, controlled removal periods has outperformed standard paired-stimulus on this exact measure.
- Should I include praise or social interaction while I run a preference assessment for toys or leisure items?
- Decide up front and stay consistent about it across the array, because it changes the answer. Adding brief social interaction during sampling has shifted which leisure item ranks highest relative to a solitary-only presentation, so a solitary assessment may understate an item's value when the real reinforcement context includes adult attention.
- Does it matter if some items in my array are bigger or presented differently than others?
- Yes, hold size and presentation magnitude constant unless magnitude is the variable you're deliberately testing. Uneven item size has shifted rank order among lower-tier items in at least one comparison, a cheap confound to eliminate by standardizing the array, not a reason to distrust preference assessment as a method.
- Can I use these procedures with a population outside the usual autism and intellectual-disability caseload, like dementia care or a general-education classroom?
- The core logic transfers, but don't assume the standard child-oriented procedure is optimal unchanged. Teacher ranking has performed comparably to a full direct assessment in a general-education sample, and adjusting pacing and instruction clarity, without changing the item set, has improved comfort and reduced confusion for adults with dementia. Both findings sit outside the autism and intellectual-disability population most of this literature is built on, so treat them as a starting adaptation and tune delivery to the population in front of you.
07The studies behind this grade
The strongest 12 of 153 constituent studies. Each links to its record in the research database and its source.
- Making Preference Assessments More Acceptable and Effective for People with Dementia
- A Comparison of the Response-Restriction Free Operant and Paired-Stimulus Preference Assessments for Children Who Exhibit Challenging Behavior
- Evaluating two iterations of a paired stimulus preference assessment
- Evaluating the effects of social interaction on the results of preference assessments for leisure items
- Comparing paired-stimulus and multiple-stimulus concurrent-chains preference assessments: Consistency, correspondence, and efficiency
- A comparison of methods for assessing preference for social interactions
- Using stimulus preference assessments to identify preferred break environments
- The effects of preference assessment type on problem behavior
- The Impact of Stimulus Presentation and Size on Preference
- Using eye gaze to identify reinforcers for individuals with severe multiple disabilities.
- Comparing preference assessments: selection- versus duration-based preference assessment procedures.
- Evaluating preference assessments for use in the general education population.