Y-BOCS Scoring and Interpretation: What Yale-Brown Scores Mean
Published August 23, 2026 by Therapy Resource Clinical Team
Where the Y-BOCS came from
Before 1989, measuring OCD severity ran into a structural problem. Most rating scales counted symptoms, so a person with twelve minor rituals scored higher than a person with one ritual that ate six hours a day. Wayne Goodman and colleagues at Yale and Brown built the Yale-Brown Obsessive Compulsive Scale to pull severity apart from content. The ten severity items ask nothing about whether the obsessions concern contamination, harm, or symmetry. They ask how much of the person's life the symptoms take. Someone with scrupulosity and someone with relationship OCD can produce the same 22.
The instrument arrived as two companion papers in the November 1989 issue of Archives of General Psychiatry. Part I covered development, administration, and reliability (Goodman et al., 1989a). Part II covered validity, including data from a fluvoxamine trial showing the total score moved with drug-induced symptom change (Goodman et al., 1989b). Nearly four decades later it is still the primary outcome measure in OCD medication trials and in most exposure and response prevention trials. The 2010 second-edition paper describes the original as the gold standard measure of OCD symptom severity (Storch et al., 2010).
Administration is a semi-structured clinical interview in two parts. The clinician first works through a symptom checklist to establish what this particular person's obsessions and compulsions actually are, then rates the ten severity items against that specific picture. The checklist supplies the content; the ten items supply the number. That structure matters for anyone whose compulsions are entirely internal, since counting, silent reviewing, and mental reassurance all score on the compulsion items. See Pure O and mental compulsions for what those look like in an interview.
What the psychometrics show
In the original reliability study, four raters jointly rated 40 patients, and the scale showed excellent interrater agreement along with high internal consistency (Goodman et al., 1989a). Across 42 pretreatment patients, every item was endorsed frequently and spread across a range of severity. That property is what keeps a summed total meaningful instead of letting one or two items carry the whole score.
The validity paper examined 81 patients across three cohorts (Goodman et al., 1989b). The total score correlated significantly with two of three independent OCD measures and correlated weakly with depression and anxiety ratings in patients whose secondary depression was minimal. That discriminant finding carries practical weight. OCD and depression travel together often enough that a severity measure contaminated by mood would be useless for tracking a course of treatment.
Evidence has accumulated well since. A reliability generalization meta-analysis covering 144 studies found a mean alpha of .866 for the total scale, a mean test-retest correlation of .848, and a mean intraclass correlation of .922 (Lopez-Pina et al., 2015). Reliability dropped in nonclinical samples where the score range is restricted, which is one reason the scale performs poorly as a general-population screener.
Storch and colleagues published a second edition in 2010. The Y-BOCS-II widened the top of the scale to address a ceiling effect where the most impaired patients bunched near the maximum, folded avoidance into the severity ratings, and rebuilt the symptom checklist. Among 130 treatment-seeking adults, the severity scale showed an alpha of .89 with test-retest and interrater intraclass correlations above .85 (Storch et al., 2010). Adoption has been slow. Most trials and most clinics still run the original, so the original bands remain the reference point for interpreting a number someone hands you.
A self-report version exists and performs better than clinicians tend to expect. Steketee, Frost, and Bogart (1996) compared self-report against the interview across clinical and nonclinical samples and found the self-rated form had excellent internal consistency and test-retest reliability, in some respects exceeding the interview. Patients endorsed more symptoms on the checklist when writing it themselves than when asked out loud, which is worth remembering when a client's interview checklist looks thin.
How the scoring works
Ten items, each rated 0 to 4, for a total range of 0 to 40. Items 1 through 5 rate obsessions and yield an obsessions subscale of 0 to 20. Items 6 through 10 rate compulsions and yield a compulsions subscale of 0 to 20. The two subscales sum to the total, and it is the total that carries the severity interpretation.
The same five dimensions repeat in the same order on both halves: time occupied (items 1 and 6), interference with functioning (items 2 and 7), distress (items 3 and 8), resistance (items 4 and 9), and degree of control (items 5 and 10). Item 3 rates distress from obsessions, item 8 rates distress from compulsions. Once you know the pattern you can read a subscale breakdown without the form in front of you.
Anchors are behavioral, which is what keeps ratings comparable between clinicians. On the time items, a 1 means less than an hour a day, a 2 means one to three hours, a 3 means more than three and up to eight hours, and a 4 means more than eight hours or near-constant intrusion. On the interference items, a 1 is slight interference with overall performance intact, a 2 is definite interference that remains manageable, a 3 is substantial impairment, and a 4 is incapacitating.
Two dimensions are scored in the direction people find backwards. On resistance (items 4 and 9), a 0 means the person always makes an effort to resist and a 4 means they yield completely and willingly. On control (items 5 and 10), a 0 means complete control and a 4 means none. The rating window is the past seven days, so any score describes a week of someone's life rather than a stable trait.
Severity bands, one at a time
The conventional interpretation bands come from the developers and from the clinical trial practice built around the scale. They apply to the total score. Subscale scores of 0 to 20 have no separate band system, and reading them as half-severity produces nonsense.
A Y-BOCS score of 0 to 7 is subclinical, meaning obsessive thoughts and checking behaviors are present at the level most people experience them: a passing intrusive image, a second look at the stove. Symptom severity alone in this range does not describe a treatable disorder. This band also matters at the other end of treatment, because a post-treatment score here is the outcome a full ERP course aims for.
A Y-BOCS score of 12 falls in the mild range under the standard bands (8 to 15), meaning obsessions and compulsions take up less than an hour on most days and cause slight interference the person can usually absorb. People here often hold jobs comfortably and hide the rituals well. Brief outpatient ERP is typically enough at this level. Most medication trials would exclude them for sitting below the entry threshold, which says more about trial design than about who deserves treatment.
A Y-BOCS score of 18 falls in the moderate range (16 to 23), meaning obsessions and compulsions take up one to three hours a day and cause definite interference with work, school, or relationships that the person is still managing. Sixteen is the number that opens the door to most randomized trials. In an outpatient practice this band is the common presentation of someone who finally calls: employed, functioning, and losing two hours a day to checking, washing, or mental reviewing.
Impairment turns substantial in the next band. A Y-BOCS score of 27 falls in the severe range (24 to 31), meaning symptoms consume more than three hours a day and impairment in work or social functioning has moved past manageable. Jobs get lost around here. Family members are usually accommodating heavily by this point, so family accommodation work belongs in the plan alongside the exposure itself. Weekly sessions are frequently too thin a dose; twice-weekly or intensive ERP is the more realistic prescription.
The top band runs from 32 to 40. A Y-BOCS score of 34 falls in the extreme range, meaning obsessions and compulsions occupy more than eight hours a day or intrude nearly constantly, with incapacitating interference. People scoring here are often housebound, and rituals may have displaced sleeping and eating. The original scale has a documented ceiling problem in exactly this stretch: two people scoring 36 can look clinically very different, which is the limitation the second edition was built to fix (Storch et al., 2010).
These bands are convention. A large consortium put them to an empirical test and came out somewhere else. Cervin and Mataix-Cols (2022) pooled data across ages 5 to 82 and proposed unified benchmarks of 0 to 13 subclinical, 14 to 21 mild, 22 to 29 moderate, and 30 to 40 severe, noting that their mild threshold of 14 sits two points below the 16 most trials use and arguing that the 16 cutoff is arbitrary. Both schemes circulate now. When a score turns up in a chart or a referral letter, check which set the writer had in mind: a 16 reads as moderate under the Goodman bands and as mild under the consortium's.
Treatment response, remission, and relapse
Percentage change on the Y-BOCS is how OCD outcomes get reported, and the field settled the thresholds through a formal consensus process. Mataix-Cols and colleagues (2016) ran a Delphi survey of international OCD experts and published operational definitions in World Psychiatry that most trials now follow.
Response means a reduction of at least 35 percent from the baseline total, paired with a Clinical Global Impression Improvement rating of 1 (very much improved) or 2 (much improved), sustained for at least one week. Partial response means a reduction of at least 25 percent but under 35 percent with a CGI-I rating of 3 or better over the same window. Put in real numbers: a client who starts at 28 and finishes at 18 has dropped 36 percent and counts as a responder, while a finish at 21 is a 25 percent drop and counts as partial.
Remission is defined by where the score lands instead of how far it traveled. The consensus threshold is a total of 12 or less plus a CGI Severity rating of 1 (normal) or 2 (borderline ill), holding for at least a week. An alternative route is a structured diagnostic interview showing the person no longer meets criteria for OCD. Recovery is the same standard sustained for a full year.
Relapse carries its own definition, and it is the piece clinicians most often skip in a psychoeducation conversation. For someone who responded, relapse means losing the 35 percent reduction with a CGI-I rating of 6 or worse for at least a month. For someone in remission, it means scoring 13 or higher on the same terms. A single rough week fails the duration criterion by design. Saying that to a client in advance takes some of the terror out of their first flare-up.
Using the score during treatment
Re-administering the full interview weekly is impractical in most outpatient practices and largely unnecessary. A workable rhythm is a baseline at intake, a repeat around session 6 or 8, and a final rating at termination. Between those points, track hours per day and interference informally in session, because that is where movement shows up first anyway.
Expect the change to arrive unevenly. Over a course of ERP the time and interference items usually shift before distress does, since a person can stop washing well before handwashing stops feeling awful. When a client reports feeling no better while their time item has gone from a 3 to a 2, they are responding. Showing them the item-level numbers is often more convincing than anything you can say about it.
The resistance items deserve careful reading. Item 4 rates how hard the person tries to resist obsessions, and it is scored so that more effort earns a lower number. Success at resisting is a separate thing the item does not capture. A client who has improved and has deliberately stopped fighting intrusive thoughts can score worse on resistance while their total falls. Factor analytic work has repeatedly found the resistance and control items loading on a factor separate from the severity items, with the resistance items the weakest of the set (Deacon & Abramowitz, 2005). Read items 1, 2, 3, 6, 7, and 8 as the cleaner severity signal, and treat resistance and control as information about the client's stance toward their symptoms.
Control ratings connect straight to what exposure is teaching. A client who scores a 0 on control is reporting that they can suppress obsessions on demand, which is usually a report about how much effort suppression is costing. Under an inhibitory learning approach, the target is willingness to let a thought sit there unanswered, so the control rating can stay flat while functioning and distress both improve. Interpret a rising control score with the same caution as a rising resistance score.
Baseline severity also sets the starting rung. A client at 30 rarely tolerates a first exposure anywhere near the top of their list, so the fear ladder needs finer gradations and more steps at the bottom than the same ladder would need at 16. Worked OCD hierarchy examples show what that spacing looks like for contamination, checking, and harm themes, and the step-by-step method is laid out in how to build an exposure hierarchy.
For pacing across the course, a graded exposure guide gives the session-to-session structure. When the feared outcome cannot be arranged in real life, which is most of the time with harm and taboo themes, an imaginal exposure script carries the exposure instead. Both belong in a plan for anyone whose obsessions subscale runs high while their compulsions subscale looks deceptively low.
Why this page explains the Y-BOCS instead of hosting it
The Y-BOCS is copyrighted, and it is built for a trained clinician to administer as a semi-structured interview. The ratings depend on judgment calls a web form cannot make: deciding whether a behavior qualifies as a compulsion, probing a vague answer about hours, telling avoidance apart from a completed ritual, catching mental rituals a client does not think of as rituals. Reproducing the ten items here would hand people a number stripped of the interview that gives it meaning.
If you want your own score, the route is an assessment with a clinician who treats OCD. If a previous provider already gave you a number, this page exists so that number is legible to you.
For free self-administered screens of related symptoms, this site hosts scored versions of several no-permission-required instruments at /assessments, including the GAD-7 for generalized anxiety, the PHQ-9 for depression, and the PCL-5 for PTSD symptoms. None of them measures OCD. They are useful for the comorbidity picture, which in OCD is most often depression, and for tracking the mood and anxiety load riding alongside the obsessions.
Common questions
What is a normal Y-BOCS score? A normal Y-BOCS score is 0 to 7. That band is labeled subclinical and covers the ordinary intrusive thoughts and double-checking most people experience. General-population scores cluster low, which is also why the scale works poorly as a screener outside clinical settings.
What Y-BOCS score indicates OCD? A score of 16 or higher is the conventional threshold for clinically significant OCD and serves as the entry criterion for most randomized controlled trials. Diagnosis itself comes from a clinical interview against DSM-5-TR criteria, with the Y-BOCS layered on top as the severity measure. Someone can score 14 and clearly have the disorder, and the 2022 consortium benchmarks would call 14 mild rather than subclinical (Cervin & Mataix-Cols, 2022).
What does a Y-BOCS score of 20 mean? A 20 is moderate. Obsessions and compulsions occupy roughly one to three hours a day, interference with work or relationships is definite, and the person is generally still holding the pieces together. That is a standard starting point for outpatient ERP.
What does a Y-BOCS score of 25 or 30 mean? Both land in the severe band of 24 to 31. Symptoms run past three hours a day, impairment is substantial rather than manageable, and family members are usually accommodating heavily. Once-weekly sessions are often too thin at this level.
How much does a Y-BOCS score need to drop to count as treatment response? At least 35 percent from baseline, plus a clinician rating of much improved or very much improved, held for at least a week (Mataix-Cols et al., 2016). From a baseline of 24, that means reaching 15 or lower. A drop of 25 to 34 percent counts as a partial response.
What Y-BOCS score counts as remission? A total of 12 or less, plus a global severity rating of normal or borderline ill, sustained at least a week. The same standard held for a full year is called recovery. Scoring 13 or higher again with a clinician rating of worse for a month meets the relapse definition.
Can I take the Y-BOCS myself? A validated self-report version exists and correlates well with the interview (Steketee et al., 1996), but the instrument is copyrighted and the checklist portion needs a clinician's read to be scored consistently. Ask a therapist who treats OCD to administer it at intake and again around session 8.
Does the Y-BOCS work for Pure O? Yes. The compulsion items rate mental rituals on the same anchors as visible ones, so counting, silent reviewing, praying to neutralize, and reassurance-seeking all score. Pure O presentations often show a high obsessions subscale next to a compulsions subscale that looks low, and a low compulsions subscale in that pattern usually means the interview did not probe for mental compulsions.
What is the difference between the Y-BOCS and the Y-BOCS-II? The second edition rescaled the severity items to spread out the most impaired patients, folded avoidance into the ratings, and revised the symptom checklist (Storch et al., 2010). Its psychometrics are strong, and its adoption is limited. Scores from the two versions are not interchangeable, so confirm which edition produced any number you are comparing over time.
References
Cervin, M., & Mataix-Cols, D. (2022). Empirical severity benchmarks for obsessive-compulsive disorder across the lifespan. World Psychiatry, 21(2), 315-316.
Deacon, B. J., & Abramowitz, J. S. (2005). The Yale-Brown Obsessive Compulsive Scale: Factor analysis, construct validity, and suggestions for refinement. Journal of Anxiety Disorders, 19(5), 573-585.
Goodman, W. K., Price, L. H., Rasmussen, S. A., Mazure, C., Fleischmann, R. L., Hill, C. L., Heninger, G. R., & Charney, D. S. (1989a). The Yale-Brown Obsessive Compulsive Scale. I. Development, use, and reliability. Archives of General Psychiatry, 46(11), 1006-1011.
Goodman, W. K., Price, L. H., Rasmussen, S. A., Mazure, C., Delgado, P., Heninger, G. R., & Charney, D. S. (1989b). The Yale-Brown Obsessive Compulsive Scale. II. Validity. Archives of General Psychiatry, 46(11), 1012-1016.
Lopez-Pina, J. A., Sanchez-Meca, J., Lopez-Lopez, J. A., Marin-Martinez, F., Nunez-Nunez, R. M., Rosa-Alcazar, A. I., Gomez-Conesa, A., & Ferrer-Requena, J. (2015). The Yale-Brown Obsessive Compulsive Scale: A reliability generalization meta-analysis. Assessment, 22(5), 619-628.
Mataix-Cols, D., Fernandez de la Cruz, L., Nordsletten, A. E., Lenhard, F., Isomura, K., & Simpson, H. B. (2016). Towards an international expert consensus for defining treatment response, remission, recovery and relapse in obsessive-compulsive disorder. World Psychiatry, 15(1), 80-81.
Steketee, G., Frost, R., & Bogart, K. (1996). The Yale-Brown Obsessive Compulsive Scale: Interview versus self-report. Behaviour Research and Therapy, 34(8), 675-684.
Storch, E. A., Rasmussen, S. A., Price, L. H., Larson, M. J., Murphy, T. K., & Goodman, W. K. (2010). Development and psychometric evaluation of the Yale-Brown Obsessive-Compulsive Scale, Second Edition. Psychological Assessment, 22(2), 223-232.
Related Resources
This article is for informational purposes only and is not a substitute for professional mental health care. If you are in crisis, contact 988 Suicide & Crisis Lifeline or call 911.