The CES-D: Scoring, the 16-Point Cutoff, and Clinical Use
Published August 20, 2026 by Therapy Resource Clinical Team
Development and history
The Center for Epidemiologic Studies Depression Scale (CES-D) was developed by Lenore Radloff at the National Institute of Mental Health and published in 1977 in Applied Psychological Measurement. Its purpose was population measurement. NIMH needed a short self-report scale that could estimate depressive symptom levels across whole communities, administered by survey interviewers with no clinical training.
Radloff assembled the twenty items from existing validated scales, selecting for the components that appeared consistently in the depression literature of the period: depressed mood, feelings of guilt and worthlessness, helplessness and hopelessness, psychomotor retardation, appetite loss, and sleep disturbance. The wording is deliberately plain. It reads as slightly dated now, which is one of the few honest complaints about the instrument.
The CES-D is in the public domain and has been translated into dozens of languages. It is one of the most used depression measures in research history, and that history gives it something newer scales cannot match: decades of epidemiologic data in the same metric, from cohorts followed for years.
It remains in wide clinical and research use today. You can take a free CES-D screening here, which handles the reverse-scored items automatically and returns a severity band with the total.
Psychometric properties
Internal consistency is high. Cronbach's alpha runs about .85 in community samples and about .90 in clinical samples (Radloff, 1977). The scale holds together across age, sex, and education groups, which was a design requirement for an instrument meant to be handed to anyone who opened the door.
The 2016 individual-participant meta-analysis by Vilagut, Forero, Barbaglia, and Alonso, published in PLOS ONE, pooled data across studies and reported sensitivity of 0.87 and specificity of 0.70 at the classic cutoff of 16 for major depression. A screen should miss few cases, and at 0.87 the CES-D misses few.
Specificity of 0.70 is the number that shapes practice. Roughly three in ten people without major depression will still score 16 or above, and in a general population where the base rate of major depression is low, most positives are false positives. Higher cutoffs, commonly 20 or more, get used when a false positive is expensive: when a positive screen triggers a full diagnostic workup, or a referral queue with a three-month wait.
Scoring and interpretation
The CES-D has twenty items and asks about the past week, phrased as how often each experience occurred. The four response options carry explicit day counts, which is unusual and genuinely useful: Rarely or none of the time (less than 1 day) scores 0; Some or a little of the time (1-2 days) scores 1; Occasionally or a moderate amount of time (3-4 days) scores 2; Most or all of the time (5-7 days) scores 3.
Total range is 0 to 60. Four of the twenty items are reverse scored, which the next section covers in detail, because that step is where most hand-scoring goes wrong.
A total of 16 or above is the classic cutoff for clinically significant depressive symptoms. Radloff set it in the original paper and it has survived nearly fifty years of use, which says more about its practical utility than about its precision.
This site's tool sorts the total into four bands: 0 to 15 is Subthreshold, 16 to 20 is Possible Depression, 21 to 30 is Probable Depression, and 31 to 60 is High Likelihood.
One structural feature deserves flagging. The CES-D contains no item about suicidal thoughts or self-harm. The PHQ-9 has Item 9, and clinicians who move between the two instruments sometimes assume the CES-D covers the same ground. It does not, so risk has to be asked directly regardless of the total. If you or someone you know is in crisis, contact the 988 Suicide and Crisis Lifeline by calling or texting 988.
The CES-D measures symptom level. It does not diagnose. A total above 16 indicates that a diagnostic conversation is warranted; major depressive disorder is established through a clinical interview covering symptom count, the two-week duration criterion, functional impairment, and the medical or substance-related causes that can produce an identical profile.
The four reverse-scored items
Sixteen of the twenty items are worded so that endorsing them signals distress ("I felt sad", "My sleep was restless"). Four run the other direction and are reverse scored: item 4 ("I felt I was just as good as other people"), item 8 ("I felt hopeful about the future"), item 12 ("I was happy"), and item 16 ("I enjoyed life").
On these four the scale runs backward. Most or all of the time scores 0, Occasionally scores 1, Some or a little scores 2, and Rarely or none scores 3. A person who was happy most of the week contributes 0 points from item 12. A person who was almost never happy contributes 3.
Radloff included them for two reasons. Positive-affect items pick up anhedonia, which straight symptom-count items can miss, and mixing item direction disrupts the acquiescent response set, the tendency to keep agreeing once a pattern of agreeing has been established halfway down a long list.
Scoring these four straight through with the rest is the most common CES-D error, and the distortion it produces is large. Someone with no depressive symptoms who answers "most or all of the time" to all four positive items should contribute 0 points from them; scored straight, they contribute 12, which can push a total of 4 up to 16 and manufacture a positive screen out of nothing. The error runs the other way too. A depressed client who rarely felt happy or hopeful loses up to 12 legitimate points, and a real 20 gets recorded as an 8. When a CES-D total contradicts the clinical picture by a wide margin, re-score items 4, 8, 12, and 16 before doing anything else.
Clinical applications
Intake screening: the CES-D works as a baseline depression measure at intake, particularly in settings that already run it for research or program evaluation and want one metric across both purposes. Most adults finish it in three to five minutes.
Treatment monitoring: the one-week window makes frequent re-administration defensible in a way the AUDIT's one-year window never allows. Weekly administration produces clean non-overlapping windows, and every two to four weeks is the common clinical cadence. Consistency matters more than frequency here, since a score collected at the same point in the treatment week is the only kind that compares cleanly to the last one.
Meaningful change: the CES-D has no single agreed reliable change index of the sort the PHQ-9 has in its five-point rule. Most clinicians treat a drop below 16 as the response target and read movement across the bands as the meaningful unit. A client going from 34 to 22 has crossed two bands, and that trajectory carries real information even though both totals remain above threshold.
Measurement-based care: collecting a symptom score on a schedule and letting it inform decisions about session frequency, modality, and referral outperforms treatment as usual. The choice of instrument matters less than the discipline of collecting it on time and looking at the trend line with the client in the room.
Severity bands in detail: what each score range means
0-15 (Subthreshold): Depressive symptoms are absent or below the level associated with clinically significant depression. In monitoring, a total that has fallen into this range from a higher band is the standard marker of response. Context does the interpreting: a score of 8 in someone who scored 30 two months ago tells a very different story than the same 8 at intake.
16-20 (Possible Depression): This band opens at the classic cutoff. The person is reporting depressive symptoms at a level that warrants a diagnostic conversation, though the modest specificity of the instrument means a meaningful share of totals in this range will not correspond to major depression on interview. Typical next steps are a structured diagnostic interview, attention to duration and functional impairment, and a repeat administration in two to four weeks to see which direction the number is heading.
21-30 (Probable Depression): Symptom burden is substantial and spread across domains, since reaching the low twenties generally requires endorsing symptoms at the 3-4 day level or higher on many items. Active treatment is usually indicated. Evidence-based psychotherapy, medication, or the combination are first-line considerations, chosen with attention to history, prior response, and what the client actually wants.
31-60 (High Likelihood): Totals above 30 indicate severe symptom burden reported on nearly every day of the past week. Combined treatment and shorter follow-up intervals are typical, along with direct assessment of safety, functioning, sleep, and whether a higher level of care is warranted. Remember that the CES-D itself asks nothing about suicidal thinking, so that question has to come from the clinician.
What does a CES-D score of 22 mean? A score of 22 falls in the Probable Depression band, which runs from 21 to 30, and sits six points above the classic cutoff of 16. It indicates depressive symptoms occurring on most days of the past week across several symptom domains, at a level where active treatment is generally the recommendation. The number establishes no diagnosis on its own; major depression requires two weeks of symptoms, and the CES-D asks about one.
Is a CES-D score of 16 bad? A 16 is the lowest total that counts as a positive screen, sitting exactly at the cutoff where pooled sensitivity is 0.87 and specificity is 0.70 (Vilagut et al., 2016). It signals that a diagnostic conversation is warranted. Because specificity at that threshold is modest, a 16 reads better as an invitation to look closer than as evidence of a disorder, and clinics that want fewer false positives raise their working cutoff to 20 or above.
Common scoring questions
How often should the CES-D be re-administered? The items reference the past week, so weekly administration produces clean non-overlapping windows and no double counting. In routine practice, every two to four weeks during active treatment is common, dropping to monthly or quarterly in maintenance. Giving it at the same point in the week each time keeps the comparison honest, since a Monday score and a Friday score can differ for reasons that have nothing to do with treatment.
Does a score above 16 mean a person has depression? No. The CES-D is a symptom-severity screen, and specificity of 0.70 at the 16 cutoff means a substantial fraction of positives will not meet criteria on interview. Major depressive disorder requires a two-week duration, a specific symptom count, functional impairment, and the exclusion of medical and substance-related causes. A positive CES-D starts that assessment. It does not finish it.
How does the CES-D compare with the PHQ-9? They cover overlapping ground by different logic. The CES-D runs twenty items to the PHQ-9's nine, asks about the past week instead of the past two weeks, and includes four positive-affect items the PHQ-9 has no equivalent for. The PHQ-9's items map one to one onto the DSM criteria for major depressive disorder, which makes a positive screen easy to translate into a diagnostic conversation; the CES-D's items come out of the epidemiologic research tradition and were selected for population measurement. For brief clinical monitoring aligned to DSM criteria and a two-week window, the PHQ-9 is usually the more practical pick. For continuity with a research protocol or a long-running cohort, the CES-D has the history.
What if the score is low and the client looks depressed? Check the reverse-scored items first, since an error on items 4, 8, 12, and 16 can suppress a total by as much as 12 points. Then consider the one-week window: a client whose worst stretch ended three weeks ago can score low on a scale that asks only about the last seven days. Minimization, cultural differences in how distress gets described, literacy and language barriers, and the wish to look improved for a clinician they like will all pull a total down. The instrument supplements clinical judgment; it never replaces it.
References
Radloff, L. S. (1977). The CES-D Scale: A self-report depression scale for research in the general population. Applied Psychological Measurement, 1(3), 385-401.
Vilagut, G., Forero, C. G., Barbaglia, G., and Alonso, J. (2016). Screening for depression in the general population with the Center for Epidemiologic Studies Depression Scale (CES-D): A systematic review with meta-analysis. PLOS ONE, 11(5), e0155431.
Related Resources
This article is for informational purposes only and is not a substitute for professional mental health care. If you are in crisis, contact 988 Suicide & Crisis Lifeline or call 911.