Hamilton Depression Rating Scale scoring sums 17 clinician-rated items into one number between 0 and 52, and that total sets the severity band. Scores of 0-7 read as normal, 8-13 as mild, 14-18 as moderate, 19-22 as severe, and 23 or above as very severe.
The HAM-D is not a questionnaire the patient fills in. A trained clinician runs a 15 to 20 minute interview, rates each domain against written anchors, and adds the items up. Nine of those items are scored 0-4 and the other eight are scored 0-2.
This guide walks through the items, the interview, and the severity bands. It then compares the HAM-D with the PHQ-9 and the Beck Depression Inventory. A scoring template is below, ready to download.
Key takeaways
Hamilton Depression Rating Scale scoring sums 17 clinician-rated items into a total between 0 and 52.
Severity bands run 0-7 normal, 8-13 mild, 14-18 moderate, 19-22 severe, and 23 or above very severe.
Nine items are scored 0-4 and eight are scored 0-2, so mood and psychomotor domains carry most of the total.
The HAM-D needs a trained interviewer and takes 15 to 20 minutes, which rules it out as a self-report screen.
Practice management software like Pabau stores each HAM-D total against the client record, so severity changes are visible at the next visit.
Download your free Hamilton Depression Rating Scale scoring template
A ready-made HAM-D-17 scoring sheet listing all 17 items with their 0-4 and 0-2 ranges. It also carries the five severity bands and space to record the total and rater at each visit.
Download templateWhat is the Hamilton Depression Rating Scale?
The Hamilton Depression Rating Scale (HAM-D or HDRS) is a clinician-administered instrument that measures how severe a patient’s depressive symptoms are. Max Hamilton published it in 1960 in the Journal of Neurology, Neurosurgery, and Psychiatry. The 17-item version, HAM-D-17, became the default in clinical research and psychiatric practice.
Why it matters: the HAM-D asks for clinician judgment rather than patient self-report. During the interview you watch the patient’s mood, listen for guilt and suicidal ideation, and rate each domain against an anchor. Observation catches presentations a checklist tends to flatten, such as psychomotor retardation.
- Type: clinician-administered rating scale
- Items: 17 domains, covering mood, guilt, suicide risk, insomnia, work impairment, psychomotor changes and somatic symptoms
- Time to complete: 15 to 20 minutes per patient
- Scoring range: 0-52, where a higher score means more severe depression
- Primary use: baseline assessment, treatment monitoring, clinical trials and research studies
The HAM-D has been used in thousands of published clinical trials, and it remains the most frequently cited depression rating scale in psychiatric research. That longevity reflects both its clinical usefulness and its psychometric track record.
How the total score maps to severity bands
The total is the sum of all 17 items, each scored 0-2 or 0-4 depending on the domain. That total maps onto severity bands that guide treatment intensity. The bands also belong in the written assessment, which the psychiatric evaluation template lays out in full.
Key point: clinical trials commonly define treatment response as a drop of 50% or more from baseline. A patient who scores 24 at baseline and 12 at follow-up has moved from very severe to mild. That shift signals the current treatment is working.
What each of the 17 items measures
Each item targets a distinct symptom domain. Knowing what an item is meant to capture helps you ask the right clarifying question. It also keeps your use of the anchors consistent from patient to patient.
Items 1-3, 7-11 and 15 are scored 0-4, while items 4-6, 12-14 and 16-17 are scored 0-2. That uneven weighting comes from the original scale, and it decides where the 52 available points actually sit.

How to run the interview and score it
The HAM-D is not a self-report tool. You run a structured or semi-structured interview, observe the patient, and rate each item against explicit anchors. Recording those ratings in psychiatry EMR software keeps the anchors, the item scores and the totals together across visits.
- Allow 15 to 20 minutes: book a dedicated assessment slot for baseline evaluations so the interview is not rushed.
- Create a safe environment: use a private, quiet room and build rapport before you reach guilt and suicidality.
- Open on mood: ask “how has your mood been over the past week?” and listen for spontaneous reports of sadness, hopelessness or anhedonia.
- Probe each domain in order: work through items 1 to 17 using the manual’s prompts. Record observed psychomotor change, sleep disturbance, appetite change and somatic complaints.
- Rate against the anchors: each item carries a 0-4 or 0-2 scale with descriptive anchors. Item 1 (depressed mood) runs from 0 = absent to 4 = communicates the mood almost entirely without prompting. The middle points cover reporting it only on questioning, reporting it spontaneously, and communicating it non-verbally.
- Sum the items: add all 17 scores and check the arithmetic before it goes in the chart.
- Record the score and the band: write the raw total and its severity band into the clinical note. Use the same field every time so the numbers stay comparable.
- Repeat at set intervals: weekly or monthly repeats show whether treatment is working and when to adjust it.
Common pitfall: rushing the interview or asking leading questions biases ratings downward. “You’re sleeping fine, aren’t you?” invites agreement. Ask open questions and let the patient describe the week before you settle on a rating.
HAM-D vs PHQ-9 vs Beck Depression Inventory: which scale to use?
The HAM-D is one of several validated depression instruments. Comparing it with the PHQ-9 and the Beck Depression Inventory shows which one suits your setting and your patient population.
Key differences: the HAM-D is clinician-administered and backed by thousands of published trials. Research still treats it as the reference instrument. It also costs training time. The PHQ-9 is quicker and dominates primary care. The Beck Depression Inventory leans on cognitive symptoms and suits therapy settings.
Many practices combine them. A PHQ-9 screen at intake flags who needs a fuller assessment, the HAM-D establishes the baseline, then a shorter instrument handles routine monitoring between reviews.
Interpreting the results in clinical practice
Once the total is calculated, the severity band points to the next clinical step. The five bands below carry different obligations, and the ones at the top of the scale carry the most urgent.
Scores 0-7 (normal): the patient sits below the threshold for a major depressive episode. Consider maintenance therapy if they were recently treated, or discharge with a relapse-prevention plan and crisis contacts.
Scores 8-13 (mild): a mild episode is present. Psychotherapy on its own, or combined with a low-dose antidepressant, may be appropriate depending on duration and functional impact.
Scores 14-18 (moderate): a moderate episode with noticeable impairment. Combined therapy and pharmacotherapy are usually indicated. Monitor suicidality and put a suicide safety plan in place if Item 3 scores 2 or more.
Scores 19-22 (severe): functional decline is likely. Urgent psychiatric consultation, medication review and intensive psychotherapy are warranted, and inpatient care may be. Agree a crisis protocol with the patient and, where appropriate, their family.
Scores 23 and above (very severe): treat this as an emergency. Arrange psychiatric evaluation the same day, run a formal suicide risk assessment, and consider hospitalization.
Limitations and pitfalls worth knowing
No single instrument does the whole job. The HAM-D has known weaknesses, and knowing them changes how you read a total.
- Interviewer variation: rater training and inter-rater reliability differ widely. Two clinicians can rate the same patient differently, and supervision reduces that spread without removing it.
- Somatic over-weighting: sleep, appetite, fatigue and sexual function all carry points. Medically ill patients, and those on medications that disturb sleep or appetite, can score high without a severe mood episode.
- Weak on atypical depression: hypersomnia, increased appetite and leaden paralysis score poorly here, so an impaired patient can look mild on the total.
- Time and cost: a 15 to 20 minute interview is clinician time. The scale cannot be automated or handed to the patient without losing validity.
- Item 17 (insight): rating a patient’s awareness of their illness is subjective, and it can record the clinician’s view rather than the patient’s.
- No recovery-side measure: the HAM-D scores symptom burden. Meaning, resilience and recovery goals sit outside it entirely.
Because of this, many practices pair the HAM-D with a self-report measure. Adding the patient’s own account of function gives the chart more than one view of the same week.
How Pabau keeps HAM-D scores on the client record
On paper, the HAM-D works well enough for a single visit. The rater completes the sheet during the interview, copies the total into a free-text note, and files the form. Reading a six-month trend then means paging back through every note.
Pabau, therapy practice management software, stores the HAM-D as a structured digital form instead. The total lands on the client record next to the appointment that produced it. Item-level answers stay attached, so a later reviewer can see which domains drove the score.
Repeat assessments then line up against each other, and a drop from 24 to 12 is visible without arithmetic. Pabau Scribe, our AI scribe, can draft the interview note around that total. The letter composer turns the same note into a referral or consultant letter.

Every Pabau subscription includes every feature, so structured assessment forms, secure storage and progress tracking are not held back for a higher tier. Onboarding is structured rather than self-serve, and your existing assessment history can be imported.
Keep every HAM-D score on the client record
Pabau stores HAM-D assessments as structured digital forms, tracks totals visit by visit, and keeps severity changes visible to the whole care team.
Conclusion
The HAM-D earns its place when every rater in the practice uses the same anchors. A move from 24 to 12 only means something if both numbers were produced the same way. Rater training, not the form itself, is what decides whether the instrument works.
The weighting is worth carrying into every reading. Sleep, appetite and weight sit in the 0-2 group, so a medically ill patient can total high without a severe mood episode. Read the item pattern alongside the total.
Download the template above, agree one set of anchors across your raters, and store each total where the next clinician will find it. Book a demo to see how Pabau keeps HAM-D scores and severity bands on the client record.
Continue your research
Need a faster screen before the full interview? The PHQ-9 and GAD-7 template gives you a two-minute self-report measure to run at intake.
Want to structure the interview itself? The psychiatric interview sets out how to sequence the questions that feed a rating scale.
Comparing depression instruments? The Major Depression Inventory maps its items directly to diagnostic criteria, which the HAM-D does not.
Tracking mood between appointments? A daily mood chart fills the weeks between formal assessments with the patient’s own record.
Frequently asked questions
What is Hamilton Depression Rating Scale scoring used for?
It measures how severe a patient’s depressive symptoms are, tracks treatment response over time, and serves as a primary outcome measure in psychiatric clinical trials. A clinician rates 17 symptom domains during a structured interview and sums them into one severity score.
Is the HAM-D self-administered or clinician-rated?
The HAM-D is clinician-rated. A trained mental health professional runs a structured or semi-structured interview, observes the patient, and scores each item against explicit anchors. It is not valid as a self-report tool.
What counts as a normal HAM-D score?
Totals of 0-7 indicate no depression, or recovery from an episode. Anything above 7 indicates mild or greater severity. A score of 8 or more typically meets the threshold for a depressive episode that warrants treatment consideration.
How often should you repeat the assessment to monitor treatment?
Frequency depends on the setting. Clinical trials often repeat the HAM-D weekly or every two weeks to catch change early. Routine practice usually repeats it monthly or quarterly. Shorten the interval during a crisis or a rapid treatment change.
Can the HAM-D be used in primary care?
It can, but it needs trained raters and 15 to 20 minutes per patient. Many primary care practices screen with the PHQ-9 instead, which takes two to three minutes, and reserve the HAM-D for complex cases or research-grade assessment.