Dune passed his evaluation. Forty minutes on the floor at Cedar & Co., three introductions, a handler with eleven years behind her, and a clean sheet at the end of it. Nine days later he put a hole in another dog's ear over a tennis ball that nobody had seen him claim. Marisol pulled the evaluation form out of the file and read it twice. Nothing on it was wrong. It just wasn't about the right thing.
That distance, between an assessment that was performed correctly and an outcome it entirely failed to anticipate, is one of the most common blind spots in pet care. It is not a staffing failure. It is a measurement failure, and it comes from asking one question when there are two.
A temperament evaluation asks what a dog did in one room on one morning. A dog personality test asks what a dog is generally like, across situations, over time. Both are legitimate. They answer different things, and most operators are only running one of them. This piece is about the second one: what the research actually supports, what it does not, and how to turn it into five lines on a pet record that your staff will read at 7am.
01 / Two instrumentsTwo tests, two different questions.
Start by separating them cleanly, because the search results for this topic mix them together badly.
A temperament test is a live observation. Staff expose the dog to a structured sequence of stimuli: handling and restraint, a novel surface underfoot, a sudden noise, brief separation from the owner, a resource presented and removed, then a controlled introduction to one or two settled dogs. Published protocols tend to run eight to twelve minutes of structured subtests, though plenty of facilities wrap it inside a three or four hour trial day. The output is a decision. Yes, yes with conditions, or no.
A personality test is a rating exercise. Somebody who already knows the dog answers questions about how it usually behaves, and the answers resolve into scores on a small number of stable traits. Two instruments dominate the field. The Dog Personality Questionnaire came out of Amanda Jones's doctoral work at the University of Texas at Austin, published in 2008. Her team pulled roughly 1,200 candidate descriptions from the existing literature, from shelter assessments and from trainers, behaviorists and veterinarians, then narrowed them down through repeated factor analysis across several thousand respondents to a 75-item form and a shorter 45-item form. The Canine Behavioral Assessment and Research Questionnaire, developed at the University of Pennsylvania, takes a related approach across fourteen behavioral domains and now holds records on tens of thousands of dogs across more than 300 breeds. A shortened version was validated in 2024 using participants from the Dog Aging Project.
| Temperament test | Personality test | |
|---|---|---|
| What it samples | One session, your building | Months of ordinary life |
| Who answers | Your staff, by observation | Owner or handler, by rating |
| Output | A decision | A description |
| Good at | Screening out the obvious | Predicting the ordinary |
| Bad at | Rare events | Anything urgent |
| Time cost | 8 to 40 min | 3 min, once |
Read that last row again. The personality profile is the cheaper of the two by a wide margin, and almost nobody runs it. That is the opportunity.
02 / The snapshot problemWhy one morning is a bad sample.
The strongest evidence against relying on a single evaluation comes from the shelter world, where the stakes are highest and the research is therefore best funded. In 2016, Patronek and Bradley published an analysis with the memorable title No better than flipping a coin. Their conclusion, reinforced by a broader 2019 review, was that the standard behavior evaluations in use showed little demonstrable inter-rater reliability, test-retest reliability or predictive validity, and that claims to the contrary generally rested on a statistical confusion between group-level correlation and accuracy for an individual animal.
The mechanism is base rates, and it is worth sitting with for a moment because it applies directly to your floor. Serious incidents are rare. When you test for a rare event with an imperfect instrument, most of the positives you generate are false ones. A test that flags 10% of dogs in a population where 2% will ever actually bite is producing four false alarms for every genuine one, even when it is working exactly as designed. That is not a flaw you can train out of your staff. It is arithmetic.
Two caveats, because the transfer is not perfect. Shelter dogs are assessed under acute stress, by strangers, with no behavioral history, and a bad result can end a life. Your intake is gentler on every one of those dimensions: the dog comes from a home, you have an owner in the lobby who can answer questions, and a “no” means a refund rather than a euthanasia decision. So do not read this as an argument for scrapping your evaluation. Read it as an argument against letting the evaluation be the only thing you know about a dog before you put it in a room with twenty others.
03 / The frameworkThe Five-Line Profile.
Here is the version that survives contact with a real business. Not a 45-item questionnaire, which no owner will complete in your lobby and no handler will ever read. Five lines on the pet record, one per validated trait, each scored 1 to 5.
- Fearfulness. How readily does this dog startle, retreat or freeze? Low scores are bombproof dogs. High scores need exits and slow introductions.
- Aggression toward people. Any history of growling, snapping or biting at humans, including over handling, food or space.
- Activity and excitability. Baseline arousal and how fast it climbs. High scores are not bad dogs. They are dogs with a room requirement.
- Responsiveness to training. Does a recall land when it matters? Can a handler interrupt a behavior mid-sequence?
- Aggression toward animals. History with other dogs specifically, separated from the human line because the two do not travel together.
Those five are not a list somebody made up in a staff meeting. They are the factors that came out of Jones's analysis as dimensions that vary independently of one another, which is the whole point. Operators who build their own forms almost always end up with twelve lines, and roughly seven of them are the same line asked twice. “Confidence” is fearfulness inverted. “Friendliness with people” is line two upside down. “Energy” is line three. Five is not a simplification of a longer list. Five is what is left when you stop double-counting.
A score with no observation attached to it is a number somebody invented. Every line takes a clause: “4, backs away from men in hats.” “2, recall lands about half the time outdoors.”
Three is your default. If a handler is scoring everything 3, they have not watched the dog yet, and you should say so kindly and ask again next week.
One more design decision, and it is the one that makes the whole thing work: every dog gets scored twice, by two different people.
04 / Two ratersThe gap between them is the data.
In 2017, a study in Applied Animal Behaviour Science did something unusually relevant to this industry. Researchers took 60 dogs and had each one rated independently by two informants: the dog's owner, and the dog's professional walker. Ten walkers covered the sample. Both raters completed two separate personality instruments.
The interesting result was not that the two questionnaires agreed with each other, though they did. It was the pattern in where owners and walkers agreed. Significant inter-rater reliability showed up on fearfulness, on aggression toward people, and on aggression toward animals. The trait with no meaningful consensus at all was motivation.
Sit with that for a second. The three dimensions where an owner and a professional handler, who see the dog in completely different contexts, independently arrive at the same answer are precisely the three that decide whether a dog is safe in a group. That has a direct operational consequence, and it runs against the instinct most of us developed the hard way.
Believe the owner on those three lines. When a client tells you at intake that their dog is nervous, or that it once snapped at a delivery driver, or that it had a scrap at the park last spring, the research says that report carries real signal. Our industry has trained itself to discount owner reports as either anxious exaggeration or wishful minimizing. On fearfulness and on the two aggression lines, that skepticism is not earned.
The other two lines are where it gets useful. Activity and excitability, and responsiveness to training, are strongly contextual. A dog that is placid in a quiet house with two adults is a materially different animal on a floor with twenty-two dogs and a door opening every nine minutes. When the owner scores activity a 2 and your handler scores it a 5, nobody is lying and nobody is wrong.
The owner isn't wrong and the handler isn't wrong. They are describing two different rooms.Marisol R., nine years running a daycare floor
05 / ApplicationRunning it on a Tuesday.
Back to Dune, because his record is the clearest example of the framework doing work.
At intake his owner scored him fearfulness 2, aggression toward people 1, activity 2, responsiveness to training 4, aggression toward animals 1. A friendly, trainable, fairly settled dog. Which, in a quiet house in Sellwood with one adult working from home, he genuinely is.
After three visits the handler scored the same five lines without looking at the owner column: fearfulness 4, aggression toward people 1, activity 5, responsiveness to training 3, aggression toward animals 2. The two aggression lines held almost exactly, as the research says they should. Activity moved three full points. Fearfulness moved two.
High arousal alongside moderate fear, with recall degrading under load, is the signature that precedes resource guarding in a crowded room. It is not a character flaw and it is not grounds for refusal. It is a set of conditions. Dune belongs in the six-dog morning group rather than the twenty-two-dog afternoon floor, and there should be no loose balls down while he is in.
None of that required a behaviorist. It required five numbers from two people and ninety seconds of reading. Had the profile existed nine days earlier, the ear would probably still be intact.
Total cost: about ninety seconds of owner time at intake and sixty seconds of handler time three weeks later. Set against one ear, one refund and one conversation with a client you had for two years, the trade is not close.
06 / Anti-patternsFour ways this goes wrong.
I have watched all four of these happen, twice in my own building.
- Turning the profile into a gate. The moment a 5 on any line starts meaning “refuse,” staff begin scoring defensively and the numbers stop describing dogs. Your temperament evaluation remains the gate. The profile only ever decides placement.
- Scoring once and filing it. Dogs change. Adolescence arrives around ten months and rearranges lines one and three. A house move, a new baby, an undiagnosed sore hip, a scare at the park: all of them move the profile. Re-score annually at minimum, and always after an incident.
- Letting one person fill in both columns. This is the failure that quietly destroys the whole method. If your handler completes the owner column from memory because the client was in a hurry, you no longer have two independent raters. You have one opinion written down twice, and the gap, which is the entire signal, drops to zero everywhere.
- Reading a 5 as a bad score. Only lines two and five have a direction where high is genuinely alarming. A 5 on activity is an excellent daycare dog placed in the right group. A 5 on training responsiveness is the dog you use to settle a new arrival. Train your staff to read the profile as a shape, not a total, because there is no total.
07 / TakeawaysTake this with you.
If you do one thing this week, add the five lines to your intake form. Not the questionnaire, not a new policy, not a staff training day. Five questions and a box for one clause each. The handler column can wait until you have a month of dogs to score against it.
The thing that surprised me most when we started doing this was how often the two columns agreed, and how loudly the few disagreements spoke. Most dogs are roughly what their owners say they are. The handful who are not are the handful who end up in your incident log, and they announce themselves three weeks ahead of time if you have written the numbers down.
Write to me and tell me what your gaps look like. I am collecting profiles across facility sizes and the pattern in the activity line is already more interesting than I expected.
– DR, Portland, between the morning group and the afternoon floor