Why the questions feel odd
The single most important design choice is that the check never asks you to rate the thing directly. Ask someone to score their team’s psychological safety out of a hundred and almost everyone lands around 85. Ego and social desirability do the answering. Ask instead, when someone makes a mistake here, does it tend to get held against them, and you get the truth, because it doesn’t feel like a judgement of the person answering.
Three moves make that work. Behavioural, not evaluative: what happens, not how good it is. Recent and specific: “in the last two weeks” beats “generally”, because the memory of a real event resists spin. And reverse scoring: about a third of the items are worded so that agreeing is the unflattering answer, which breaks the autopilot of nodding along.
There are 32 items across the six conversations, safety and engagement, and each one maps to a single facet of a single construct, so the result is specific enough to act on. In the team version, items written from the leader’s seat are reframed to the team member’s seat, so both of you are reading the same team through comparable questions.
How it’s scored
Every answer becomes a number from 0 to 100. Nobody ever sees an item as a score, only as a question. Each area is the average of its items and lands in one of three bands: By design (70 and up), Taking shape (50 to 69), or By default (under 50). Culture strength is the average of the six conversations. Safety and engagement are reported separately, as the foundation and the outcome.
Safety also acts as a gate. When it’s thin, the report says the whole read is probably optimistic, and it shapes what you’re told to do first. The priority itself follows a simple default: fix the foundation if safety is thin, otherwise start with the lowest conversation, because that’s usually where the most is to be gained. It’s a starting point, not a rule. Building from a strength is a perfectly good call when it carries more leverage.
With your team
On your own, you get an informed read. Two of the three gaps, perception and spread, don’t exist until your team answers as well. When they do, they take the same read anonymously. Individual answers are never shown, and results only unlock once at least four people have responded, so nobody can be picked out. The team score is the average of its members. The perception gap is your score minus theirs, area by area. The spread is how far apart your people are, which is the operational measure of climate strength.