Guide
How to evaluate team performance — and find the one thing to fix first
Most teams are not short of feedback. They are short of a decision. This guide walks through a way of evaluating performance that ends with one constraint worth attention now, rather than a list of nine things that all look urgent.
Start with a result, not a rating
A useful evaluation begins with one result that is not moving: a resolution time stuck at two days, a pipeline that converts at half the expected rate, a launch date that slips every quarter. Ratings of people describe symptoms. A stalled result gives you something you can test a fix against.
Write down three things before you evaluate anything:
- The result as it stands today, in a number or an observable fact.
- The result you want, and by when.
- What has already been tried, and what the evidence says happened.
Why most performance reviews find the wrong problem
Four failure patterns account for most wasted effort:
- Averaging everything. A survey with twenty questions returns twenty mediocre scores and no direction.
- Treating look-alikes as the same thing. A team without the tools to do the work looks exactly like a team without the skill, and exactly like a team that has quietly stopped caring. The three need opposite responses.
- Fixing the loudest complaint. The most vocal problem is rarely the binding one.
- Adding instead of removing. More people, more meetings, more targets. If the constraint is coordination, adding load makes it worse.
Step one: scan wide across four forces
Before going deep, check which part of the system is weakest. Every result depends on four forces working together:
- Offer — is what the team is being asked to deliver clear, valuable and coherent?
- Actors — do the people involved have the capability, motive and authority to deliver it?
- Context — do the tools, time, information and environment allow it?
- Binding — is anything holding the commitment in place once attention moves elsewhere?
Rate each force honestly against evidence, not intent. The lowest force is not your answer — it is your direction for the second step.
Step two: go deep on the weakest force
Inside the weakest force, separate the look-alikes. Ask whether the problem is the aim (nobody agrees what good looks like), the will (people can but won't), the skill (people would but can't yet), the means (the work is impossible with what they have), the sync (individually fine, collectively out of step), the grip (nothing holds the change after week two), or the loop (nobody sees the result of their own work in time to correct it).
One of these will score lowest. That is your constraint. Name the second-lowest too: if the two are close, treat the diagnosis as a hypothesis and gather more evidence before spending money.
Step three: turn the constraint into one next step
A good evaluation output fits on one page and contains four things:
- The constraint, stated plainly.
- Two or three concrete actions that address that constraint specifically.
- One thing to stop doing, to free the capacity the actions need.
- One observable signal, with a review date, that tells you whether it is moving.
If the signal has not moved by the review date, do not push harder. Return to the evidence and diagnose again — a wrong diagnosis pursued with discipline is still wrong.
A short worked example
A support team's resolution time stays at 51 hours after the team doubles in size. The wide scan shows Context as the weakest force. The deep dive puts Loop lowest: agents never learn which of their fixes actually stopped the customer coming back, so the same faults are re-solved every week. The next step is not more agents — it is a weekly repeat-contact review with the product team, stopping the daily volume stand-up to make room for it, with repeat contacts as the signal.
Do it now, in about ten minutes
The free diagnostic on this site runs both steps for you. You describe the stalled result, answer short rating questions, and get the top three constraints, the actions for each, what to stop, and the signal to watch. The scoring is plain arithmetic — the same answers always produce the same diagnosis, so you can re-run it with colleagues and compare where you disagree.