Sales coaching is working when the behaviors you coached show up on the next calls, and the deal and revenue numbers follow a few weeks later. Measure it in that order: behavior first, deal signals second, outcomes last. Activity counts (calls made, coaching sessions held) tell you the coaching happened, not that it changed anything. Here's a measurement setup a sales manager can run without a data team.
Why activity metrics don't answer the question
Most coaching dashboards count inputs: one-on-ones held, calls reviewed, training modules completed. Those numbers go up whenever a manager is diligent, and they say nothing about whether a rep sells differently afterwards. A rep can sit through eight coaching sessions and still skip the budget question on every discovery call.
Win rate has the opposite problem. It is the number everyone wants, and it moves too slowly and for too many reasons to tell you whether coaching caused the change. A quarter of win rate data covers maybe a dozen closed deals per rep, with pricing, territory, and product changes all mixed in. If win rate is your only coaching metric, you will wait two quarters to learn something you could have seen in three weeks.
Measure in three layers
The layers sit in the order the effect travels:
- Behavior on calls. Did the rep do the thing you coached? Ask the budget question, handle the pricing objection with the approved response, book the next step before hanging up. This changes within days of a good coaching conversation.
- Deal signals. Did the deals those calls belong to get healthier? Economic buyer identified, decision process mapped, mutual plan with dates. This changes within a few weeks.
- Outcomes. Stage conversion, cycle length, ramp time for new hires, win rate. This changes over a quarter or two.
Coaching that is working shows movement at layer one first, then two, then three. If layer one never moves, stop looking at layers two and three, because nothing downstream will change.
Step 1: pick three behaviors per rep, not thirty
Every rep gets a coaching focus of three observable behaviors for the next six weeks. Observable means a manager could watch a call, or read its transcript, and mark yes or no. "Be more consultative" is not observable. "Asks about the decision process before proposing next steps" is.
Pick the three from the rep's actual calls, not from the methodology poster. Say a rep runs twelve discovery calls a week and quantifies business impact in two of them. That is the behavior to work on. Write it down as a line the rep agrees with: "Get a number on the cost of the problem in every discovery call."
Step 2: baseline before you coach
Score the rep's last fifteen to twenty calls on the three behaviors before the first coaching conversation. Without a baseline you are guessing at improvement, and reps know it. A baseline of "you asked about timeline on 4 of your last 18 discovery calls" makes the coaching concrete and gives you the number to beat.
This is the step that gets skipped, because scoring twenty calls by hand takes a manager most of a day. If you have call scorecards generated automatically from transcripts, the baseline takes ten minutes. If you don't, do it anyway for the three behaviors only, and accept that it costs a day once.
Step 3: score every call the same way
The rubric has to be the same for the baseline, the coaching window, and the follow-up, or you are comparing different things. Keep it to yes, partial, or no for each behavior, scored per call. Store it somewhere you can chart. A spreadsheet with one row per call is enough.
Resist adding behaviors mid-window. Managers see a new problem on a call and want to add it to the list, and six weeks later the rep has eleven focus areas and no progress on any of them. New problems go on the list for the next window.
Step 4: run a six to eight week window and read the trend
Six weeks is long enough for a behavior to move from "remembers when reminded" to "does it without thinking", and short enough that the manager stays engaged. Review the per-rep chart weekly. You are looking for the trend line, not any single week.
Three outcomes are possible, and each means something different. The behavior moves and holds: coaching worked, move to the next three. The behavior moves in the first two weeks and then drops back: the rep can do it and stopped, which is a motivation or workload problem rather than a skill problem. The behavior never moves: either the rep does not understand what you are asking for, or does not agree it matters. Ask the rep which it is. Both are fixable, but not with more of the same coaching.
Step 5: connect behavior to the numbers that matter to your CRO
Once you have two or three windows of behavior data, line it up against the lagging outcomes. Reps who moved on "quantifies business impact" should show better conversion from discovery to proposal a month or two later. New hires whose behaviors moved in their first ninety days should reach quota sooner than the previous cohort did.
Keep the comparison honest. You are not proving causation to a statistician; you are checking whether the story holds together. If behaviors moved and deal signals moved and outcomes did not, look at what else changed: pricing, territory, a competitor launch. If behaviors did not move but outcomes improved, do not credit the coaching.
A rhythm that fits a manager's week
Weekly, fifteen minutes: look at the behavior chart for each rep and pick one call to discuss in the one-on-one. Monthly, one hour: review deal signals by rep and check whether the behavior trends are showing up in pipeline health. Quarterly, half a day: line behavior and deal data up against outcomes, decide the next three behaviors per rep, and drop anything from the rubric that never predicted anything.
Where this goes wrong
- Measuring the coach instead of the coached. "Sessions held" measures the manager's effort. Behavior on the next call measures whether it landed. Track the second one.
- Too many behaviors. More than three per rep per window and nothing moves. If you cannot pick three, you have not diagnosed the rep yet.
- Coaching from memory. Managers remember the two calls they joined and coach to those. Score every call, or the sample is whatever the manager happened to see.
- Waiting for win rate. It is the last thing to move and the noisiest. Use it to confirm what the behavior data already showed, not to decide what to coach.
- No baseline. Improvement claims without a starting number turn into arguments in the one-on-one.
Where Aircover fits
Aircover scores every call against your rubric (MEDDPICC, SPICED, BANT, or your own) from a live transcript, so the baseline and the weekly behavior chart exist without a manager listening to recordings. It also coaches the behavior during the call, surfacing the discovery prompt or the approved objection response while the rep is talking, which is the shortest path from "coached" to "did it". The live coaching page shows how the in-call and post-call pieces fit together.