Selected work

Three systems, and the operational problems behind them.

Each started as something broken on a floor. A quality program that only ever reviewed the calls that worked. Contact rates falling at every attempt with no evidence for why. Numbers that took an hour to defend in a meeting. All three ran in production on a Fortune 500 program, and the interfaces below are reconstructions built to the real design specs and filled with synthetic data.

Three tools, one mindset.

Quality · Caliber AI

Every call scored, not a sample.

An effective QA reviewer needs about two hours to score one hour of calls, and that capacity goes to validating sales. The blindspot it leaves is the shorter calls and the opportunities that never closed. This scores every call instead, and the scorer is built so it can't invent what it didn't see.

Node.js · Ollama · MSSQL · PHP
Caliber AIProgram A · Telecom
Quality — brand home

Program A / Outbound

8,412 calls scored · 214 agents · 19% sale rate
8,412Calls scoredrolling 30 days
96%Population coveragewas ~2% sampled
19%Sale ratedialer-derived · 1,598
214Agents with scored calls3 sites
31Auto-failsred flag · review queue
32%Compliance rateapplicable items only
Behaviors% of scored calls
Active engagement41%
Ownership52%
Effort reduction28%
Needs discovery17%
Solution framing31%
Customer appreciation58%
Effectively demonstratedPartially demonstratedDid not demonstrate

Needs discovery sits at 17%, the weakest behavior on the board, and every later step leans on it. It also carries the heaviest weight in the rubric — 15 of the 60 behavior points. An agent who never found the need has nothing to tie the offer to.

Complianceapplicable items
ItemYesNoN/A
Products positioned91%6%3%
Products explained84%11%5%
Rate disclosure62%31%7%
Overall bill changes44%9%47%
Monthly recurring charges78%14%8%
Non-recurring charges39%12%49%
Term and agreement71%18%11%

N/A means the item could not apply to that call, so it drops out and its weight renormalizes across the rest. Rate disclosure fails on 31% of the calls where it did apply, which puts it in regulatory territory rather than coaching.

Call-flow steps11 steps · full population · click a header to sort
#StepEffectivelyPartiallyDid not demonstrate
1Positive greeting74%21%5%
2Set the agenda66%24%10%
3Discover needs19%36%45%
4Link the need to a product27%41%32%
5Assume the sale48%29%23%
6Engage the objection22%33%45%
7Re-frame and re-close18%30%52%
8Confirm the order57%26%17%
9Deliver compliance63%22%15%
10Order summary34%28%38%
11Close with appreciation69%20%11%

Discovery, engaging the objection, and re-closing are the three weakest steps, and they run back to back. Sort by Did not demonstrate and they surface together at the top. A call that never discovered anything has nothing to re-frame against when the objection lands.

Objectionsclick any count to open those calls
TypeCountHandled effectively
2,24021%
1,38871%
98143%
79437%
75259%
35747%
29833%
23155%
17466%
633%

"Not interested" comes up more than any other objection and gets recovered 21% of the time. Price, which every training deck is built around, is already handled on 71% of calls. The coaching budget is pointed at the objection reps solved years ago.

Productsclick any count to open those calls
ProductPitchedSold
92636
31722
3123
30115
25624
2236
21527
15627
1423
1417

The scorer pulls the offer and the price the agent quoted, so one product shows up at several price points, and as unspecified when nobody named a figure. Those rows convert worst. 312 pitches of an unnamed internet tier returned three sales.

Competitive landscapeclick any count to open those calls
CarrierMentionsAs current providerAs objection basis
92278721
61244025
54745216
19415412
1599623
72642
63532
34233
2825
853

Roughly 85% of competitor mentions are the customer telling you who they are with. About 2% become the basis of an actual objection. Train every rep to "handle the competitor" on every mention and you have spent the coaching hour on the wrong 85%. Carrier labels genericized.

Reconstruction · synthetic data

Four surfaces in the dark bar above: the population dashboard, one scored call, a supervisor's coaching digest, and the rubric definition itself. No real agent, customer, client, or caller information appears anywhere on this site.


The interesting parts are the ones I left out.

Happy to walk through any of it in detail. Architecture, trade-offs, and the bugs I'm least proud of.