- Outcome
- Breakpoint
- Evidence
OPEN SOURCE · MEASURED · OPEN RUBRIC
Evidence over vibes.
Same call. Typical AI coaching: 32. QC 4.1 Light: 100. Every finding tied to a quote — or it does not ship.
Gong analyzes after. Balto scripts contact centers. Light tells a closer the exact line — open, free, bilingual.
The industry duel
Same anonymized fixture call. Left = unstructured “be more confident” coaching. Right = QC 4.1 Light. Scored by an open rubric.
- Outcome
- Breakpoint
- Evidence
Before / After
No API key. Anonymized fixture fragment — then the forensic read.
Open challenge
Beat 100 on evidence grounding. 50 synthetic calls. Public rubric. PRs welcome.
- Clone the repo. Run the eval corpus.
- Score your agent with eval/score_report.py
- If you beat the reference average — open a PR.
python3 eval/generate_corpus.py python3 eval/run_baseline.py # reference: 100 · typical AI: 32
Measure quotes against the transcript — not vibes.
Leaderboard
Eval baselines when present.
Run python3 eval/run_baseline.py to generate playground/leaderboard-data.js
Try the playground
Replay the fixture demo. No API keys — ever. For your own calls, use the CLI or install the skill locally.
This public playground is spectacle + proof. Paste-your-key was removed on purpose — use CLI/MCP for real transcripts.