QC 4.1 Light

OPEN SOURCE · MEASURED · OPEN RUBRIC

Evidence over vibes.

Same call. Typical AI coaching: 32. QC 4.1 Light: 100. Every finding tied to a quote — or it does not ship.

Gong analyzes after. Balto scripts contact centers. Light tells a closer the exact line — open, free, bilingual.

The industry duel

Same anonymized fixture call. Left = unstructured “be more confident” coaching. Right = QC 4.1 Light. Scored by an open rubric.

Typical AI
32

Outcome
Breakpoint
Evidence
QC 4.1 Light
100

Outcome
Breakpoint
Evidence

Before / After

No API key. Anonymized fixture fragment — then the forensic read.

The call that felt fine

        
Forensic read

Open challenge

Beat 100 on evidence grounding. 50 synthetic calls. Public rubric. PRs welcome.

  1. Clone the repo. Run the eval corpus.
  2. Score your agent with eval/score_report.py
  3. If you beat the reference average — open a PR.
python3 eval/generate_corpus.py
python3 eval/run_baseline.py
# reference: 100 · typical AI: 32

Measure quotes against the transcript — not vibes.

Leaderboard

Eval baselines when present.

Run python3 eval/run_baseline.py to generate playground/leaderboard-data.js

Try the playground

Replay the fixture demo. No API keys — ever. For your own calls, use the CLI or install the skill locally.

This public playground is spectacle + proof. Paste-your-key was removed on purpose — use CLI/MCP for real transcripts.