Meeting Assistant Quality

MimicScribe’s meeting assistant is tested against 128 scenarios covering sales discovery calls, customer success reviews, executive meetings, engineering standups, interviews, and more.

How does the assistant score?

CategoryScoreScenarios
Hallucination resistance97%6
Discovery quality90%26
Commitment & requirements tracking91%28
Interpersonal awareness94%12
Action items extraction90%35
Timezone resolution100%7
Question detection89%7
Long-meeting quality (90+ min)96%7

Average briefing latency is about 0.9 seconds — roughly half what it was on the previous model. The full methodology and raw results are published on GitHub.

Does it help me ask better questions?

The assistant also helps you understand the other party’s situation, not just track what’s been said. The examples below are outputs from the test suite.

Spotting workarounds

When someone mentions a manual process or homegrown tool, the assistant recognizes it as a signal that something is missing and suggests exploring it.

Customer mentions exporting dashboard data to spreadsheets every Monday:

  • Ask what an ideal cross-product dashboard looks like.
  • Quantify the impact of Marcus’s manual spreadsheet work.

Prospect reveals their engineer spends 15 hours/week maintaining an internal sync tool that breaks every release:

  • Ask what happens if replacement isn’t live in two months.
  • Explore ML pipeline impact if Yusuf’s time is freed.

Quantifying impact

Instead of accepting “it’s broken” at face value, the assistant suggests questions that surface the business cost.

Prospect says “our reports are always wrong”:

  • Ask what’s driving the two-day reconciliation delay.
  • Ask how much time Finance and Ops waste monthly.

Ideal-state questions

When someone has described their problem in detail, the assistant suggests asking what the perfect solution would look like.

Customer describes 45-minute daily data reconciliation between two systems:

  • Ask what an ideal automated solution would look like.

Beta tester says “the export is clunky”:

  • Ask what an ideal sharing workflow looks like.

Surfacing what hasn’t been discussed

When key requirements like budget, decision-makers, or timeline haven’t come up, the assistant raises them naturally.

20 minutes of technical discussion, budget and authority never mentioned:

  • Ask who else needs to approve the Q4 budget.

Interview depth

In interview settings, the assistant catches when candidates give surface-level answers and suggests specific follow-ups.

Candidate proposes Redis with consistent hashing but doesn’t address failure modes:

  • Ask how they would handle hot keys specifically.
  • Ask about cache stampede prevention for expired keys.

Does it track the objective I set?

Yes. If you set an objective for a meeting, the assistant tracks whether it has actually been discussed, not just mentioned. An uncovered topic moves to the top of the list as time runs short.

It distinguishes a topic that was named in passing from one explored with specifics, numbers or commitments.

Do I need to write detailed prep notes?

No. In our tests a single line like “Help me understand their situation before suggesting solutions” was enough to produce discovery-oriented suggestions, with no names, companies or background given.

Adding a name and company improves polish. Adding deal context such as budget status or prior conversations produces more targeted suggestions. It works with minimal input too.

How we test

Each scenario includes a realistic multi-speaker transcript (15-30 exchanges) with domain-specific language, and is run 3 times to measure consistency. We use both deterministic checks (keyword presence, bullet count, output structure) and an independent LLM judge that evaluates semantic quality — whether suggestions are grounded in the transcript, forward-looking, and free of fabricated details.

The scores above are a few points lower than earlier snapshots, for two honest reasons. First, the whole assistant pipeline moved from Gemini 3 Flash — a preview model that was retired — to Gemini 3.1 Flash Lite, a smaller and faster model; briefing latency roughly halved as a result. Second, the test suite grew and the rubric tightened: talking points now fail if they restate something the meeting already covered instead of advancing the conversation. We publish the lower numbers rather than leave the retired model’s numbers up.

Scenarios cover:

  • First discovery calls where the prospect’s problem isn’t yet understood
  • Late-stage negotiations where commitments need tracking
  • Customer success reviews with unexpressed needs
  • User research sessions with vague feedback
  • Engineering standups with unowned blockers
  • Executive meetings approaching a decision deadline
  • Long meetings (90+ minutes) with a running summary