Skip to content

Four design decisions, and what backs each one.

Frank explains itself through how it’s built. Here are the four decisions that define it, what backs each one, and how they show up in a real session.

If what you’re after is how it compares with a chat, a workshop or a video LMS, that’s in compare. Here are the underlying reasons, in the order they matter.

What we decided, and why

01

The comparison is against your own baseline.

There is no correct voice in Frank’s head. In the first session, how each person speaks gets recorded: their rhythm, their pauses, their tone at rest. Everything that comes after is measured as change against that, not against an average or an external ideal.

It’s a product decision and a matter of principle. A strong accent, a disfluency or an unusual timbre aren’t noise: they’re the starting point. And comparing each person to themselves is what lets the report be private and useful, instead of a ranking.

The evidence

For a comparison against the baseline to make sense, the engine has to measure the same thing the reference tools measure. Acoustic engine fidelity: .99 on the pitch family, .97 on loudness, .94 on harmonic quality.

Internal measurement. Pearson correlation against openSMILE eGeMAPS v02 and Praat, RAVDESS corpus, 24 speakers.

This is engine fidelity, not outcome validation. It does not imply these signals predict performance.

02

The pressure is by design.

A counterpart that waits, validates and gives in trains nothing. Frank interrupts, holds the objection even when you give it a good argument, leaves the silence open, and can end the meeting. The friction is the same one that shows up in the room afterwards.

It isn’t pressure for its own sake. Frank deliberately recreates a middle zone: safe enough to fail, real enough to matter. That’s why there’s no coach during the session and no real-time advice: the analysis arrives at the end, once it’s over.

The evidence

Performance rises with arousal and then collapses: the curve is an inverted U, and learning happens in the middle zone.

Yerkes and Dodson, 1908.

30% of a negotiation’s outcome is decided in the first five minutes of speech. That study didn’t measure the arguments: it measured the dynamics of the conversation.

Curhan and Pentland, 2007. J. Applied Psychology.

03

The conversations a company can’t record.

A sales call can be recorded, replayed and corrected. Hard feedback, a conflict between two people on the team, letting someone go: those can’t. They’re private by nature, and they’re exactly the ones that weigh most on a career.

That’s where Frank lives. A place where those conversations can be practiced out loud, with a counterpart that answers back, with no real person on the other side and no mistake costing anyone anything. The report belongs to the person who practiced. The organization sees aggregated data.

The evidence

Structured roleplay predicts job criteria with r ≈ .37, in a meta-analysis of more than 12,000 cases. It’s the evidence that simulating the situation, not just studying it, is a valid method.

Gaugler et al., 1987. J. Applied Psychology.

In practice

That meta-analysis backs the method. What Frank adds is the measure: each person’s delta against their own baseline, session after session.

04

Native Latin American Spanish, not translated.

Frank is written in Latin American Spanish, not translated from English. The objections, the silences and the way a meeting gets cut short are the ones from here. English runs in the same program.

A single program can hold stories in both languages, and the same person can practice in either. Because the comparison is against their own baseline, switching languages doesn’t move the bar: it moves the reference.

In practice

It shows in the script: the objections and the silences are written in the Spanish spoken in the room, not translated. The counterpart’s voice answers in under 1.5 seconds and the session runs in the browser, in whichever language each person picks.

What a buyer can verify.

Every report shows the delta against that person’s baseline, with signal quality and the margin of uncertainty stated on the same screen. And the method is published in full: what a pilot looks like from the inside, week by week, so you can review it before the demo.

Three terms we use, and what they mean.

The rest is in the glossary.

Individual baseline
How a person speaks at rest, recorded in their first session. Everything else is measured as change against that.
Paralinguistic analysis
How something was said, not what was said: rhythm, pauses, energy, variation in pitch. Frank describes it; it doesn’t diagnose emotions or predict performance.
Moment
A point in the session that gets marked on the timeline: an interruption, a long silence, the close. The report is read moment by moment.

The reasons are here. The proof is in one session.

30 minutes, nothing to install, with a scenario already written for your company.

Check it with your own microphone.

Listen to Frank

Thirty minutes, nothing to install. The demo is run by one of the four people who build Frank.