If what you’re after is how it compares with a chat, a workshop or a video LMS, that’s in compare. Here are the underlying reasons, in the order they matter.
What we decided, and why
01The comparison is against your own baseline.
There is no correct voice in Frank’s head. In the first session, how each person speaks gets recorded: their rhythm, their pauses, their tone at rest. Everything that comes after is measured as change against that, not against an average or an external ideal.
It’s a product decision and a matter of principle. A strong accent, a disfluency or an unusual timbre aren’t noise: they’re the starting point. And comparing each person to themselves is what lets the report be private and useful, instead of a ranking.
The evidence
For a comparison against the baseline to make sense, the engine has to measure the same thing the reference tools measure. Acoustic engine fidelity: .99 on the pitch family, .97 on loudness, .94 on harmonic quality.
Internal measurement. Pearson correlation against openSMILE eGeMAPS v02 and Praat, RAVDESS corpus, 24 speakers.
This is engine fidelity, not outcome validation. It does not imply these signals predict performance.
02The pressure is by design.
A counterpart that waits, validates and gives in trains nothing. Frank interrupts, holds the objection even when you give it a good argument, leaves the silence open, and can end the meeting. The friction is the same one that shows up in the room afterwards.
It isn’t pressure for its own sake. Frank deliberately recreates a middle zone: safe enough to fail, real enough to matter. That’s why there’s no coach during the session and no real-time advice: the analysis arrives at the end, once it’s over.
The evidence
Performance rises with arousal and then collapses: the curve is an inverted U, and learning happens in the middle zone.
Yerkes and Dodson, 1908.
30% of a negotiation’s outcome is decided in the first five minutes of speech. That study didn’t measure the arguments: it measured the dynamics of the conversation.
Curhan and Pentland, 2007. J. Applied Psychology.
03The conversations a company can’t record.
A sales call can be recorded, replayed and corrected. Hard feedback, a conflict between two people on the team, letting someone go: those can’t. They’re private by nature, and they’re exactly the ones that weigh most on a career.
That’s where Frank lives. A place where those conversations can be practiced out loud, with a counterpart that answers back, with no real person on the other side and no mistake costing anyone anything. The report belongs to the person who practiced. The organization sees aggregated data.
The evidence
Structured roleplay predicts job criteria with r ≈ .37, in a meta-analysis of more than 12,000 cases. It’s the evidence that simulating the situation, not just studying it, is a valid method.
Gaugler et al., 1987. J. Applied Psychology.
In practice
That meta-analysis backs the method. What Frank adds is the measure: each person’s delta against their own baseline, session after session.
04Native Latin American Spanish, not translated.
Frank is written in Latin American Spanish, not translated from English. The objections, the silences and the way a meeting gets cut short are the ones from here. English runs in the same program.
A single program can hold stories in both languages, and the same person can practice in either. Because the comparison is against their own baseline, switching languages doesn’t move the bar: it moves the reference.
In practice
It shows in the script: the objections and the silences are written in the Spanish spoken in the room, not translated. The counterpart’s voice answers in under 1.5 seconds and the session runs in the browser, in whichever language each person picks.