Speaker Hub uses necessary cookies for secure accounts. With your permission, first party analytics help us improve signup and onboarding. We do not use advertising or cross site tracking cookies. Read the Cookie Policy.
The Speech Analyzer score does not come from a chatbot. It comes from a brain-scan model with a name. · Speaker Hub
The Speech Analyzer score does not come from a chatbot. It comes from a brain-scan model with a name.
Every Speech Analyzer report is built on a named research model trained on real fMRI data, plus two separate transcription engines. Here is what TRIBE v2, Parakeet, and Whisper actually do, and why knowing that changes how much weight to put on the number.
•6 min read
Upload a talk to Speech Analyzer and a number comes back. Most speakers stop there, treat it like a verdict from an anonymous AI, and either trust it too much or dismiss it as a black box. Neither reaction fits what is actually running under the report, because the model is not anonymous. It has a name, a paper, and a license, and all three are printed at the bottom of every report you get.
The score comes from a model trained on real brain scans, not on talk transcripts
The model doing the prediction is called TRIBE v2, built by researchers at Meta AI. It is not a general chatbot repurposed for a new job. It was trained on more than 1,000 hours of fMRI data collected from 720 people while they watched and listened to video, and its job is to predict how a person's brain would respond to new video, audio, and text it has never seen. An earlier version of the same architecture won the Algonauts 2025 brain encoding competition, the field's benchmark for this exact task.
TRIBE v2 gets there by combining three separate models built for three separate senses: Llama 3.2 for the words, V-JEPA 2 for the video, and Wav2Vec-BERT for the audio track, fused into one system that reads your talk the way a brain takes in all three at once. That is a meaningfully different claim than "an AI watched your talk and scored it." It is closer to "a model trained to match real neural recordings estimated what a similar recording would look like for your talk." Still an estimate. Not a guess pulled from nowhere. You can read the model card yourself on Hugging Face.
The license is why the tool is free, not a marketing choice
TRIBE v2 ships under a Creative Commons BY-NC 4.0 license. The NC stands for non-commercial, and it is not a suggestion. A tool that charges for output generated by an NC-licensed model is using that model outside its terms. That is the actual reason Speaker Hub's Speech Analyzer has no paywall, no credit limit, and no upsell sitting between you and a report: keeping it free is how the analyzer stays inside the license it was built on, not a growth tactic that might change later.
It also means the "BY" half of the license applies. Meta requires attribution with a link back to the model and its license wherever it is used, so the credit lives on the analyzer page and on every report itself, not buried in a terms-of-service page nobody reads.
Two transcription engines, not one, and both can get your words wrong
Before TRIBE v2 sees your talk, something has to turn the audio into text. That job runs on NVIDIA's Parakeet TDT 0.6B v3, a speech-to-text model released under Creative Commons BY 4.0. If Parakeet cannot process a given file, the analyzer falls back to OpenAI's Whisper, released under the MIT license.
Neither transcript is guaranteed to be right. A mumbled word, an unusual name, or background noise from the venue can all turn into a wrong word on the page, and that wrong word can end up quoted in a coaching note that follows from it. This is the actual mechanical reason the report ties every finding to a timestamp in your original recording instead of just printing a sentence and asking you to take it on faith. Two transcription engines lower the odds of a bad transcript. They do not remove the odds.
What actually changes once you know this
Read the score as a prediction from a model trained on real brains, not a survey of hypothetical listeners. That is a stronger foundation than a language model guessing at engagement from a transcript alone, and it is also still a prediction. Both things are true at once, and holding both is the point.
Check the clip before you act on a coaching note. Two transcription models is redundancy, not a guarantee. If a note sounds off, it may be describing a word the audio engine misheard rather than a word you said.
Do not expect the analyzer to ever charge you for a report. The free-forever framing is not a promise Speaker Hub is making on faith. It is a condition of the license the underlying model runs on.
None of this changes what to do with a single report: find the one moment worth rehearsing, fix it, and run the next version through Compare to see if it actually moved. What it changes is how much weight to put on the number while you do that. A prediction trained on 720 real brains responding to real video is worth taking seriously. It is still not a measurement of the room you are about to walk into.
What model actually produces the Speech Analyzer's response score?
TRIBE v2, a model built by Meta AI researchers and trained on more than 1,000 hours of real fMRI brain scans from 720 people. It combines a text model (Llama 3.2), a video model (V-JEPA 2), and an audio model (Wav2Vec-BERT) into one system that predicts how a viewer's brain is likely to respond to your talk, moment by moment.
Why is Speaker Hub's Speech Analyzer free if it runs on a research model?
TRIBE v2 is released under a Creative Commons BY-NC 4.0 license, and the NC means non-commercial. Speaker Hub keeps the analyzer free specifically to stay inside that license. Charging for reports built on it would step outside the terms Meta released the model under.
Which model writes the transcript under my report?
Usually NVIDIA's Parakeet TDT 0.6B v3. If it cannot process a file, OpenAI's Whisper runs as a fallback. Both can make errors, which is why every finding in a report links back to a timestamp in your own recording instead of asking you to trust the text on its own.
Analyze produces one number for one talk. Compare and Talk Library are the other two-thirds of the tool, and they are what turn that number into evidence of whether you actually got better. Here is what each part does and how to use them.
The Speech Analyzer pairs your recording with a transcript and an estimated response timeline, and none of that is a measurement of how the room actually felt. Here is what the number is built from, what to check before you act on it, and why Compare matters more than any single score.
Two psychologists spent years studying um and uh and concluded they are words, not mistakes. Here is what to cut instead, and the one filler that actually costs you.