Business

Reading emotions at scale

—

Amelia Zhao

Outset's emotion analysis performs in line with published results on a widely used emotion-recognition benchmark.

What the system detects

Outset labels what a participant says, moment by moment, with six basic emotions (anger, disgust, fear, happiness, sadness, surprise) plus neutral. Popularized by the psychologist Paul Ekman [1], this framework is widely used in emotion-recognition research. We use a named scale rather than a proprietary sentiment score because it is auditable: any label we report can be checked against the interview it came from.

What we tested

We tested our own approach, not a model on its own: the six emotions plus neutral as we define them, and the instructions we give the model for recognizing them. We evaluated them together on MELD [2], a widely used dataset for emotion recognition from video, across all 2,610 clips in its test set.

Nothing in our setup was tuned for this benchmark. The system read each clip the way it reads a live interview, working from expression and what is said.

How the result compares

Outset scores 0.53 on MELD, inside the range published for the strongest systems tested on this benchmark: leading commercial models and the largest open ones, none of them trained on the training set of MELD.

Figure 1. Outset's weighted F1 on MELD against that published range.

What an F1 score measures

An F1 score combines two things: how often the system notices an emotion when it is really there, and how often it is right when it says so. It runs from 0 to 1, where 1 would mean catching every instance and never mislabelling one.

A higher score means the system does a better job recognizing emotions across the test clips. Any individual label can still be wrong.

Figure 2. Using anger as the example: the system is judged both on how much of the real anger it finds and on how often it is right when it calls something anger. F1 rewards doing both, so a system cannot score well by guessing freely or by staying quiet.

Why we report F1, not accuracy

Accuracy can hide poor detection of rare emotions. Fear appears in only 1.9% of clips, so always answering "no fear" would be 98.1% accurate on that question while detecting none of it. Its F1 for fear would be zero. We report F1 to account for both missed emotions and incorrect labels. The full results, including the complete confusion matrix, are available for technical review.

Which emotions it reads best

Table 1. F1 score by emotion.

The system performs best on the emotions that appear most often in the data, and weakest on the two that are rarest: fearand disgust, where a label is most worth checking against the clip.

How to use it

Use emotion analysis to decide which moments to watch, then watch them before drawing conclusions. That might be a pattern across a study, such as a checkout flow that consistently draws anger, or a single outlier worth a closer look.

Where we're extending this work

The benchmark above uses public clips, not real interviews. We're applying the same discipline to how the system performs on actual study interviews, the setting research happens in. We plan to publish those results when the analysis is complete.

Detailed methodology

The analysis ran zero-shot on a frontier multimodal model. Each clip was scored once, in a single pass over the complete MELD test set: 2,610 clips, with no failed calls. The system received the video with its original audio, with no labels and no prior conversational context.

Both headline figures are computed from the per-emotion scores in Table 1. The weighted F1 of 0.5262, shown rounded in Figure 1, averages them in proportion to how often each emotion occurs, and is the standard metric for this benchmark. Macro F1 averages them equally however rare an emotion is, and comes to 0.4232 (95% confidence interval 0.395 to 0.450). We track the macro figure internally because it is the number that would expose a gain on common emotions masking a loss on rare ones.

The comparison range covers zero-shot results reported for leading commercial and open models on the same benchmark [3, 4], taking 72B-class models as the open comparators: a weighted F1 of 0.37 to 0.63. The studies used different combinations of audio, video, transcripts, and conversation history, so these comparisons serve as a point of reference rather than an identical test setup. Our predictions and scoring scripts are retained and available for technical review.

References

  1. Ekman, P. (1992). An argument for basic emotions. Cognition and Emotion, 6(3–4), 169–200.

  2. Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., & Mihalcea, R. (2019). MELD: A multimodal multi-party dataset for emotion recognition in conversations. Proceedings of ACL 2019. arXiv:1810.02508

  3. Murzaku, J., & Rambow, O. (2025). OmniVox: Zero-shot emotion recognition with Omni-LLMs. arXiv:2503.21480

  4. Zhang, H., Li, Z., Xu, H., Zhu, Y., Wang, P., Zhu, H., Zhou, J., & Zhang, J. (2025). Can large language models help multimodal language analysis? MMLA: A comprehensive benchmark. NeurIPS 2025, Datasets and Benchmarks Track.

Interested in learning more? Book a personalized demo today!

Book Demo

About the author
Amelia Zhao

Engineering Manager

Amelia is an Engineering Manager at Outset

Subscribe to our newsletter

Enter your contact details to get the latest tips and stories to help boost your business. 

Subscribe to our newsletter

Enter your contact details to get the latest tips and stories to help boost your business. 

Subscribe to our newsletter

Enter your contact details to get the latest tips and stories to help boost your business.