Scores
Every measurement contains lists of scores, each pairing a name with a probability. The audio upload and realtime endpoints return expressions and voice_attributes for each window of speech. The video endpoints return expressions and descriptions for each face. Audio measurements and Video measurements list the names each array can contain.
Score objects
Lowercase, and may contain spaces and punctuation, as in aesthetic appreciation and surprise (positive). New names may be added over time, so treat one you do not recognize as a new prediction.
From 0 to 1, rounded to four decimal places.
Reading probabilities
The probability in each array denotes a different quantity, and each array includes only the scores that pass its own cutoff.
- Scores are independent. Each
probabilityis its own estimate for one name. Scores do not sum to 1, several names can score highly at once, and they are not meant to be compared with each other. - An empty array is a valid result. Arrays vary in length, and an empty array means that no score passed the cutoff.
- Arrays are sorted by descending probability. In voice arrays, scores with equal probability are ordered by the model’s underlying score, then by name. In face arrays, they are ordered by name.
- The audio and video endpoints use different expression names. A name in one may not exist in the other.
- Face
confidencemeasures detection. It is the detector’s confidence that a bounding box contains a face, separate from the expression scores.
What expression scores denote
An expression score denotes what a voice or face conveys to the people who hear or see it. Expression scores are grounded in judgments from people in many countries, and Hume’s publications describe the studies that collected them. A score is not a direct readout of what the person feels, because emotional experience is subjective and its expression varies from person to person and from one setting to another.
- The studies follow semantic space theory, which maps the kinds of emotion and how they relate from judgments of thousands of naturalistic expressions. Three of its findings explain how the scores work:
- Expressions convey many distinct emotions. Across studies of experience and expression, 20 to 25 distinct kinds of emotion emerge, far more than the six that earlier research focused on, so the models score many expression names.
- Specific emotion concepts describe expressions best. Concepts such as
amusementanddoubtcapture what people perceive more precisely than broad dimensions such as valence and arousal, so the names are emotion concepts. - Emotions blend. An expression often conveys several emotions at once, and emotions shade into one another along gradients, so a window of speech or a face can have several expressions, each with its own score. Closely related names, such as
laughterandhysterical laughter, often appear together.
- Voice
expressionsandvoice_attributesare separate dimensions.expressionsdescribe what the speech conveys, andvoice_attributesdescribe how the voice and the audio sound. The two often go together, as whenboredomcomes with amonotonevoice, and they can also diverge: a person can be furious and still whisper. - Face
descriptionsname visible facial actions, and faceexpressionsname what the face conveys. The same facial action can convey different emotions to different people, and different actions can convey the same emotion. - The voice and the face are separate channels. A person’s voice and face can express different things at the same moment, so voice and face results for the same moment can differ.
- Setting and culture shape expression. How people express an emotion depends on the setting. Across cultures, expressions largely share their meanings, with some differences in meaning and in how intensely people display them. Test an application in each population it serves.
- Conclusions about feelings are strongest across many measurements. In Hume’s research on facial expression, expressions averaged across many people predicted what those people reported feeling far more accurately than one person’s expression did.
- Commercial applications must follow the ethical guidelines of The Hume Initiative.
Using scores in an application
Research and applications built on expression scores generally move through four stages.
- Exploration. Look for patterns in your data: differences between users or study participants, changes over time, and differences between stages of a study or product experience.
- Prediction. Use scores to predict outcomes you already know matter, such as customer satisfaction or mental health. Check whether expression and language together predict an outcome better than language alone. If expression predicts an outcome, track how it changes over time to find the critical moments for a user.
- Improvement. Use what predicts well to change how the application works:
- Act on a prediction directly. If expression and language predict whether two people will get along, the application can pair them up.
- Apply statistics or machine learning to the data you gather.
- Describe the scores in words and add them to a language model prompt, such as “The user sounds calm but a little frustrated.”
- Fine-tune a model, such as an AI tutor, using the expressions that predict student performance and well-being.
- Testing. Make expression part of every A/B test, so each change is measured by how often users laugh or express frustration, interest, or boredom, alongside engagement and retention.

