Scores

What scores mean and how to read them.

Every measurement contains lists of scores, each pairing a name with a probability. The audio upload and realtime endpoints return expressions and voice_attributes for each window of speech. The video endpoints return expressions and descriptions for each face. Audio measurements and Video measurements list the names each array can contain.

Score objects

Score
{
"name": "amusement",
"probability": 0.7124
}
name

Lowercase, and may contain spaces and punctuation, as in aesthetic appreciation and surprise (positive). New names may be added over time, so treat one you do not recognize as a new prediction.

probability

From 0 to 1, rounded to four decimal places.

Reading probabilities

The probability in each array denotes a different quantity, and each array includes only the scores that pass its own cutoff.

Arrayprobability isIncluded when
Voice expressionsThe measured chance that the expression is truly present, according to human ratersJudged present, against a threshold set for each expression
Voice voice_attributesCalibrated to the expected intensity a human rater would give the quality0.725 or above, except nasal and breathy, which are never included
Face expressionsThe chance a human rater would name the expression for this faceAbove 0.1
Face descriptionsThe chance that the facial action is presentAbove 0.1
  1. Scores are independent. Each probability is its own estimate for one name. Scores do not sum to 1, several names can score highly at once, and they are not meant to be compared with each other.
  2. An empty array is a valid result. Arrays vary in length, and an empty array means that no score passed the cutoff.
  3. Arrays are sorted by descending probability. In voice arrays, scores with equal probability are ordered by the model’s underlying score, then by name. In face arrays, they are ordered by name.
  4. The audio and video endpoints use different expression names. A name in one may not exist in the other.
  5. Face confidence measures detection. It is the detector’s confidence that a bounding box contains a face, separate from the expression scores.

What expression scores denote

An expression score denotes what a voice or face conveys to the people who hear or see it. Expression scores are grounded in judgments from people in many countries, and Hume’s publications describe the studies that collected them. A score is not a direct readout of what the person feels, because emotional experience is subjective and its expression varies from person to person and from one setting to another.

  1. The studies follow semantic space theory, which maps the kinds of emotion and how they relate from judgments of thousands of naturalistic expressions. Three of its findings explain how the scores work:
    1. Expressions convey many distinct emotions. Across studies of experience and expression, 20 to 25 distinct kinds of emotion emerge, far more than the six that earlier research focused on, so the models score many expression names.
    2. Specific emotion concepts describe expressions best. Concepts such as amusement and doubt capture what people perceive more precisely than broad dimensions such as valence and arousal, so the names are emotion concepts.
    3. Emotions blend. An expression often conveys several emotions at once, and emotions shade into one another along gradients, so a window of speech or a face can have several expressions, each with its own score. Closely related names, such as laughter and hysterical laughter, often appear together.
  2. Voice expressions and voice_attributes are separate dimensions. expressions describe what the speech conveys, and voice_attributes describe how the voice and the audio sound. The two often go together, as when boredom comes with a monotone voice, and they can also diverge: a person can be furious and still whisper.
  3. Face descriptions name visible facial actions, and face expressions name what the face conveys. The same facial action can convey different emotions to different people, and different actions can convey the same emotion.
  4. The voice and the face are separate channels. A person’s voice and face can express different things at the same moment, so voice and face results for the same moment can differ.
  5. Setting and culture shape expression. How people express an emotion depends on the setting. Across cultures, expressions largely share their meanings, with some differences in meaning and in how intensely people display them. Test an application in each population it serves.
  6. Conclusions about feelings are strongest across many measurements. In Hume’s research on facial expression, expressions averaged across many people predicted what those people reported feeling far more accurately than one person’s expression did.
  7. Commercial applications must follow the ethical guidelines of The Hume Initiative.

Using scores in an application

Research and applications built on expression scores generally move through four stages.

  1. Exploration. Look for patterns in your data: differences between users or study participants, changes over time, and differences between stages of a study or product experience.
  2. Prediction. Use scores to predict outcomes you already know matter, such as customer satisfaction or mental health. Check whether expression and language together predict an outcome better than language alone. If expression predicts an outcome, track how it changes over time to find the critical moments for a user.
  3. Improvement. Use what predicts well to change how the application works:
    1. Act on a prediction directly. If expression and language predict whether two people will get along, the application can pair them up.
    2. Apply statistics or machine learning to the data you gather.
    3. Describe the scores in words and add them to a language model prompt, such as “The user sounds calm but a little frustrated.”
    4. Fine-tune a model, such as an AI tutor, using the expressions that predict student performance and well-being.
  4. Testing. Make expression part of every A/B test, so each change is measured by how often users laugh or express frustration, interest, or boredom, alongside engagement and retention.

Next steps