We tested 12 vision-capable large language models on 70 labeled face photographs, five times each, to assess how accurately they identify the emotion on a human face. The highest score is 67%, no two models in the test are statistically distinguishable, and the cohort reads anger at 22%.
In addition, we explore ten leading emotion AI tools and share our hands-on insights.
Benchmark on emotion recognition
Each model saw the same 70 images five times, one image per request, for 4,200 calls and 4,195 valid classifications. See the methodology for the label set and scoring rules.
Emotion recognition benchmark results
The scores span 51% to 67%, and no two models separate. Comparing every pair on the answer each model gave most often across its five passes, four of the 66 exact McNemar tests have uncorrected p-values below 0.05, and none survives Holm correction. With 70 images, the 16-point spread between the first and last is one band, and the order within it carries no meaning.
The best score has not improved measurably in seven months.
GPT o4 Mini High scored 69% on these same 70 images in January 2026 in a single pass, compared to 67% from Gemini 3.7 Flash over five passes. The two runs used different versions of the test code, so the 1.5-point gap is not a decline; it is a result that has not moved.
Accuracy by emotion
The seven emotions do not fail together:
- Happiness: read correctly 89% of the time across the cohort
- Fear: 38%
- Anger: 22%, and the lowest-scoring emotion for 10 of the 12 models
Anger sits eight points above guessing. Balanced classes put a random guess at 14.3%, so:
- Anger clears the chance by eight points; happiness clears it by 75 %
- The best model on anger reaches 34%
Nothing forces a model to reach the baseline: two score zero when the image is withheld.
Models also disagree with themselves. Asked the same image five times:
- MiniMax M3 repeats its own most common answer 81% of the time
- Claude Fable 5 repeats it 97% of the time
A single-pass number on this task therefore carries the model and the run-to-run variation together.
That instability disappears in the images that the cohort gets wrong. Three images drew no correct answers from any model across any of the five passes (60 answers), zero matches.
The models did not scatter across the remaining six labels. They converged on one:
- Anger-labeled image: sadness in 59 of 60
- Fear-labeled image: sadness in 47 of 60 answers
Convergence this tight across 12 independent models points at the label as readily as at the models.
Cost, latency and accuracy
Cost spans 62x across the cohort while accuracy spans 16 points.
- GPT-5.6 Luna classifies 1,000 images for $0.06 and scores 53%
- Claude Fable 5 charges $3.83 for the same 1,000 and scores 63%.
Those ten points do not survive the Holm correction.
Slower answers score no better.
- Qwen3.8 Max averages 14.5 seconds per image with a 95th percentile of 46 seconds and scores 63%
- Claude Sonnet 5 averages 3.1 seconds and scores 64%
The figure is round-trip time through OpenRouter, so it carries routing as well as the model’s own deliberation, but a 4x spread on a one-word answer is too large to come from routing alone.
Methodology of emotion recognition benchmark
Each model receives one image and the same instruction: pick exactly one word from a list of seven emotions. The labels are “happy”, “sad”, “angry”, “fear”, “disgust”, “surprise,” and “neutral”, shuffled per request so no label holds a fixed position.
An answer counts as correct only if it matches the dataset label exactly. Anything else, including a sentence, a refusal, or a word outside the list, is recorded as an invalid response and is not scored as correct. Five of the 4,200 calls were invalid: MiniMax M3 returned a truncated word four times, and Kimi K3 answered with a paragraph once.
Every model is called through OpenRouter, one image per request, with no conversation history and no per-vendor prompt. This matters more than it sounds:
In an earlier run on this dataset, some models were called through their vendor APIs and others via OpenRouter, and every OpenRouter model scored between 7% and 21% while every direct-API model scored between 51% and 63%.
To confirm that the images reach the models, we run the same prompt with the image removed:
- Every model gains between 37 and 64 points when the image is present.
- Without it, eight of the twelve land within a point of the 14.3% baseline, and seven of those answer with a single label for all 70 items.
- Claude Fable 5 refuses to guess on all 70, and Claude Sonnet 5 on 68 of them, so both score zero.
Dataset
We use a part of the Facial Emotion Detection dataset, which includes a set of labeled images showing different human emotions.1 Each image contained facial expressions representing common emotional states such as happiness, sadness, anger, fear, and surprise.
Cost
The completion budget is 16,000 tokens for every model, and reasoning tokens are drawn from it.
Qwen3.8 Max is the only model that comes near it: on one image, its reasoning spent ranged from 297 to 2,115 tokens across repeated calls, and two of its answers were cut off entirely at an earlier 4,000-token budget.
The two truncated answers were rerun at the wider budget because an answer that hits a budget wall reflects the budget rather than the model. No other model in the cohort exceeded 300 reasoning tokens.
Known limits
Five of the 70 images carry a label that all 12 models contradict, and on three of them no model produced the label in any pass. Model agreement does not prove a label is wrong, since the models are correlated rather than independent annotators, but it does mean no model should be expected to approach 100% on this set.
We ran the whole benchmark twice. Two independent five-pass runs of the same 12 models on the same 70 images, under the same settings:
- Each model moves by 1.2 points on average, and by 4.9 points at the extreme
- Nine of the twelve models change rank between the two runs
- Three positions hold: first, second and last
Repeating the experiment reaches the same place the significance tests do, without running a test.
Three things this run does not establish:
- The tests report failure to detect a difference, not equivalence. No two models are shown to be equal; they are shown to be unseparated at this sample size.
- A photograph with one ground-truth label is a laboratory task. Naming the emotion on a still face is not the same problem as reading a customer on a support call.
- The two dedicated emotion AI products previously covered in this benchmark, Hume and Imentiv AI, are not in this run. We are re-testing both and will publish their scores under the same scoring rules as the models.
Affective computing tools comparison
Hume Expression Measurement
Hume Expression Measurement is an emotion AI tool that helps identify and measure human emotions. It works through a single app and uses four types of data: voice, images, video, and facial expressions. Together, these offer a deeper and more detailed look at how people express emotions.
Real-life experience
This emotion recognition software may not always be 100% accurate, but it captures emotional nuances effectively, especially through speech patterns. However, it’s not perfect. Sometimes, it may not detect basic emotion from vocal bursts. Still, the emotional results often feel realistic and nuanced.
Hume is best for users who want a detailed and responsive look at emotional behavior, not simple labels like “happy” or “sad.” The web application for the emotion recognition software is extremely user-friendly.
Key features
- The software provides a real-time analysis for emotions, sentiment, and toxicity for a given text.
Figure 1. Hume Expression Measurement text analysis for emotions
Figure 2. Hume Expression Measurement text analysis for sentiment
For more information on sentiment analysis, check our sentiment analysis articles.
- This emotion recognition software also detects emotions from videos, images, and audio documents. Users might upload documents, or they may prefer to use their own camera and speakers for emotion detection.
Hume analyzes speech, images, and videos using several features:
- Facial expression: Detects facial movements to understand facial emotions like joy, anger, or sadness.
- Vocal burst: Measures how someone sounds, whether calm, excited, stressed, etc.
- Speech prosody: Tracks changes in tone, pitch, and rhythm. This helps identify the emotional tone of what someone is saying.
Figure 3. Hume Expression Measurement video analysis for speech prosody
Mangold Observation Studio
Mangold Observation Studio is a comprehensive platform designed for advanced, sensor-driven research. It brings together many data sources, video, audio, facial expressions, physiological signals, and more, into one synchronized system.
Key features
- Video and screen recording: Captures participants’ behavior and screen activity for full context.
- Sensor integration: Supports EEG, eye tracking, heart rate, skin response, and muscle activity.
- Speech analysis: Converts spoken words into text automatically.
- Surveys and annotations: Add participant feedback or tag key moments during sessions.
- Multimodal design: Unlike tools that focus on one data type (like facial expression), Mangold combines over 120 sensor types in one platform.
- Scalable setup: Supports unlimited participants and devices at once, with time-synced recordings.
- Full network control: All devices can be managed from a central station.
- Modular and customizable: Researchers can build their own setup and integrate with external tools using an API.
Visage SDK
Visage SDK is a facial emotion recognition software that helps businesses track and analyze faces in real-time. It uses advanced computer vision to understand people’s emotions, age, gender, and identity.
Key features
- Online & offline support: Works both online (in the cloud) and offline (on your device), so you’re not always dependent on an internet connection.
- Privacy-first: Ensures that no personal data, like names or photos, is stored or processed without your consent.
- Unity integration: Integrates with Unity for creating face filters or interactive experiences in games.
Applications
- Virtual try-ons: Use face recognition to let customers try on glasses, makeup, or other products virtually.
- Driver monitoring: Detect unsafe driving behavior, such as drowsiness or distraction, to enhance road safety.
- Passenger monitoring: Track passengers’ well-being in cars or public transportation to improve safety and comfort.
- Augmented reality (AR): Create fun, engaging experiences like beautification filters or realistic face masks for social media or apps.
Imentiv AI
Imentiv AI is an emotion detection software that helps users understand how people feel, speak, and behave in video, audio, and text content. It combines artificial intelligence with psychological expertise to analyze human emotion and personality in real time.
Real-life experience:
Imentiv AI helps users analyze emotions from video content. You can upload a full video or focus on a specific frame. The tool looks at facial expressions, voice tone, and the transcript to understand emotional cues.
The analysis seems accurate and covers a wide range of emotional signals. In addition to basic insights, the platform also offers psychological evaluations. These can be scheduled through an appointment system.
Figure 4. Imentiv AI personality trait analysis
Key features
- Multimodal analysis: Analyzes video, audio, and text together. This gives a fuller picture of emotional reactions.
- Face and voice tracking: Detects multiple faces in each video frame. Matches voices to faces or analyzes them separately. Shows which person is speaking and when.
- Emotion graph: Shows real-time facial emotions on a dynamic circular graph. The Emotion Wheel gives a clear visual of how emotions change.
- Personality trait analysis: Uses the OCEAN model (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) to summarize the personality traits of people in the video. Results are shown as a simple color-coded bar chart.
- Psychologist review: Trained psychologists review the AI results to find hidden biases and emotional triggers. This adds valuable insight to AI analysis.
RightFlow
RightFlow is an emotion AI tool that analyzes facial expressions to understand how people feel during their experience with a brand. It helps businesses capture emotions like happiness, anger, fear, or surprise to improve marketing, customer service, and product design.
Key features
- Hot zone detection: Identifies where people spend time and what grabs attention.
- People count: Tracks how many people interact with a space or product.
- Demographic analysis: Captures age and gender to understand audience differences.
- Attention analysis: Measures head and eye movement to learn what customers focus on.
Unlike tools focused on emotion detection, RightFlow combines emotion data with customer counting, demographic tracking, and physical safety features. It’s designed for public spaces, stores, or events where real-time, contact-free analysis matters.
MoodMe Face AI Emotion Detection Engine
MoodMe’s Face AI Engine is a tool that reads facial expressions to detect emotions in real time. It works directly on the user’s device, with no internet connection or cloud processing needed.
Key features
- Demographic detection: The engine can estimate gender, age, ethnicity, and hair type. This helps apps better understand who is interacting with them.
- Face matching: MoodMe includes a built-in tool for face identification. It can match a face to stored templates locally for secure identity checks.
- Unbiased and inclusive: The AI is trained on diverse data to avoid favoring any group. This ensures fairer results across different faces and expressions.
- Privacy-first: All processing happens on the user’s device. Faces are never stored or sent to the cloud. This protects privacy and meets strict data regulations.
Affectiva AFFDEX
The Smart Eye Group provides software for analyzing emotion and products with a wearable design. Affectiva AFFDEX 2.0 is a toolkit aimed at analyzing the facial expressions of individuals in real-time. It is designed to analyze facial action units (AU) and head pose to track faces, and detect emotions.
Key features
- Multiple face tracking: The tool can process multiple face at the same time.
- Facial expression: AFFDEX 2.0 recognizes 9 basic emotions from facial landmarks (e.g., outer eye corners, nose tip, and chin). The tool does not process speech to categorize emotions.
- Blink rate: It detects some other facial expressive metrics (e.g. blink, valence, and attention).
Hume Empathic Voice Interface (EVI)
Hume’s Empathic Voice Interface (EVI) is a speech-to-speech AI system that makes conversations sound more human. It lets users create, clone, and control voices that respond in real-time with emotion and personality.
Real-life experience
In tests, conversations with EVI felt lifelike and engaging. Emotion detection worked well. Users could guide the tone and setting, although this feature didn’t always perform perfectly.
In short, Hume’s Empathic Voice Interface combines fast response, emotional depth, and high control, making conversations with AI sound closer to real human interaction. The web interface of the conversation platform is simple and intuitive to use.
Figure 6. Hume EVI analysis of conversation with AI
Key features
- Custom voice: Supports over 100,000 custom voices, each with unique traits. You can even create voices like a “calming British matriarch” or an “excited Caribbean musician” by typing a prompt.
- Clone a voice: Upload an audio sample to create a digital version of your own voice.
- Real-time conversations: Responds in about 300 milliseconds, about as fast as a human.
Hume Octave
Hume Octave is a voice-based language model that understands the meaning behind words. The company claims that it helps to create conversation with better emotion, rhythm, and tone.
Real-life experience
Octave often found the right voice for a prompt. It helped improve voice descriptions and matched tones well. However, the final voice sometimes sounded flat or artificial, like a weak acting performance. Still, the tool showed strong potential in capturing different speaking styles.
In short, Hume Octave brings meaning to voice. It helps users create more lifelike, expressive speech that fits both the words and the moment, and it is easy to use.
Key features
- Low latency: Starts speaking in 200 milliseconds with Instant Mode.
- Custom voices: Create voices from scratch, use your own voice, or pick from many pre-made options.
- Expression control: Add acting-style instructions to shape how the voice delivers each line.
- Unique Voices: With a simple prompt, build voices like a “sarcastic medieval peasant” or “calm science teacher.”
Revoicer
Revoicer is an AI-powered text-to-speech software with emotion recognition technology that turns written text into realistic voiceovers. It claims to create audio content with emotional tones that sound more human and less emotion AI technology.
Key features
- Emotional voices: Revoicer can speak in tones such as cheerful, sad, angry, friendly, whispering, or excited.
- Wide language support: It works in English and over 40 other languages, including French, German, Arabic, and Mandarin.
- Custom options: Users can change the voice’s pitch, speed, and tone. They can also add pauses or emphasize specific words.
- Many voices: The tool offers more than 80 voices, including male, female, and child voices. Users can also choose from different English accents like American, British, Australian, or Indian.
Evaluation criteria
To evaluate each Emotion AI tool fairly, we used the same set of criteria across all platforms. These include:
- Accuracy of emotion detection: How well the tool identifies emotions such as happiness, anger, or surprise from facial expressions, voice, or text.
- Multi-modal capabilities: Whether the tool can analyze multiple input types (e.g., video, audio, text) together or separately.
- Ease of use: How intuitive the interface is for non-technical users, including setup and everyday use.
- Real-time feedback: Whether the platform can provide instant insights during live interactions or recordings.
- Depth of insights: Quality and detail of the emotion analytics, including behavioral patterns, attention tracking, and demographic breakdowns.
Further readings
- Sentiment Analysis Benchmark
- Top Methods for Audio Sentiment Analysis
- Open Source Sentiment Analysis Tools
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{phd2026,
author = {PhD., Ezgi Arslan,},
title = {{Top Emotion AI Tools Tested}},
year = {2026},
month = aug,
howpublished = {\url{https://aimultiple.com/emotion-ai-tools}},
note = {AIMultiple. Retrieved August 27, 2026}
}Results and timestamps of 10 data points. Download the data used in this article as a ZIP file containing one CSV file.






Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.