Inter-2 Audio

Hear more

Social signals

Each of the 6 signals as it starts and ends, with a probability (high, medium or low) and a rationale naming the cues the model heard.

Engagement

Engaged, neutral or disengaged, sent the moment it shifts, so you always know where the person stands.

Conversation quality

A 0–100 score for how well the conversation is going, built from clarity, authority, energy, rapport and learning. For the whole session and over time.

Get more insight from the way people speak.

AI has learned to understand our words, but speech carries information far beyond the transcript. The way someone pauses, changes their tone, emphasises a word or moves through a sentence can reveal confidence, hesitation, uncertainty, engagement and more. Inter-2 Audio brings these signals into AI by interpreting what is said alongside how it is delivered, using audio alone.

Signals detected using Inter-2 Audio, https://api.interhuman.ai/v2/upload/analyze.

More than what was said

A transcript captures the words, but it leaves out much of what happens around them. Inter-2 Audio combines lexical information with acoustic cues including intonation, loudness, emphasis, rhythm and timing, allowing AI to identify social signals from the sound of speech itself.

What makes this particularly powerful is how much information can be extracted without seeing the speaker at all. Inter-2 Audio can identify signals such as confidence, hesitation, uncertainty and frustration from audio alone, despite having none of the visual information available to a multimodal model. It considers different cues together, allowing the broader delivery of speech to inform what an individual pause, word or change in tone means.

Audio opens more possibilities

A huge amount of human interaction happens through audio. Phone calls, voice interfaces, interviews, customer conversations, remote meetings and recordings can all contain valuable social information without a camera ever being involved.

Inter-2 Audio makes that information accessible, opening social-signal intelligence to environments where video is unavailable, unnecessary or simply not practical. The model reaches 54.67 Micro-F1, and 58.08 EW-F1, alongside 82.72% engagement accuracy, on 839 held-out audio clips, ranking first across every reported measure.

EW-F1

Inter-2 Audio reaches 58.08% EW-F1 at 5.48× reference inference speed, the fastest of every model tested.

Evaluated zero-shot on 839 held-out clips against a six-signal vocabulary, with engagement scored separately as a three-class engaged / neutral / disengaged task. 235 of those clips (≈28%) carry no positive social-signal label and are scored through Neutral-F1 instead. Visually dominant signals such as Skepticism and Confusion are excluded, since not all social signals can be inferred from audio alone.

It is also built to deal with the variability of real-world audio, remaining robust across tested conditions with room reverberation and different levels of background noise.

What this makes possible

The value of Inter-2 Audio is ultimately in what becomes possible when AI can access social information through sound alone. Customer conversations can reveal changes in confidence or engagement, research platforms can capture how participants respond during interviews, and voice-based AI can work with more than the words being recognised.

By bringing social signals to audio, Inter-2 Audio gives products another layer of information to work with, even without a camera.