Tavus Griffin Explained: The Human Interaction Model

Tavus Griffin explained: how Griffin-Lite uses full-duplex video, what the Human Interaction Model means, and what the VideoFDB results actually show.

By VioEvo Editorial•Published 2 ottobre 2026•Reading time 10 min

Tags

tavus-griffin
griffin-lite
human-interaction-model
videofdb
ai-video

Create a reference-driven video while you explore Griffin's approach.

Tavus introduced Tavus Griffin on October 1, 2026, as what it calls the first Human Interaction Model (HIM). The announcement is easy to misread as another AI avatar release. Tavus Griffin is aimed at a different problem: making a face-to-face conversation with an AI feel continuous, responsive, and visually aware.

Griffin listens while a person is speaking, watches the video feed, decides whether to wait or respond, and generates speech and expressive video at the same time. The public preview is Griffin-Lite. It is not currently a general-purpose video generator, and it is not available as a normal customer-facing Tavus model.

The video above is Tavus's official Griffin introduction. See the Tavus Griffin announcement for the technical description, evaluation methodology, and release limitations.

The short version

Griffin is a real-time conversational system built around a full-duplex video-to-video loop. In a conventional conversational avatar stack, speech recognition waits for a turn, a language model produces a response, text-to-speech generates audio, and an avatar renderer produces the face. Each stage hands a finished result to the next one.

Griffin keeps listening and watching while it is responding. At short intervals, its conversational model can hold the floor, yield, interrupt, backchannel, change expression, or react to something visible in the user's video. Tavus presents this as a unified interaction model rather than a chain of separate turn-based components.

The distinction matters because the hardest part of a face-to-face conversation is not producing a sentence. It is timing a glance, pause, nod, interruption, or brief acknowledgement while the other person is still speaking.

What is Tavus Griffin's Human Interaction Model (HIM)?

Human Interaction Model is Tavus's name for a model class designed to understand and generate face-to-face interaction in real time. The phrase is new, and it should be treated as Tavus's proposed category rather than an established industry standard.

The category describes a broader capability than an AI avatar that reads a script. A HIM is expected to account for the visual and vocal signals around the words: gaze, facial expression, pauses, gestures, timing, and changes in the conversational floor. It must also generate its own voice and appearance in a way that responds to those signals.

That is why Tavus describes Tavus Griffin as a full-duplex video-to-video model. The model does not simply convert a completed text response into an animated face. It receives audio and video, makes a conversational decision, and drives streaming speech and video generation continuously.

HIM is useful as a label for this research direction, but it does not mean that Griffin has solved human interaction in general. The benchmark results and live study show meaningful progress under specific evaluation conditions; they do not establish that Griffin is equivalent to a person or reliable for every conversational setting.

How Tavus Griffin-Lite is built

Tavus describes Griffin as a two-part system.

The first part is a continuous conversational modeling engine. It receives the person's audio and video, tracks the state of the exchange, and emits controls for what the system should say and how it should behave. Those controls include conversational timing as well as expressive signals such as emotion, gaze, and gesture.

The second part is an audio-visual generation engine. Streaming speech generation produces the voice, while streaming video generation produces the face and movement. Perception, decision-making, and generation run concurrently, so the model can react while it is listening and while it is speaking.

For the Griffin-Lite video generator, Tavus reports the following design:

  • a reference image of the character
  • streaming audio
  • streaming controls for gesture, gaze, and emotion
  • an autoregressive diffusion generator that predicts one latent at a time
  • 720p output produced in 320 ms latent chunks

On H100 hardware, Tavus reports an average audio-to-video latency of 0.43 seconds for the video generator. That number describes the time from incoming audio to the visible effect of that audio in the generated video. It is not a claim that the complete conversation system responds in 0.43 seconds from end to end.

The one-latent-at-a-time design is important. It avoids waiting for future audio before rendering the next piece of video, which lets an interruption or a change in speech affect the face sooner. Older frames can be kept in a compressed history so the system can continue a long stream without treating every frame as a new generation request.

Source: Tavus technical approach and Griffin-Lite evaluation.

What Tavus Griffin can do in a conversation

Tavus's demonstrations focus on behaviors that ordinary talking-head systems handle poorly:

  • acknowledge someone with a short backchannel while they are still talking
  • stop or change course when the person interrupts
  • wait through a meaningful pause instead of treating every silence as a handoff
  • use gaze, facial expression, and gestures as part of the response
  • react to visible context, such as an object the person is showing or an action they are performing
  • maintain a conversational thread while the other person changes pace or direction

These examples are not just cosmetic improvements. In a turn-based pipeline, the avatar normally has no mechanism for adding a nonverbal response during the user's turn because its motion is driven by the speech it has already generated. Griffin's design gives the conversational system a separate way to signal attention or take the floor before a complete spoken response is ready.

The Tavus Griffin video Turing test claim

Tavus reports that Griffin-Lite passed what it calls a real-time video Turing test. In the study, participants were told they would speak with another participant for one minute. They were actually speaking with a PAL powered by Griffin-Lite, which generated the face, voice, and responses in real time.

The reported result was:

SystemParticipants who believed they spoke with a real person
Griffin-Lite26 of 54 (48%)
Phoenix-4.5 + Sparrow-2 + Raven-11 of 41 (2.4%)

The result is notable, but its scope is narrow. The sample was small, each interaction lasted one minute, and the study was run by Tavus using its own protocol. The result supports the more precise statement that Tavus reports 48% of participants believed Griffin-Lite was a real person in this one-minute study. It does not show that 48% of people would mistake Griffin for a human across longer conversations, different populations, or independent tests.

Tavus also reports that participants who identified the system as AI still rated it highly on whether it was listening. That is useful context: perceived conversational quality and mistaken identity are related, but they are not the same measurement.

What NVIDIA's VideoFDB benchmark shows

The Tavus Griffin announcement has a second source of evidence: VideoFDB, a benchmark created by NVIDIA and David AI for full-duplex audio-visual-to-audio-visual conversation.

VideoFDB uses 237 clips from real video calls covering 11 nonverbal conversational dynamics. It separates two questions that are often mixed together:

  1. Perception: Does the system understand what is happening in the other person's audio and video?
  2. Generation: Does the system produce an appropriate voice, face, and nonverbal response of its own?

The published leaderboard reports these overall scores for Tavus Griffin Lite:

TrackGriffin Lite overall scoreTiming alignment
Perception3.73 / 573.8% / 2,232 ms
Generation3.83 / 562.8% / 1,892 ms

On the Generation track, Griffin Lite scored 1.03 points above the next-best compared system, Gemini 2.5 + Anam at 2.80. On the Perception track, it scored 0.29 points above the strongest reported baseline, MiniCPM-o 4.5 at 3.44. Human references scored 3.92 for Generation and 4.20 for Perception, so the model remains below the human reference on both tracks.

The benchmark is valuable because it tests the part of interaction that a conventional avatar benchmark tends to leave out. It measures whether a system uses a pause, gaze shift, laugh, or facial expression to decide what to do next, rather than only whether its words are intelligible.

NVIDIA's analysis also identifies a structural weakness in cascaded speech-to-avatar systems. They can produce a fluent spoken response, but they are effectively turn-based and cannot reliably add appropriate nonverbal cues while the user is still holding the floor. VideoFDB reports cascade latencies of roughly 2.8 to 3.5 seconds behind human ground truth in the evaluated setting.

Source: NVIDIA VideoFDB benchmark and leaderboard.

Tavus Griffin is not a conventional AI video generator

The phrase “video-to-video” can make Griffin sound closer to an ordinary video creation tool than it is. The two systems solve different problems.

GriffinVioEvo's video generation workflows
Real-time face-to-face interactionPre-rendered video creation and transformation
Continuously reads audio and video during a conversationUses prompts, images, references, or existing footage as generation inputs
Produces a responsive AI face, voice, and nonverbal behaviorProduces clips for creative, marketing, social, and production workflows
Griffin-Lite is a research previewAvailable video creation workflows can be started directly on VioEvo

If you came here looking for Griffin access, Griffin API documentation, or Griffin pricing, there is an important limitation: Griffin-Lite is not currently available for ordinary customer use. Tavus says it is available to select trusted testers as a research preview and that broader release depends on additional safety and disclosure work. Tavus also says Griffin is not yet on its regular platform.

If your goal is to create a character video today, the closest VioEvo workflows are different in kind but useful in practice. Start with image-to-video when one character or product image should become a short clip. Use reference-to-video when several reference images need to guide the same scene. Use video-to-video when an existing motion source should be transformed.

What Griffin changes in AI video research

Griffin's main contribution is not a new way to make an eight-second clip. It is a different definition of what a video model should do when the input never stops arriving.

Most AI video systems are judged on a finished clip: visual quality, prompt adherence, motion, consistency, and audio-visual alignment. Griffin adds another target: whether an AI character can share the timing of a live exchange. That requires the system to preserve information about what is happening now, what just happened, and whether the other person is yielding the floor.

The distinction creates two related but separate paths for the field:

  • AI video generation creates controlled visual assets from prompts, references, and source footage.
  • Human Interaction Models create ongoing, responsive exchanges in which perception and generation run together.

They will increasingly borrow techniques from each other. Better streaming generation can improve interactive avatars. Better temporal and social understanding can improve characters in generated scenes. But a benchmark result in one area should not be presented as proof of performance in the other.

Frequently Asked Questions

What is Tavus Griffin?

Tavus Griffin is a real-time conversational model that Tavus describes as the first Human Interaction Model, or HIM. It is designed to understand and generate face-to-face audio-visual interaction, including speech, pauses, gaze, facial expression, and gestures.

What is Griffin-Lite?

Griffin-Lite is the research preview of Tavus's first Human Interaction Model. It is the version used in the reported one-minute face-to-face study and in the VideoFDB evaluation described by Tavus and NVIDIA.

Is Griffin available through an API?

Not for ordinary customers at the time of publication. Tavus says Griffin-Lite is available to select trusted testers as a research preview, while Griffin is not yet on the regular Tavus platform. Availability, API access, pricing, and any waitlist terms may change.

What is the Human Interaction Model (HIM) in Tavus Griffin?

Human Interaction Model is Tavus's term for a model class that understands and generates face-to-face interaction continuously. It is a proposed category, not yet a universally adopted industry standard.

What is the VideoFDB benchmark?

VideoFDB is an NVIDIA and David AI benchmark for full-duplex audio-visual conversation. It evaluates perception and generation separately across 237 real video-call clips and 11 nonverbal conversational dynamics.

Is Griffin an AI video generator?

Griffin generates video as part of a live conversational system, but it is not a general-purpose prompt-to-video or image-to-video product. Its focus is real-time interaction with an AI character, not the creation of standalone video clips.