← All posts / Models

The Face That Fooled Half the Room: Tavus Griffin Passes the Video Turing Test

Tavus's Griffin-Lite convinced 48% of study participants it was human on live one-minute video calls, and tops NVIDIA's VideoFDB leaderboard — but the details cut both ways.

The Face That Fooled Half the Room: Tavus Griffin Passes the Video Turing Test

On October 1, 2026, Tavus introduced Griffin, which it calls the world’s first Human Interaction Model (HIM) — a full-duplex, video-to-video system that listens, watches, decides when to speak, and generates its own face, voice, and gestures in real time. The company’s headline claim is blunt: Griffin is the first model to pass the real-time video Turing test. In a live study, 48% of participants who talked with Griffin for one minute believed they had been talking to a real human.

That number deserves scrutiny, and it gets scrutiny below. But even after discounting the marketing, what Tavus has demonstrated is a genuine architectural leap in how machines hold conversations — and the independent benchmark data behind it is strong enough that the industry should pay attention.

What Griffin actually is

Most “video agent” systems today are cascades: automatic speech recognition feeds a language model, which feeds a text-to-speech engine, which finally drives an animated avatar. Each stage waits for the previous one to finish. The result is the familiar turn-taking rigidity — the system can’t nod while you speak, can’t interrupt you gracefully, and leaves dead air whenever it needs to think.

Griffin collapses that pipeline into two tightly coupled engines running concurrently:

  1. Continuous Conversational Modeling — perceives incoming audio and video, and reassesses the conversation at regular sub-second “mini-turns,” deciding what to say, how to say it, and whether to say anything at all. Crucially, it emits expressive controls — emotional tone, stance, facial expression, gesture — not just words.
  2. Audio-Visual Generation — a pair of streaming generators (speech and video) that convert those controls into voice and pixels as they arrive, rather than waiting for a complete utterance.

Because perception never stops while the model speaks, a mid-sentence decision — a pause, a glance away, a smile, an interruption — shows up in the voice and face within the same mini-turn. There is no hand-off between systems. The model can backchannel (“mm-hm”) while you’re still talking, stop the instant you cut in, and react to things it sees, like a Rubik’s cube being solved off-camera.

The numbers behind the claim

On the video side, Griffin-Lite generates 720p video in 320 ms chunks, one latent at a time, from a single reference image. Its VAE compresses time 8×, so at 25 fps each latent covers eight frames. On H100 GPUs, average audio-to-video latency is 0.43 seconds — half that of the next fastest published streaming diffusion model. Speech is generated through a custom codec called Tavec, a convolutional autoencoder that maps 48 kHz audio into a compact continuous latent (40 values per frame, 100 frames per second, no codebooks), with a fully causal decoder that can stream audio in packets as small as 10 ms. Voice cloning requires only about 10 seconds of reference audio.

The independent validation comes from NVIDIA’s VideoFDB benchmark, which scores full-duplex audio-visual conversation on two tracks. On the generation track (does the model produce the right behavior — fluency, affect matching, nonverbal timing), Griffin-Lite scored 3.83 out of 5, more than a full point ahead of the next-best system (Gemini 2.5 + Anam at 2.80) and only 0.09 below the human reference at 3.92. On the perception track (does the model understand the moment, using visual grounding rather than words alone), Griffin-Lite scored 3.73, leading all 15 evaluated models including Gemini 2.5 Flash Native (3.17) and OpenAI’s gpt-realtime (2.97). It also posted the highest takeover-rate alignment on both tracks — 62.8% and 73.8% — meaning its decisions about when to speak best matched human conversational timing.

The caveats are real

The 48% figure comes from Tavus’s own study, not an outside evaluator. The setup: 54 participants recruited through an unnamed research platform were told they’d have a one-minute call with “another participant” about what they were looking forward to that year. The partner was Griffin-Lite. Twenty-six of the 54 — 48% — said afterward that their partner was a real person. By comparison, Tavus’s previous production stack (Phoenix-4.5 + Sparrow-2 + Raven-1) fooled just 1 of 41 participants, or 2.4%.

Three details temper the headline. Participants were primed to believe the partner was human, which inflates pass rates in any Turing-style test. Those who believed averaged 79% confidence; those who said AI averaged 81% — the skepticism wasn’t weak. And people who suspected AI tended to decide within the first 20 seconds, while the conversation’s lowest self-reported rating (4.9 on a 7-point scale) was for “flowing naturally.”

The benchmark’s latency column is the other honest counterweight: Griffin’s median response delay on the perception track is 2,232 ms, slower than the open-source MiniCPM-o at 720 ms and the human reference at 1,400 ms. Griffin wins on quality of perception; it is not yet the fastest on the block.

Tavus is also being deliberately cautious with access. Griffin-Lite is a research preview restricted to select early testers, with no published pricing and no general customer access until the company ships the disclosure features it says are required. A more powerful model is promised to follow.

Why it matters

The leap from 2.4% to 48% in a single generation is the signal that survives scrutiny. Even granting every methodological quibble, no previous conversational video system has come close to convincing half its interlocutors it was human in live, open-ended, face-to-face conversation. The architectural lesson is that duplex, joint audio-visual generation — not better cascades — is what closes the gap.

The implications split cleanly into promise and peril. On the promise side, agents that can read expressions, honor pauses, and share turn-taking naturally unlock genuinely ambient computing: remote support that feels present, language practice partners, accessibility interfaces for people who struggle with text. On the peril side, a model that passes a video Turing test is, definitionally, a deception engine until disclosure is enforced. Tavus seems aware — access is gated on safety work — but the capability now demonstrably exists, and it will not stay confined to one vendor’s careful rollout forever.

The next twelve months will likely be a race between the “more powerful Griffin” Tavus has promised and the disclosure, watermarking, and consent infrastructure needed to keep synthetic faces honest. The 48% says the capability side of that race is already moving very fast.