Crypt0's NewsCrypt0's News

AI

Tavus Shipped Phoenix 4.5, and the AI Human Finally Got a Whole Body

The AI human just got a body. Tavus has shipped Phoenix 4.5, its biggest step forward in real time human rendering, and the headline change is architectural: the model now generates the entire frame in a single pass. Face, head, neck, shoulders, torso, clothing, and surroundings all render together in one continuous image, so expression can move freely across the whole frame. Phoenix 4 generated a roughly 512 by 512 pixel region around the face and composited it onto recorded body footage. That boundary is gone.

The numbers tell the story of a model built for live conversation. Audio to video latency sits at 134 milliseconds, which Tavus calls the fastest on the market, about 25 percent faster than the systems it compares against. Lip sync improved 16 percent over Phoenix 4. In human preference testing, even zero shot Phoenix 4.5 scored 1002 Elo against 929 for fine tuned Phoenix 4, rising to 1069 once fine tuned. Identity fidelity held level while frame and video quality improved. Milliseconds matter when the goal is a conversation that feels alive, and Phoenix 4.5 is winning on every axis that matters for presence.

Creation got dramatically easier too. Phoenix 4.5 is the first Phoenix model that renders a new identity while training happens in the background: a preview is ready in about a minute from a single image or video, while fine tuning runs behind the scenes and takes over roughly two hours later, half the four hours Phoenix 4 required. The model is also far more forgiving of real human appearance. Long hair, glasses, earrings, and headbands used to cause training failures. In Tavus's tests, 64 percent of faces that failed on Phoenix 4 trained successfully on 4.5. All existing stock faces were upgraded, 50 new stock faces were added, and for the first time the system renders cartoons, anime, and stylized characters in real time while preserving their look.

Tavus is explicit about why presence matters: in education, healthcare, coaching, sales, and support, how someone listens, reacts, and explains is part of what makes the interaction work. Builders replacing audio only interactions with face to face conversational AI report 2 to 3 times higher engagement across sales and healthcare, 40 percent higher knowledge retention in learning and development, and 50 percent faster ramp plus 3 times higher rep conversion in sales training. Presence, in other words, is becoming part of the product itself.

The angle everyone else will miss: this is the moment AI humans cross from demo to infrastructure. A model that renders a full bodied conversational human from one photo in about a minute, speaking more than 30 languages, has crossed from research result to deployment primitive for tutors, coaches, concierges, interviewers, and support agents. The same capability raises the bar for provenance and consent tooling, which will now grow up alongside the rendering tech.

For builders, the constraint on video agents just moved from rendering to imagination. For everyone else, expect the next tutor, coach, or support rep you meet online to have a face, a voice, and body language that moves with its words. The uncanny valley is filling in fast, and the winners will be the teams that pair this presence with genuine usefulness.

Sources

New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.

← Back to Crypt0's News