There’s a version of digital humans in XR that you’ve definitely seen and quietly dismissed: the uncanny valley avatar with dead eyes that a company put in a product video and called “revolutionary.” Fair enough to be sceptical. But the gap between that and what’s actually possible now has widened considerably, and some of the enterprise use cases are genuinely holding up in production.

So what’s the actual state of photorealistic digital humans in XR heading into the second half of 2026?

MetaHuman: The Tool That Changed the Accessibility Equation

Epic Games’ MetaHuman Creator didn’t make digital humans good — they were already technically impressive in cinema. What it did was make them accessible. Before MetaHuman, creating a believable digital human required a team of character artists, weeks of work, and budgets that only studios could justify. MetaHuman Creator is browser-based, produces Unreal Engine-ready characters in hours, and the results are legitimately convincing in real-time rendering contexts.

The latest iteration supports full facial performance streaming, which is where things get interesting for XR. You can drive a MetaHuman’s facial expressions using standard webcam input or iPhone Face ID hardware, meaning a real presenter can be replaced — or augmented — with a digital human in real time inside a headset environment. The lip sync, blinking, and micro-expressions are close enough to feel natural at conversational distance.

For enterprise VR training, this opens up scenarios that were previously too expensive. A MetaHuman customer, patient, or colleague can respond dynamically during a roleplay simulation rather than playing from a fixed animation set. Several UK training providers working in healthcare and customer service have started shipping this approach.

Where the Tech Actually Works

The honest answer is that digital humans in XR work best in constrained, high-value scenarios rather than as general interfaces. Here are the contexts where they’re earning their keep.

Healthcare and empathy training. VR training simulations that put clinicians, social workers, and police in conversations with distressed or unwell digital patients have a track record now. The HeadSpace platform and various NHS training partnerships have run pilots with measurable outcomes on communication skills. The digital human doesn’t need to be perfect — it needs to be “good enough to trigger an emotional response,” and that bar is lower than people expect.

Virtual onboarding and learning. Replacing a talking-head video with an interactive digital human presenter who can respond to questions, adjust pacing, and stay on-screen throughout a module increases engagement metrics in most documented trials. Companies including Accenture and Deloitte’s UK training divisions have deployed this.

Customer service in constrained domains. Digital human kiosks in retail, banking, and hospitality environments handle a narrower set of conversations than a general AI chatbot, but the embodied presence increases customer willingness to engage and trust the interaction. UneeQ and Soul Machines both have UK deployments in this space.

What doesn’t work: using digital humans as a replacement for real human connection in high-stakes emotional contexts, expecting uncanny valley tolerance from all demographics equally, or trying to make a digital human handle genuinely open-ended conversation without heavy guardrailing.

NVIDIA’s Audio2Face and the Voice Problem

The hardest part of a convincing digital human isn’t the face. It’s the voice. Specifically, it’s synchronising voice with facial motion in a way that doesn’t feel automated. NVIDIA’s Audio2Face, now integrated into Omniverse and licensable as a standalone service, takes audio input and generates facial animation that matches it with high fidelity — blending lip sync with natural emotional expression markers that humans subconsciously read.

Pair this with a quality text-to-speech voice (ElevenLabs and similar services have become genuinely difficult to distinguish from real speech at normal listening distances) and you have a complete real-time digital human pipeline that doesn’t require a live operator behind it. The latency on the full chain — LLM response generation, TTS, Audio2Face animation, rendering — is still a noticeable beat on current hardware, but for asynchronous or semi-scripted interactions it’s workable.

The Render Budget Reality

If you’re doing XR for mobile platforms — Meta Quest 3, Samsung Galaxy XR — a photorealistic MetaHuman character at close range is a render budget problem. The full-quality MetaHuman render is designed for PC-class or next-gen console hardware. On standalone headsets, you’re making tradeoffs: lower polygon counts, reduced texture resolution, simplified hair, and most significantly, cutting back on the subsurface scattering that gives human skin its life.

The practical outcome is that mobile XR digital humans in 2026 sit in a slightly lower fidelity band than you’d see in a PC VR or mixed reality headset context. Still convincing in most scenarios; just not at the level of a rendered cinematic. Most enterprise deployments that care about fidelity are targeting tethered headsets or PC VR precisely for this reason.

What’s Coming

The trajectory is pointing toward digital humans that aren’t just rendered but are genuinely conversational — not just playing audio with matching facial animation, but understanding context, remembering prior interactions, and adjusting their persona to the user. The underlying LLM and memory infrastructure to support that exists. The latency and cost to run it in real time at scale doesn’t quite work yet, but that’s a 12-to-24-month problem rather than a fundamental limit.

For organisations evaluating XR investments now, the question isn’t whether digital humans will be a useful interface layer — they will be — but whether the use case justifies the current implementation cost versus a simpler avatar approach. For high-stakes training and brand-critical customer interactions, the answer is increasingly yes.