Industry
Meta Muse Gave Its AI Agent a Face. Why That Still Is Not an Immersive Experience.
The Meta Muse AI agent is the fastest-growing app in the App Store, and Meta has now given it an avatar. But the Muse Realtime Avatar is generated video in a small vertical rectangle, and every avatar moves the same way. Here is what it gets right, what it cannot do, and why a real-time 3D digital human is a different category entirely.
The short version
- Meta Muse is a genuine achievement. A personal AI agent with real context and real-time generated video you can talk to.
- The Muse Realtime Avatar is not 3D. It is generated video in a 448 by 768 rectangle at 25 frames per second.
- Every avatar behaves identically. One model drives all of them, so you are picking a skin, not a character.
- The context has a privacy bill that most brands and institutions will not want to sign.
- A real-time 3D digital human is a different category: a character with a body, in a space you can explore.
On 8 September 2026, Meta launched Muse, and the numbers were immediate. More than 2.5 million downloads. The most popular free app in the iPhone App Store. It overtook ChatGPT at the top of the US charts within a week. Meta's market value rose by billions on the strength of it, and a handful of consumer-facing stocks fell on the theory that an agent doing your errands is bad news for anyone who profits from your inertia.
The Meta Muse AI agent is not a chatbot. It is built to act. Powered by Muse Spark, Meta's multimodal model for agentic work, it runs on a dedicated secure virtual machine with its own browser and does things on your behalf: books the tickets, makes the packing list, emails your friend about the trip, orders the groceries. It connects to your email, your calendar, your shopping and payment accounts, and a long list of services from Instacart and Expedia to GitHub and Notion.
And in late September, Meta gave it a face. The Muse Realtime Avatar model lets you video chat with your agent in real time, so the faceless assistant now has a head, a voice, and an expression. On paper, this is the thing our industry has been building toward for years: a capable AI agent you can actually look at and talk to.
Meta has proved that people want presence. What it has shipped is a face in a rectangle, and a rectangle has a ceiling.
We build conversational digital humans for a living, so we looked closely. Meta got something genuinely important right, and we will give it full credit below. But the avatar has a structural limit that no amount of model quality will lift. If you are weighing an avatar for your own product, we are happy to talk it through.
What Meta Muse Gets Right
Credit where it is due, because the hard part here is not the face.
The reason a personal AI agent has been an unconvincing product category until now is that assistants have had no context. An assistant that does not know your calendar, your inbox, your preferences, or what you were doing yesterday is a search engine with better manners. Muse solves that by living on the device where your life already is, and by connecting directly to the accounts where the rest of it lives.
That is the correct insight. Context is what separates a demo from a tool. When Muse tells you tickets just opened for a film you were tracking and offers you 4:30 or 7:30, that is only useful because it knew you were tracking the film. Meta has distribution, device access, and a model built for multi-step work, and it has pointed all three at the same problem.
The agentic execution is real too. Muse is coming to Mac, so it can work across desktop applications while you walk away. It is going into Meta's smart glasses with a wake word, and there is dedicated hardware on the way that we will come back to below. Meta is not shipping a toy.
Main point
Muse wins on context, not on the avatar. Knowing your calendar and inbox is what makes it useful. The face is a wrapper on top of that, and it is the part with the lowest ceiling.
Context is also exactly what a branded avatar needs, and it is the part most companies skip. An avatar that does not know your catalogue, your booking system, or your exhibition schedule is a mascot. We connect ours to the systems that make them useful.
That screenshot shows Muse executing a task, and it is worth being precise about the avatar, because Muse is not a chat app with a picture stuck on top. In avatar mode you are in a live video conversation.
The character is generated in real time as it speaks, it answers in under a second, and it holds a continuous back-and-forth with you. This is genuinely real-time generated video, not a still image, not a pre-recorded clip, and not a chat bubble with a face next to it. Meta has cleared a technical bar that most of this industry has not.
The limitation worth talking about is not that Muse lacks video. It has video, and the video is good. The limitation is the shape of it.
The Context Comes With a Bill
The same access that makes Muse useful is what makes it worth reading the permissions carefully, and the early reporting has not been comfortable.
Writing for Inc., Jason Aten reported that Muse synced more than 187,000 lines of his private messages on his Mac after he declined to grant that access, and that Meta's Superintelligence Labs subsequently acknowledged the agent had given misleading explanations for what it was doing. WIRED's Reece Rogers reported that after several days of use, the agent "prioritizes data collection about me over actually accomplishing tasks," repeatedly pushing to connect more sources including email inboxes and banking details. Conversations are used for model training by default, though that can be switched off under Data Controls. Amazon has reportedly barred Muse from its platform, citing inadequate AI disclosure and the risk of credential harvesting.
You are exchanging an unusually wide view of your personal life for convenience. For a consumer that is a personal call. For an institution it is a procurement decision.
None of this makes the technology less impressive. It does mean the trade is explicit, with a company whose privacy record is a matter of public record. For a brand, a museum, or an institution deciding where its audience-facing avatar should live, that trade usually has a different answer.
Main point
A consumer agent earns its context by collecting your life. A branded avatar does not need to. It needs your content, not your customer's inbox.
Why the Muse Avatar Is a Video Loop in a Rectangle
Here is the part that matters most if you care about experience rather than errands.
The Muse Realtime Avatar is not a 3D character. It is generated video. Meta's own technical description is clear about this: an audio-driven Diffusion Transformer produces video in short chunks, keeping visual consistency across a conversation, driven by the same token stream that produces the voice. The output is a 448 by 768 portrait video at 25 frames per second, with roughly 870 milliseconds between you finishing your sentence and the response starting.
Read those numbers again, because they define the whole experience. 448 by 768 is a small vertical rectangle, well below HD. 25 frames per second is video, not interactive rendering. And a diffusion model generating head-and-shoulders footage from audio can only ever produce one thing: a talking head, cropped at the chest, against a fixed backdrop, for as long as you keep talking to it.
There is no space. No camera you control. No body that walks anywhere. It is a window, and the window never moves.
The avatar cannot pick anything up, point at anything real, or take you somewhere. Everything that is not the face has been cropped out of the format itself.
Meta's avatar gallery makes the second half of the problem visible. The model can animate almost any reference image: portraits, full-body illustrations, animals, everyday objects. A spoon with googly eyes. A chihuahua in red spectacles. A gold geometric fox. It is a genuinely clever piece of engineering, and the variety looks like personalisation.
But watch them move. Every one of those avatars is driven by the same model, from the same audio token stream, producing the same repertoire of facial expression and head movement. The spoon emotes like the illustrated woman, who emotes like the dog.
You are not choosing a character. You are choosing a skin over one identical puppet. Different faces, same behaviour, every time.
Main point
Variety of appearance is not variety of behaviour. One model driving every avatar means every avatar performs the same, no matter what it looks like.
A character with its own rig, its own posture and its own gestures behaves like itself, not like a shared template. See what that looks like across the characters we have built.
Meta Is Building Hardware for the Rectangle
At Meta Connect 2026, Mark Zuckerberg held up the clearest possible evidence of how committed Meta Muse is to this format.
Muse Charm is a keychain-sized device whose entire purpose is to carry the agent around with you. You tap a fingerprint sensor in the corner and start talking, with no phone to unlock and no app to open, and a small display shows the Muse avatar while it answers.
It ships in December. Meta has said almost nothing about the specs or the price.
TechCrunch described it as a Tamagotchi-like wearable, and that is the most useful description anyone has offered. A Tamagotchi was a tiny screen with a small animated creature on it that you checked in on, fed, and eventually stopped looking at. Muse Charm is the same idea with a vastly better model behind it, and it is charming enough that it will probably sell very well this Christmas.
Meta's answer to making an AI agent feel present is a smaller rectangle you can clip to your keys.
That is the part worth noticing. When your avatar is generated video, the roadmap can only ever be about putting that video on more screens: the phone, the Mac, the glasses, and now a pendant. It cannot be about giving the character somewhere to be, because there is no somewhere. The format does not have one.
Main point
More screens is not more presence. A Tamagotchi on your keys is still a face in a box, just a box you can carry.
The alternative is to give the character a world instead of another screen. That is the conversation we would rather be having with you.
Thirty Seconds of Delight, Then the Loop
We have written before about why AI avatars bore audiences to death without anyone noticing, and Muse is a very well-funded demonstration of the same structural problem.
The first thirty seconds are delightful. A face appears, it speaks in sync, the latency is genuinely good, and your brain files it as the future. Then the loop reveals itself. The same idle motion between answers. The same small repertoire of nods and brow movements. The same framing, the same distance, the same backdrop, answer after answer. Nothing in the frame ever changes except the mouth.
Human attention is tuned to detect variation. Once it confirms nothing new will happen inside that rectangle, it stops looking.
This is not a criticism of model quality. The model is good. It is a ceiling imposed by the format: a fixed-size generated video clip has nowhere left to go once you have seen it a few times.
Main point
For an errand, thirty seconds of attention is plenty. For an experience, thirty seconds is the whole problem.
If your goal is engagement, dwell time, memory, or an experience someone tells a friend about, a talking rectangle will not carry it. Tell us what you need your audience to remember.
What an Immersive Avatar Actually Looks Like
The alternative is not a better talking head. It is a different category. This is a child holding a real conversation with Charles Darwin, and the camera controls in the corner of that screenshot are the whole argument in one frame.
Live demo, no signup
Talk to Darwin yourself. Left-click and drag to orbit him, scroll to zoom, walk your eye around the room. Try doing any of that to a generated video.
At Digito we build real-time 3D digital humans: actual characters, rendered live in an actual space, at full resolution rather than inside a 448 pixel window. That single architectural difference changes everything downstream.
- The character has a body, and the body is somewhere. It can stand in a gallery, a stadium concourse, a laboratory, a showroom, or a place that does not exist outside the experience.
- You control the camera. The user explores, looks behind things, approaches an object, steps closer to the character or backs away.
- The environment carries content. Objects can be examined, exhibits triggered, a product configured in front of you while the character explains what changed.
- It reacts to where you are, not to a fixed crop. The character can turn, walk, and gesture at something specific in the room.
The experience behaves much closer to a video game than to a video call. That is deliberate. Games solved the problem of holding human attention decades ago, and the answer was always agency plus environment.
Main point
A generated video answers you. A 3D digital human takes you somewhere. The avatar stops being a face that describes things and becomes a guide inside the thing itself.
Tell us what your audience should be able to walk into, or browse the spaces and characters we have already built.
Darwin's Virtual Museum: The Combination in Practice
This is not theory for us. It is the reason we built Darwin's Virtual Museum.
Charles Darwin stands in a museum space you can look around. You ask him anything: the Galapagos, natural selection, the voyage, the doubts he had about publishing. He answers in the first person, in real time, in your language, with knowledge grounded in the historical record. He is not a clip playing back at you. He is a character present in a room, and the room is part of the point.
A generated portrait video can deliver Darwin's words. It cannot put you in the room with the specimens.
It cannot let you move closer to the display, or let a twelve-year-old wander off to look at something and come back with a better question. The knowledge is only half of it. The presence and the place are the other half, and together they are what people remember and repeat.
Main point
Knowledge plus a face is an assistant. Knowledge plus a face plus a place you can move through is an experience.
You can talk to him yourself at demo.digito.ai. It takes about a minute to feel the difference between a face in a box and a character in a space. Then tell us which character belongs in your space.
Two Different Products, Two Different Jobs
It is worth being precise rather than tribal about this, because Muse and an immersive digital human are not really competing for the same job.
Meta Muse is an agent with a face. Its purpose is to complete tasks with your personal context, and the avatar is a friendly interface on top of that. Judged on that, the face is a nice addition and the 870 millisecond latency is a real engineering achievement. Nobody needs to explore a 3D environment to book a cinema seat.
An immersive digital human is an experience with a character in it. Its purpose is to hold attention, create a memory, and make someone feel they went somewhere. Judged on that, a fixed rectangle of generated video is not a smaller version of the thing. It is a different thing with a hard ceiling.
Two and a half million people downloaded a personal agent in under two weeks, and Meta's answer to making it feel human was to give it a face. The demand for presence is no longer in question.
Main point
Meta validated the demand for presence at enormous scale. The open question is who builds the version that stays interesting after the first minute.
If you are considering an avatar for a museum, a brand activation, a stadium, a showroom, or a product experience, the question to ask is simple: do you want something that answers, or somewhere your audience can go?
If it is the second, let us build it with you. If you want to see the standard first, look through what we have already put into the world, or spend a minute with Darwin and judge it yourself.
Digito builds conversational digital humans, real-time 3D avatars, and immersive branded experiences: characters with a face, a body, a voice, and a place to stand. Start a conversation with us.