Realtime avatars that render on the user's iPhone.
There is no datacenter in this loop. Audio goes in and animated frames come out, on the Apple Neural Engine, at 25 fps. We never receive a frame, so there is nothing on our side to count.
iOS only today. Custom characters are a concierge bake, not self-serve. We render a mouth region onto a living host, not arbitrary full-body scenes. All three are on this page — the Limits section is in the nav for a reason.
A server draws every frame. Or the phone does.
A streamed avatar runs a GPU session for as long as your user is talking: it animates, renders, encodes, and ships you H.264. That session is the thing you are billed for. We ship the model once and the phone does all of it — so there is no session, no stream, and no unit of consumption on our side to meter.
Ship the SDK
Add the package. Models arrive over an authenticated OTA channel on first launch rather than sitting in your binary, so your download stays small and the weights are not in a decompiled .ipa.
Feed it audio
Any source — your TTS, our on-device TTS, or a live voice stream. A ~6.8 MB audio encoder drives the mouth geometry. Bring whatever LLM you like; we never see that call.
Frames appear
The renderer composites a 320 px mouth region onto a living idle host at 25 fps on the ANE. No round trip, no jitter buffer, no session for anyone to bill.
Published list rates, with the receipts.
Every figure below is the vendor's own published rate for their realtime product, retrieved 26 July 2026, linked to the page it came from. Where a vendor does not publish a rate, we say so rather than estimate.
| Provider | List rate / streamed minute | Renders on | Renders with no network | Source, retrieved 26 Jul 2026 |
|---|---|---|---|---|
| Yoob | n/aNo streamed minute exists. Plans are $0 / $499 / $2,500 per month. | User's device (iOS) | Yes — rendering and speech.Your LLM still needs a link. | this page |
| Tavus | $0.37 Starter$0.32 Growth | Vendor GPU | No | tavus.io/pricing |
| Beyond Presence | €0.35 Starter€0.175 Scale. Euro is the published figure. | Vendor GPU | No | beyondpresence.ai/pricing |
| Runway Characters | $0.202 credits per 6 s at $0.01/credit, plus $0.02 per session. | Vendor GPU | No | docs.dev.runwayml.com |
| HeyGen LiveAvatar | $0.10 Lite$0.20 Full. Stated list credit value; the entry plan works out to $0.127. | Vendor GPU | No | help.heygen.com |
| Anam | $0.16 Starter$0.11 Professional. Monthly plan prices render as placeholders on their page; we did not read them. | Vendor GPU | No | anam.ai/pricing |
| bitHuman | $0.01 self-hosted$0.04 cloud-hosted. | Device or cloud | No — one billing heartbeat per minuteTheir docs, not our characterisation. | docs.bithuman.ai |
| Spatius | $0.009 Starter$0.007 Scale. The cheapest verified rate here. | Vendor GPU | No | spatius.ai/pricing |
| Simli | Not publicly listedPricing page returned 404 when checked. Third-party estimates span more than 5×, so we publish none. | Vendor GPU | No | simli.com |
| D-ID | Not publicly listedPricing page did not render to our fetch. Their docs do confirm usage is rounded up to the nearest 15 seconds. | Vendor GPU | No | d-id.com/pricing |
How we compare. Each figure is the vendor's own published list or pay-as-you-go rate for their realtime streaming product at the vendor's entry paid tier, in the vendor's own currency, retrieved 26 July 2026. Most are overage rates that apply after a bundled monthly allowance, so a customer inside their allowance pays a lower blended rate, and actual pricing varies with volume and negotiated terms. Beyond Presence publishes in euros; we print euros.
What we left out, and why. Offline video-generation products are not comparable to a realtime stream, so Hedra's Character-3 and HeyGen's Video Agent endpoint are not here. Where a vendor does not publish a realtime rate we say so — guessing at a competitor's price is how you earn a demand letter, and it would undermine the only thing this table is for.
bitHuman is the row worth reading twice. They render on-device too, and they still charge $0.01 a minute, because their SDK posts a billing heartbeat every minute. On-device rendering is not the rare thing. Not having a meter is.
Our own costs are on this page too. Plans are $0 / $499 / $2,500 per month, a custom character is $3,000 once, and your LLM provider bills you directly for tokens — we do not resell or mark that up. Spot an error? pricing@yoob.com, and we will correct it.
Arithmetic over the rates above.
Our column goes up when you add apps and characters. Competitor columns use whichever of their published tiers would actually be cheapest at your volume. No savings figure is computed for you — the totals are here, do the subtraction yourself.
What Yoob is not.
A competitor would raise these in your first call. Here they are first.
iOS only
The Neural Engine is what makes this fit in tens of megabytes at 25 fps with no server, and that pipeline does not port for free. Android is on the roadmap and we will not print a date we cannot hold. Web is further out. If you need Android this quarter, Spatius ships SDKs for both and bitHuman ships Android — use them, and come back when we ship.
Characters take days, and cost $3,000
There is no self-serve path. Runway makes a character from one image because they render on their GPUs; we have to land an identity inside a model small enough to run on your user's phone, then check it across the full pose range before it reaches you. Stock characters ship same-day and Dev includes one.
A mouth region, not a scene
We composite a 320 px mouth onto a living idle host. That constraint is precisely why it fits on the device. If you need a full-body character walking around an arbitrary environment, this is the wrong product and we would rather you know now.
What is and isn't local
- ANIMATIONAudio in, mouth motion out — on the device. There is no animation server in the loop. This is the part that costs everyone else money.
- RENDERFrames composited on the Neural Engine. No video stream, no WebRTC, no jitter buffer.
- VOICEOn-device TTS. If you bring a cloud voice instead, that part is a network call — your choice, not our constraint.
- BRAINWhatever LLM you point us at. Cloud model, cloud call, your key, your bill. We are not going to tell you the whole stack is offline when your brain is in someone's datacenter.
Numbers we have not published yet
- QUALITYWe have not put Yoob side by side with a streamed competitor on the same script. Until we do, assume a datacenter-rendered head can look better and ask us for a device build.
- BATTERYNo measured curve yet. A render costs battery; so does an hour of continuous H.264 decode over a saturated radio. We will publish both, measured on the same device, rather than argue.
- BUNDLE≈32 MB of avatar models. On-device voice is a separate ≈87 MB. The real thinned download for your app depends on what you include, and we will measure it for yours.
We don't sell minutes.
Minutes are what you charge for when rendering costs you money. Rendering doesn't cost us anything, so we charge for the two things that do: supporting your application, and training each identity.
- Unlimited rendered minutes
- Up to 100k MAU
- 1 stock character, no watermark
- SDK updates and iOS-version support
- Free re-bakes when we improve the model
- Email support
- Unlimited rendered minutes
- Up to 2M MAU
- Unlimited apps
- 5 characters included
- Priority bake lane, SLA
- Shared Slack channel
- Unlimited rendered minutes
- Unlimited MAU
- Private characters
- All network calls off by default
- Model-weight and source escrow
The ones that decide it.
You'll add per-minute pricing after your Series A, right?
Fair question. Note first what cannot change: a build you have already shipped cannot be re-priced. The renderer is compiled into your binary and it does not ask us for permission to draw a frame.
Beyond that it is a commitment, not a mechanism — bitHuman renders on-device and meters anyway with a once-a-minute heartbeat, so we are not going to insult you by claiming metering is impossible. Ours is contractual: no per-minute, per-session, per-render or per-MAU charge on any tier; ninety days' notice on any list-price change; and apps you have already shipped stay on the terms they shipped under.
Other providers render on the client too. How is this different?
Some do, and the table says so. Two shapes are worth separating. bitHuman renders on-device and still bills $0.01 a minute, because their SDK posts a billing heartbeat. Others stream compact animation state and rasterise it in the browser — better than shipping H.264, but a server is still generating the motion for every frame of every session.
The expensive, hard part is turning audio into motion. In those systems it happens in a datacenter. In ours it happens on the phone, and no packet about it reaches us.
Is it really zero network, or zero-ish?
Zero-ish, and here is where the "ish" is. Face rendering and speech never touch the network — no audio, no video, no frames, no transcripts. Your LLM does need a link, and that call is yours: your provider, your key, your bill, and we never see it. Model delivery is one authenticated fetch on first launch.
Ask us for a packet capture from a real session and run your own. That is the only version of this answer worth anything.
So the user pays per minute — in battery.
Partly true, and we would rather measure it than argue. We have not published a curve yet, and we are not going to invent one for a landing page. What we can say is that the honest comparison is not against zero: a streamed avatar spends the same user's battery on continuous H.264 decode and a busy radio for those same minutes, plus their data. We will publish both sides, measured on the same device, including the oldest one we support.
What happens in minute 20?
The phone gets warm and iOS throttles. That is physics, and any vendor telling you otherwise is not measuring. Running on the Neural Engine rather than the GPU is what makes sustained sessions viable — it draws materially less power and does not fall off the way a GPU does under continuous load. When thermal pressure gets serious the renderer steps down and then falls back to the idle host loop with audio continuing, rather than dropping frames unpredictably. You can override that policy.
Your models sit in my app. Can someone extract them?
Yes, eventually, and we will not tell you otherwise — a determined attacker can pull a Core ML model out of any iOS app, and compiling it is not protection. What we do is raise the cost and make theft traceable: weights are never in your binary, they arrive over an authenticated channel and are encrypted at rest, and each licensee's copy carries a quality-neutral fingerprint so a leaked model identifies its source.
What an attacker actually gets is a quantised renderer for one identity — not the pipeline that produced it, and not anyone else's character.
What does $499 a month buy if it runs on my hardware?
A reasonable challenge, and worth answering plainly rather than hiding. It buys SDK updates and iOS-version support — an OS release that changes Core ML behaviour is our problem, not yours; re-bakes of your character when we improve the model, at no extra cost; and support with an actual human. If you stop paying, what you have already shipped keeps working. We would rather you leave and come back than be locked in by something that breaks.
What happens to my app if Yoob shuts down?
Apps already in the App Store keep rendering — the models are on the device and the renderer does not check in. New installs need one provisioning call for the character key, so for Enterprise we ship pre-provisioned keys plus model-weight and source escrow, and that call never happens.
Compare that to the alternative: if a streamed vendor goes dark, every avatar they power stops mid-sentence.
No meter. Shipping on iOS today.
Tell us what you're building and we'll send a device build, so you can check every number on this page yourself.