How to add visemes and blinks to a VRM — without sculpting shape keys

A rigged model from Tripo, Meshy or a marketplace walks and waves but can't say a word. Giving it a face that moves means either sculpting shape keys by hand, or building them from the model's own geometry.

To give a rigged model a moving face, upload the .glb (or an existing .vrm) to Riggle and spend about twenty seconds marking the face: trace a polygon around the lips, click a point inside the mouth opening, tag the upper and lower lip, and click each eye centre. Riggle turns those marks into a real lip mask on the mesh, zips the mouth shut into a closed rest pose, and generates the five VRM visemes (aa, ih, ee, oh, ou), blink with independent blinkLeft and blinkRight, happy, angry and surprised, plus eye-bone lookAt when the model has eye bones. You get a VRM 1.0 and a video of every expression playing back. This is the VRM expression set that drives lip sync, blinking and reactions — not the 52-shape ARKit perfect-sync rig, which is on our roadmap.

Two different jobs both get called 'face rigging'

'Perfect Sync' is the VTuber term for full iPhone-grade facial tracking. It means the model carries the 52 ARKit blendshapes — browInnerUp, eyeBlinkLeft, jawOpen, mouthSmileLeft and 48 more — and that each one is mapped into a VRM expression (a BlendShapeClip in VRM 0.x, an Expression in VRM 1.0). Apps like VSeeFace, Warudo, VMagicMirror, Vear and Luppet look those clips up by name, which is why a near-miss on naming produces the classic 'we recognized 45 of your 52 shapes' message.

But most avatars don't need all 52 to be alive on screen. The VRM expression set — the five visemes aa/ih/ee/oh/ou, blink, and the emotion presets — is what drives lip sync from audio, idle blinking and reactions inside a VRM app. If your goal is 'my avatar should talk, blink and react', that set is the whole job. If your goal is 'my eyebrow should follow my real eyebrow through a TrueDepth camera', you need the full 52 and nothing less will do.

Either way, there are two sub-jobs underneath that constantly get confused: (a) creating the shape-key deformations on the mesh, and (b) naming and wiring them into VRM clips. A model can have one without the other, and the symptoms look identical from the app's side.

Why automatic mouth rigging keeps failing

The reason nobody hands you a one-click face rigger for arbitrary meshes isn't laziness. On continuous geometry there is no falloff that can tell a lower lip from a chin. Drive the mouth open with a distance field and the jaw, cheeks and chin smear open with it — the model looks like it's melting, not speaking. We tried four different procedural approaches while building Riggle and every one of them failed in exactly that way.

Which is why the ecosystem's standard advice has always been 'open Blender and sculpt the shape keys by hand': roughly a day per model, repeated every time you edit the mesh. The way out isn't a cleverer detector. It's a few seconds of human input at precisely the point where the machine is blind — and full automation everywhere else.

The fast path for an AI-generated or marketplace model

  1. Upload a rigged model. A .glb from Tripo, Meshy, Mixamo or your own pipeline, or an existing .vrm you want re-rigged. The body rig has to be there already — Riggle maps skeletons, it doesn't create them. It normalises the scale to human height, fixes the model's facing, bakes an A-pose to T-pose, maps the bones onto the VRM humanoid skeleton and renders an orthographic view of the face.
  2. Mark the face — about twenty seconds. Trace a polygon around the lips, click a point inside the mouth opening, tag the upper lip and lower lip, click each eye centre. That's the entire manual step.
  3. Riggle rebuilds the mouth. Your outline becomes a real lip mask: the mesh is welded across glTF UV seams (which fracture edge connectivity on export), the traced region is flood-filled across the surface, the mouth-cavity rim is found, and the lips are zipped shut along a seam biased toward the upper lip so the jaw does the closing, like a real mouth. Rest pose becomes a closed mouth, and the visemes are genuine re-opens of your model's own cavity geometry.
  4. Watch it, then download. Every expression is rendered into a short preview video with per-expression stills, and each shape is pixel-diffed against the neutral render so a dead morph can't slip through. Happy with it? Download a VRM 1.0. Not happy? Re-mark and run it again — re-running is free, and tokens are only spent on downloads.

What comes out, precisely: aa, ih, ee, oh, ou, blink plus independent blinkLeft and blinkRight, happy, angry, surprised, and eye-bone lookAt when the rig has eye bones. VRM 1.0 only — no 0.x export.

Open-mouthed models work best — and that's not a footnote

Riggle's visemes re-open a mouth that already exists, so the input genuinely matters. A model generated with its mouth open carries real cavity geometry — teeth, tongue, an inner lip rim — and that geometry is the asset the whole approach is built on. If you control the prompt in Tripo or Meshy, ask for an open mouth.

A sealed or painted-flat mouth still works: Riggle carves a cavity from your outline instead. The result is coarser, and you will see exactly how coarse in the preview video before you spend anything. Heavily stylized and furry meshes are the same story — preview first, and re-mark if your lip trace was too generous.

Do you actually need all 52?

It's worth asking honestly, because the answer is usually no. The 52 ARKit shapes exist to mirror a human face in fine detail — every brow segment, every cheek puff, every individual eyelid state — driven live from an iPhone's depth camera. If you are performing subtle facial acting to an audience that will notice a raised inner eyebrow, that detail is the product.

For most avatars it isn't. What reads as 'alive' on stream is a mouth that matches the audio, eyes that blink on their own, and a face that can register an emotion — the five visemes, blink and independent winks, and happy/angry/surprised. That's the VRM expression set, it's what VRM apps drive from audio and idle timers, and it's what Riggle builds.

So the honest boundary: Riggle does not generate the 52 ARKit shapes — that's roadmap, not product. What it does is take a model whose face doesn't move at all and give it one that does, in about a minute, with a video preview before you spend anything. If that's the gap you're staring at, it's the whole fix.

Give your model a face that moves

Upload a rigged GLB, mark the lips and eyes in about twenty seconds, and watch every expression play back before you download a validated VRM 1.0. Your first model is free.

FAQ

Can I add visemes and blinks without Blender?

Yes. Riggle runs in your browser: upload, mark the face, watch the preview, download a VRM 1.0. Blender does the actual work on our servers, but you never open it.

Does Riggle do perfect sync / all 52 ARKit blendshapes?

No. Riggle produces the VRM 1.0 expression set — five visemes, blink plus independent left and right winks, happy, angry, surprised, and eye-bone look-at. Full 52-shape ARKit perfect sync is a larger job and it's on our roadmap, not in the product today. If your model's face doesn't move at all today, that expression set is the difference between a mannequin and an avatar.

Do I need a model that's already rigged?

Yes. Riggle maps an existing humanoid skeleton onto the VRM bones; it does not create skeletons. Rig the body in Tripo, Meshy or Mixamo first, then bring the .glb here — or bring an existing .vrm to re-rig.

How long does it take?

About twenty seconds of your attention, plus roughly a minute of processing. The part that disappears is the slow one: hand-sculpting shape keys, or waiting a day on a studio service.

Related