Talking avatars

Bring a single portrait to life

Combine an avatar, a script, and a voice. Set up the avatar once and change the script for each video. Choose an AI avatar to stay off camera, or clone your own voice.

Created with xiats · AI-generated avatars, no live filming

See the results first

Every video below was generated with xiats. All avatars are AI-generated, with no live filming.

AI-generated avatarAI avatar presentationAn AI-generated avatar with automatic voiceover and lip sync, ready as a finished video
Everyday settingsStreet-style presentationAnimate and lip-sync a single portrait, with no live filming
Business settingsNighttime business presentationChoose the avatar, background, and style to match your content, regardless of existing footage
Single-image animationWarm-lit home presentationAn animated, lip-synced AI portrait with no live actor or real-person portrait licensing

Three steps to a video

Combine an avatar, a script, and a voice. Set up your avatar once; change only the script for each new video.

1

Choose an avatar

Upload a video of yourself facing the camera, or generate an AI character if you prefer to stay off camera.

Real footage: width and height 640–2048 pixels, fixed front-facing camera, about 30 seconds
2

Add a script

Extract speech from platform share links, upload a video for transcription, or ask AI to write a script.

Edit the script anytime and regenerate a video after even a one-line change
3

Choose a voice and generate

Pick a library voice or clone your own. The system adds voiceover and lip sync to produce the finished video.

Price shown before generation; full automatic refund on failure
  • Only your own assetsUpload an authorized avatar or generate one with AI. The Platform does not collect other people's portraits for you.
  • Temporary output storageGenerated files are temporary and retained for about 24 hours. Download them promptly.
  • Credentials managed on the serverPlatform credentials are maintained in the backend and persisted in the database, without hardcoding or exposure.
  • Two key modesPlatform mode uses official keys out of the box. BYOK uses your key and bills your cloud account.

Why use xiats for talking avatars?

Create more than lip sync: manage your avatar, voice, script, and cost in one place.

No need to appear on camera

Four presets are available: an articulate female presenter, a professional male narrator, an energetic creator, and a 3D virtual character. All are AI-generated without real-person portrait licensing.

Clone your voice

Use a library voice or clone your own. Pair the same avatar with different voices and choose your presentation style.

No need to write from scratch

Extract speech from Xiaohongshu, Douyin, Bilibili, or Kuaishou links, or transcribe an uploaded video. AI can rewrite it into a more natural spoken script.

Know the cost before generation

Voiceover is billed by character count and lip sync by output duration. A confirmation dialog shows “Voiceover ¥X + Lip sync ¥Y = Approximately ¥Z” before generation, with no hidden charges.

Full automatic refunds on failure

Failed jobs are not charged. Reserved funds return to your wallet automatically, without contacting support.

One avatar, many videos

Record your avatar footage once. Change the script for each new video and generate in minutes, without setting up lighting and cameras again.

What people create

When the same person needs to deliver different messages repeatedly, talking avatars save time.

Product presentations

Create a presentation for each product using the same avatar, and prepare a batch for launch day.

Daily educational content

Turn topics into scripts and let an avatar present them, so daily publishing does not depend on filming.

Company communications & training

Keep product introductions, onboarding, and internal announcements consistent. Update the script to update the video.

Test different versions

Try variations in script and voice to see which performs best, paying only for the generated versions.

How avatar generation is priced

  • Voiceover: Billed by the script's character count.
  • Lip sync: Billed by actual output duration, approximately equal to the voiceover length, with a minimum charge.
  • Before generation, a dialog shows “Voiceover ¥X + Lip sync ¥Y = Approximately ¥Z” and your current balance. You are charged only after confirmation.
  • Failed generation is fully refunded to your wallet automatically, without contacting support.

See the workspace price list for current rates. Sign in and open Talking avatars for a live quote. Generation is charged by usage from your wallet, independently of membership. View subscriptions。

Frequently asked questions

Create your first talking video

Upload avatar footage or choose an AI avatar, add a script, and select a voice.