LTX 2.3 Talking and Singing Lip Sync in ComfyUI | Portrait2Video Sync

详情

模型描述

Turn faces into singing, talking videos with perfect lip sync magic.

Who it's for: creators who want this pipeline in ComfyUI without assembling nodes from scratch. Not for: one-click results with zero tuning - you still choose inputs, prompts, and settings.

Open preloaded workflow on RunComfy

Open preloaded workflow on RunComfy (browser)

Why RunComfy first
- Fewer missing-node surprises - run the graph in a managed environment before you mirror it locally.
- Quick GPU tryout - useful if your local VRAM or install time is the bottleneck.
- Matches the published JSON - the zip follows the same runnable workflow you can open on RunComfy.

When downloading for local ComfyUI makes sense - you want full control over models on disk, batch scripting, or offline runs.

How to use (local ComfyUI)
1. Load inputs (images/video/audio) in the marked loader nodes.
2. Set prompts, resolution, and seeds; start with a short test run.
3. Export from the Save / Write nodes shown in the graph.

Expectations - First run may pull large weights; cloud runs may require a free RunComfy account.


Overview

This workflow lets you transform portrait images into talking or singing videos with natural mouth synchronization. It helps designers and creators bring static characters to life for storytelling, music videos, or interactive content. Built around advanced audio separation, it ensures clear vocals and precise lip movement. You can generate cinematic-quality results effortlessly, matching face expression and speech rhythm. Ideal for digital human animation, virtual singers, and performance testing.

Important nodes:

Key nodes in Comfyui LTX 2.3 Talking and Singing Lip Sync workflow

LTXVImgToVideoInplaceKJ (#1762)

Converts the reference portrait into an initial latent video. Adjust IMG STRENGTH. Set 0 for T2I mode (#1722) to balance identity lock versus freedom of motion. For tight likeness and stable head pose, keep the strength higher; for more performative motion, reduce it. If you set it to 0, treat the workflow like text-to-video and rely more on the prompt and audio.

Mel-Band RoFormer Sampler (#1599)

Separates vocals from a song to feed cleaner speech features into the model. Use USE VOCALS ONLY (#1616) when the instrumental is dominant or when consonants are getting lost; keep it off when you want the full mix to influence expressiveness. The final render uses the original trimmed audio, preserving your intended soundtrack. See the model details in the ComfyUI extension and weights linked above.

LTXV Audio VAE Encode (#1605)

Encodes the selected audio into the LTX audio latent space that drives mouth shapes and micro-expressions. Trim obvious silences before encoding for snappier alignment. Pair with the start-time offset if your lyric entry does not begin at 0 seconds.

LTX2 NAG (#508)

Applies noise-augmented guidance tailored for LTX 2.3 to sharpen audio-to-lip coupling. If lips feel under-animated, raise its guidance slightly; if teeth or mouth shapes look overdriven, lower the effect. Small, incremental changes work best because this control interacts with your prompt strength and sampler settings. Reference implementation lives in the LTX repositories.

SamplerCustomAdvanced (#161)

Runs the denoising process with your chosen KSamplerSelect (#154), CFG Guider (#153), and scheduler. Lower classifier-free guidance generally favors natural motion and identity retention; higher values increase prompt dominance but can stiffen lips. If you switch samplers, keep the rest of the settings intact for a fair A/B comparison.

Notes

LTX 2.3 Talking and Singing Lip Sync in ComfyUI | Portrait2Video Sync - see RunComfy page for the latest node requirements.

此模型生成的图像