ComfyUI guide · Tested 4 October 2026 · RTX 4090
SadTalker GFPGAN: enhancer ON vs OFF
Compare face detail before and after GFPGAN restoration using the same source portrait, narration and RTX 4090. The main pair uses size 512 on both sides.
Same input, restoration OFF and ON
Only enhancer changes in the main pair: none versus gfpgan. Both use size 512, crop, still OFF, batch size 2, expression scale 1.0 and pose style 0. The synchronized player shows GFPGAN OFF on the left, ON on the right and uses one audio track.
Size 512 · GFPGAN OFF
Size 512 · GFPGAN ON
Narration transcript
This demonstration showcases how SadTalker transforms a single image into a realistic speaking avatar with synchronized lip movements and expressive facial animation.
Shared test narration; English captions are available in every player.
Inspect eyes, skin, mouth and face edges
In the sampled close-ups, GFPGAN ON gives eyes and facial hair more defined edges and produces smoother-looking skin than OFF. The mouth and teeth also look different, but motion and expression differ between these independent renders. The sharper presentation is visible here; exact identity recovery and improved temporal stability were not measured.
Look across the whole clip, especially around teeth, eye shape, facial hair and the jaw boundary. These are visual observations from one input, not scored identity, lip-sync or temporal-consistency measurements.
Measured render time and sampled VRAM
We rendered four combinations once each to cover both questions with minimal GPU time. There were no extra warm-ups or repeated trials. The node starts a new worker for every render; model-loading time is included. Values are individual measurements, not medians or a universal speed ranking.
Swipe the table sideways to view time, resolution and VRAM.
| Size / enhancer | Render time | Final video | GPU baseline | Sampled GPU peak | Peak above baseline |
|---|---|---|---|---|---|
| 256 · OFF | 27.88 s | 256 × 256 | 396 MiB | 4799 MiB | 4403 MiB |
| 512 · OFF | 55.41 s | 512 × 512 | 396 MiB | 10047 MiB | 9651 MiB |
| 256 · ON | 77.22 s | 512 × 512 | 396 MiB | 4895 MiB | 4499 MiB |
| 512 · ON | 119.39 s | 1024 × 1024 | 396 MiB | 10207 MiB | 9811 MiB |
At base size 512, OFF took 55.41 s and ON took 119.39 s: ON added 63.99 s, or 2.15× the total elapsed time in this pair. The sampled whole-GPU peak rose from 10,047 MiB to 10,207 MiB. GFPGAN doubled final dimensions to 1024 × 1024; it did not change the base render setting.
memory.used about every 0.5 seconds. It includes the worker, ComfyUI, driver and any other GPU allocations. “Above baseline” is the observed device peak minus the measurement immediately before that run. Sampling can miss short peaks; these values are not PyTorch allocator peaks or guaranteed minimum card requirements.Server execution_start → execution_success timing includes input loading, worker/model startup, rendering and export; queue wait and provisioning are excluded. Fresh filenames invalidate input caching, and successful outputs were checked for video/audio streams and hash integrity.
Exact setup, sampled data and evidence
- GPU
- NVIDIA GeForce RTX 4090, 24564 MiB, 580.95.05
- Worker
- Python 3.10.21 · PyTorch 2.1.2+cu121 · CUDA 12.1
- CPU thread limits
- OMP, MKL, OpenBLAS and NumExpr limited to 4 each in the isolated worker.
- Inputs
- 1117 × 1034 source portrait; 11.26-second narration; identical bytes for every run.
- Run order
- 256 OFF → 512 OFF → 256 ON → 512 ON
- SadTalker revision
- cd4c0465ae0b54a6f85af57f5c65fec9fe23e7f8
- Custom node revision
- 4c231660fea4e1f2748d03dcf38ea438fd9f884f
- ComfyUI revision
- 73c9bad4d21e7addbe1d13bc92eee0f1431b017d
- Seed
- Not exposed by the template; expressions and blinking can vary.
Get times, VRAM samples and hashesGet original outputs, workflows and logs
Use the four-thread worker launcher instructions to match these CPU limits. The workflow alone does not set worker CPU threads.
Reproduce the comparison in ComfyUI
Start with our ComfyUI setup guide. This workflow extends its LoadImage / LoadAudio template with two SadTalker nodes sharing the same inputs.
- Get the comparison workflow and drag it into ComfyUI.
- Upload the portrait and the narration. Both nodes share them.
- Keep
size = 512on both. Set the left enhancer tononeand the right togfpgan. Keep crop, still OFF (displayed as animated), batch 2, expression 1.0 and pose 0. - Click Run. Generated MP4s and logs are saved under
ComfyUI/output/sadtalker/<run-id>/. Repeating an unchanged workflow can return cached results.

ComfyUI API commands
Get the two API graphs, place the supplied portrait/audio filenames in ComfyUI’s input folder, then queue them:
512-off API graph512-on API graph
for file in 512-off-api.json 512-on-api.json; do
python -c 'import json,sys; print(json.dumps({"prompt":json.load(open(sys.argv[1]))}))' "$file" > prompt.json
curl -X POST http://127.0.0.1:8188/prompt -H 'Content-Type: application/json' --data-binary @prompt.json
doneUse your actual server address. The rented machine runs internally on port 18188; 8188 is the usual local example. Fresh input filenames were used in the benchmark to avoid cached timings.
When does GFPGAN help, and when can it hurt?
Try ON when the rendered face looks soft
It may make eyes, eyebrows and facial hair look clearer. Compare the restored face against the original portrait and watch the whole clip, particularly the mouth and teeth. The tested integration uses GFPGAN v1.4 with 2× final upscaling.
Prefer OFF when restoration changes the face
Keep the unenhanced output if ON introduces waxy skin, altered features, unnatural teeth or distracting frame-to-frame changes. These are review criteria, not claims that every artifact occurred in this sample. Restoration can generate plausible detail rather than recover exact missing information.
Weigh extra processing against the visible benefit
Use the OFF/ON pair at your intended playback size. If the difference is small there, the extra time and larger output may not be useful. GFPGAN is a face-restoration stage, not a control for head pose or lip-sync timing.
Limits and failed runs
4 of 4 configurations completed successfully. No render errors occurred in these four runs. This is one portrait, one narration and one rented GPU. There are no repeat-trial error bars, scored lip-sync or identity metrics, or out-of-memory boundary tests.
GFPGAN OFF and ON are separate SadTalker runs. The template has no seed control, so part of a frame-level difference may come from expression, blinking or head motion. This is an end-to-end ComfyUI toggle comparison; it does not isolate restoration by processing the exact same generated frame twice.
If a render fails or memory runs out, inspect its worker log and reduce batch size to 1. That setting was not tested here. Changing batch size, preprocessing, still mode, hardware or worker CPU threads makes the timings less directly comparable.
Implementation references: the tested enhancement/upscaling code, GFPGAN’s official restoration project, and the tested ComfyUI node.