ComfyUI guide · Tested 4 October 2026 · RTX 4090

SadTalker GFPGAN: enhancer ON vs OFF

Compare face detail before and after GFPGAN restoration using the same source portrait, narration and RTX 4090. The main pair uses size 512 on both sides.

GFPGAN is a generative face-restoration step. Sharper-looking facial detail is not proof that it matches the original person or improves lip sync.

Same input, restoration OFF and ON

Only enhancer changes in the main pair: none versus gfpgan. Both use size 512, crop, still OFF, batch size 2, expression scale 1.0 and pose style 0. The synchronized player shows GFPGAN OFF on the left, ON on the right and uses one audio track.

Both outputs are scaled to equal 512-pixel display panels using bicubic interpolation. Each is one fresh render, not a selection of the best attempts.
Get the comparison MP4

Size 512 · GFPGAN OFF

512 × 512 · 25 fps · 11.24 s
Get original output

Size 512 · GFPGAN ON

1024 × 1024 · 25 fps · 11.24 s
Get original output
Narration transcript

This demonstration showcases how SadTalker transforms a single image into a realistic speaking avatar with synchronized lip movements and expressive facial animation.

Shared test narration; English captions are available in every player.

Inspect eyes, skin, mouth and face edges

Approximately the same normalized region is cropped from each output: x 24%, y 35%, width and height 52%. Both crops are enlarged to 512 × 512 with bicubic interpolation; this display adds no genuine image detail. Expression and head pose can vary because the template has no seed control.

In the sampled close-ups, GFPGAN ON gives eyes and facial hair more defined edges and produces smoother-looking skin than OFF. The mouth and teeth also look different, but motion and expression differ between these independent renders. The sharper presentation is visible here; exact identity recovery and improved temporal stability were not measured.

Look across the whole clip, especially around teeth, eye shape, facial hair and the jaw boundary. These are visual observations from one input, not scored identity, lip-sync or temporal-consistency measurements.

Measured render time and sampled VRAM

We rendered four combinations once each to cover both questions with minimal GPU time. There were no extra warm-ups or repeated trials. The node starts a new worker for every render; model-loading time is included. Values are individual measurements, not medians or a universal speed ranking.

Swipe the table sideways to view time, resolution and VRAM.

Four individual size and enhancer measurements on the same GPU
Size / enhancerRender timeFinal videoGPU baselineSampled GPU peakPeak above baseline
256 · OFF27.88 s256 × 256396 MiB4799 MiB4403 MiB
512 · OFF55.41 s512 × 512396 MiB10047 MiB9651 MiB
256 · ON77.22 s512 × 512396 MiB4895 MiB4499 MiB
512 · ON119.39 s1024 × 1024396 MiB10207 MiB9811 MiB

At base size 512, OFF took 55.41 s and ON took 119.39 s: ON added 63.99 s, or 2.15× the total elapsed time in this pair. The sampled whole-GPU peak rose from 10,047 MiB to 10,207 MiB. GFPGAN doubled final dimensions to 1024 × 1024; it did not change the base render setting.

What the VRAM number means: we sampled NVIDIA’s whole-device memory.used about every 0.5 seconds. It includes the worker, ComfyUI, driver and any other GPU allocations. “Above baseline” is the observed device peak minus the measurement immediately before that run. Sampling can miss short peaks; these values are not PyTorch allocator peaks or guaranteed minimum card requirements.

Server execution_start → execution_success timing includes input loading, worker/model startup, rendering and export; queue wait and provisioning are excluded. Fresh filenames invalidate input caching, and successful outputs were checked for video/audio streams and hash integrity.

Exact setup, sampled data and evidence
GPU
NVIDIA GeForce RTX 4090, 24564 MiB, 580.95.05
Worker
Python 3.10.21 · PyTorch 2.1.2+cu121 · CUDA 12.1
CPU thread limits
OMP, MKL, OpenBLAS and NumExpr limited to 4 each in the isolated worker.
Inputs
1117 × 1034 source portrait; 11.26-second narration; identical bytes for every run.
Run order
256 OFF → 512 OFF → 256 ON → 512 ON
SadTalker revision
cd4c0465ae0b54a6f85af57f5c65fec9fe23e7f8
Custom node revision
4c231660fea4e1f2748d03dcf38ea438fd9f884f
ComfyUI revision
73c9bad4d21e7addbe1d13bc92eee0f1431b017d
Seed
Not exposed by the template; expressions and blinking can vary.

Get times, VRAM samples and hashesGet original outputs, workflows and logs

Use the four-thread worker launcher instructions to match these CPU limits. The workflow alone does not set worker CPU threads.

Reproduce the comparison in ComfyUI

Start with our ComfyUI setup guide. This workflow extends its LoadImage / LoadAudio template with two SadTalker nodes sharing the same inputs.

  1. Get the comparison workflow and drag it into ComfyUI.
  2. Upload the portrait and the narration. Both nodes share them.
  3. Keep size = 512 on both. Set the left enhancer to none and the right to gfpgan. Keep crop, still OFF (displayed as animated), batch 2, expression 1.0 and pose 0.
  4. Click Run. Generated MP4s and logs are saved under ComfyUI/output/sadtalker/<run-id>/. Repeating an unchanged workflow can return cached results.
Real ComfyUI settings for SadTalker GFPGAN: enhancer ON vs OFF, with shared portrait and audio inputs
The comparison template loaded on the rented machine. Open to inspect settings; this screenshot is a settings view, not a timed output.
ComfyUI API commands

Get the two API graphs, place the supplied portrait/audio filenames in ComfyUI’s input folder, then queue them:

512-off API graph512-on API graph

Queue the matched pair
bash
for file in 512-off-api.json 512-on-api.json; do
  python -c 'import json,sys; print(json.dumps({"prompt":json.load(open(sys.argv[1]))}))' "$file" > prompt.json
  curl -X POST http://127.0.0.1:8188/prompt -H 'Content-Type: application/json' --data-binary @prompt.json
done

Use your actual server address. The rented machine runs internally on port 18188; 8188 is the usual local example. Fresh input filenames were used in the benchmark to avoid cached timings.

When does GFPGAN help, and when can it hurt?

Try ON when the rendered face looks soft

It may make eyes, eyebrows and facial hair look clearer. Compare the restored face against the original portrait and watch the whole clip, particularly the mouth and teeth. The tested integration uses GFPGAN v1.4 with 2× final upscaling.

Prefer OFF when restoration changes the face

Keep the unenhanced output if ON introduces waxy skin, altered features, unnatural teeth or distracting frame-to-frame changes. These are review criteria, not claims that every artifact occurred in this sample. Restoration can generate plausible detail rather than recover exact missing information.

Weigh extra processing against the visible benefit

Use the OFF/ON pair at your intended playback size. If the difference is small there, the extra time and larger output may not be useful. GFPGAN is a face-restoration stage, not a control for head pose or lip-sync timing.

Limits and failed runs

4 of 4 configurations completed successfully. No render errors occurred in these four runs. This is one portrait, one narration and one rented GPU. There are no repeat-trial error bars, scored lip-sync or identity metrics, or out-of-memory boundary tests.

GFPGAN OFF and ON are separate SadTalker runs. The template has no seed control, so part of a frame-level difference may come from expression, blinking or head motion. This is an end-to-end ComfyUI toggle comparison; it does not isolate restoration by processing the exact same generated frame twice.

If a render fails or memory runs out, inspect its worker log and reduce batch size to 1. That setting was not tested here. Changing batch size, preprocessing, still mode, hardware or worker CPU threads makes the timings less directly comparable.

Implementation references: the tested enhancement/upscaling code, GFPGAN’s official restoration project, and the tested ComfyUI node.