Settings tool · Rules updated 7 October 2026

SadTalker Settings Finder

Choose your input types and the result you want. Get a starting configuration for framing, head motion, base size and face restoration, with a reason for each choice.

Runs in your browser. No files are uploaded and no video is generated. Recommendations follow fixed rules; they are not predictions of quality, render time or GPU memory.

Choose descriptions of your inputs. This tool does not inspect your image or audio.

Try an example

Each example fills the three choices and shows its recommendation. Portrait-specific advice remains untested where noted.

How the choices are made

  1. Framing: a wider portrait or a retain-scene goal selects full. Otherwise, an already framed face selects resize; other image types select crop. Retaining the scene takes precedence over already-framed input.
  2. Head motion: steady presenter and retain-scene goals select still ON. The other goals select OFF. Still keeps the head steadier while the mouth continues animating.
  3. Base size: quick preview selects 256. Other goals select 512. This is the face renderer size, not necessarily the final video dimensions.
  4. Restoration: enhancer stays none unless you choose try face restoration. That goal selects gfpgan for photographic inputs; illustrations and difficult faces keep none. A soft photo alone does not force enhancement.
  5. Audio: the selected audio type changes the advice, not these four settings. Noise, speech rate and language require checking the recording and output rather than assuming a rendering toggle will fix them.
  6. Evidence: an exact match against the eight captured settings combinations is labeled separately from an untested combination. Even a match is evidence from one studio source, not a test of your media.

Check a short render first

Our controlled comparisons used one studio presenter portrait, one narration and an RTX 4090 on 4 October 2026. The finder adds no new GPU measurements. Other faces, noisy or fast recordings, illustrations and non-English speech have not been scored in those comparisons.

Inspect mouth timing, likeness, eyes, teeth and jaw/neck boundaries. GFPGAN may generate altered detail. Full composites the face into the source scene rather than animating the body. No setting guarantees correct face detection, identity or lip sync.

If memory runs out, try a smaller base size, turn enhancement OFF and reduce batch size to 1. That fallback is guidance, not a benchmarked minimum GPU requirement. Use the memory troubleshooting guide.

Inspect the benchmark proof

The Best Settings guide brings the results together and lists the 24-combination test matrix. Eight combinations have evidence; sixteen remain unrun.