A portrait photograph contains no motion data. Producing a video in which that face appears to speak requires the software to invent every frame after the first one, which is why results vary so widely between one source image and another.
For people comparing these tools, the useful questions are practical. What happens between upload and export? Which parts of the photograph affect quality? How much does the script change lip sync? The answers also explain why most poor results fall into three familiar failure modes.
What happens between the photograph and the video
The process runs in two stages that are often described as one.
The first stage is facial landmark detection. The software locates a set of reference points on the face: the corners of the mouth, the edges of the lips, the jawline, the eye contours, and the bridge of the nose. These points define a deformable mesh. Everything the finished video does to the face is a transformation of that mesh.
The second stage is audio driven animation. Text is converted to speech, the resulting audio is analysed for phonemes, and each phoneme is mapped to a mouth shape. The mesh is then deformed frame by frame to match the sequence.
This explains a behaviour that confuses first time users. The output quality depends far more on the landmark detection than on the audio. If the software cannot place the reference points confidently, no amount of adjusting the voice or the script will fix the result, because the errors are geometric rather than acoustic.
It also explains why visible articulation matters. Sounds produced with clearly visible lip movement, such as rounded vowels and bilabial consonants, render more convincingly than sounds formed further back in the mouth. The University of Iowa’s Sounds of Speech project documents the articulation of individual sounds with animated diagrams, and the distinctions it illustrates are the same ones these tools have to approximate.
The photograph does most of the work
Image requirements across these tools are consistent enough to treat as a single checklist, and they are stricter than most product pages suggest.
The face must be roughly frontal. Landmark detection degrades sharply as the head turns, because points on the far side of the face become occluded and the software has to estimate their position. A three quarter portrait that looks flattering as a still often produces a mouth that appears to slide during speech.
The image needs genuine resolution rather than an upscaled file. A phone photograph of a printed picture is a common failure case: the file dimensions look adequate while the actual detail around the mouth is insufficient for a clean mesh.
Group photographs are rejected or produce unpredictable results, since the software must decide which face to animate. Hats and sunglasses interfere with landmark points on the brow and eye contours. Heavy filtering removes the texture gradients that detection relies on, which is why a heavily retouched portrait can perform worse than a plain one.
Leadde’s published requirements follow the same pattern, specifying a high definition image that reflects the subject’s current appearance and excluding group shots, hats, sunglasses, pets in frame, heavy filters, low resolution files and screenshots. Tools that animate photos to speak with AI generally converge on this list because the underlying constraint is the same.
Script wording changes the output more than voice selection does
Most guidance on these tools concentrates on voice choice. Sentence construction has the larger effect.
Short sentences with clear boundaries produce cleaner lip sync than long subordinated ones. A phoneme sequence with natural pauses gives the animation discrete resting positions, and the mouth returns to a neutral shape between them. Continuous speech without punctuation produces a mouth that never settles.
Proper nouns are the other common problem. Place names, product names and abbreviations are frequently mispronounced, and a mispronunciation is also a wrong mouth shape. Most platforms allow a pronunciation to be corrected once and applied across an entire script, which is worth doing before generating rather than after.
Where a script is being written for an audience that will read as well as listen, captions are worth enabling from the first version. Captions for prerecorded video are a WCAG success criterion, and retrofitting them across a set of published videos is more work than turning them on at the start. Leadde provides nine caption styles with adjustable colour, font, weight, size and alignment, applied across a whole video.
Three failures account for most poor results
Poor output usually traces back to one of three causes, and none of them are fixed by switching tools.
The first is an unsuitable source image, described above. This is the most common single cause and the easiest to correct, because it requires a different photograph rather than different software.
The second is expression that does not match the content. Several platforms offer two animation modes. A standard mode handles mouth movement and blinking. A more expressive mode infers emotion and body language from the narration and generates matching facial movement, though it typically carries a duration limit. Leadde caps its Expressive IV Engine at sixty seconds per video. Applying expressive animation to procedural or factual content produces a presenter who appears to react to information that carries no emotional content, which reads as strange even when viewers cannot articulate why.
The third is length. Attention to a static frame with an animated face declines quickly. A talking photograph works for short segments and degrades over several minutes, regardless of technical quality. Content that runs long generally needs additional visual material rather than a longer animation of the same face.
Where the technology still falls short
Side profiles remain unreliable, and no current tool handles them well.
Speech disfluency is absent. Generated narration contains no false starts, no hesitation and no self correction, which makes it recognisable as synthetic to attentive listeners even when the animation is convincing.
Regional accent control is coarse. A dialect selector in these tools generally means a national standard variety rather than a specific regional accent. Leadde lists 88 languages and 175 dialects, which covers most practical requirements without delivering the precision that the word dialect implies.
Free tiers apply watermarks in most cases, including Leadde’s. Free access is sufficient for evaluating whether the output is usable, and removing the watermark requires a paid plan on nearly every platform in this category.
Finally, these tools generate video from supplied content and do not record screens. Any material whose value lies in showing an interface requires a screen recording instead.
Test with the photograph you will actually use
A useful trial starts with an ordinary photograph, a script containing numbers and proper nouns, and the duration planned for the finished clip. Those inputs test the image mapping and narration under realistic conditions. A flattering studio portrait reading a warm greeting is an easy case for every tool in this category, so it reveals very little about how the software will handle day-to-day work.
