When a friend leans across a loud izakaya table to tell you something, where do your eyes go?
Most people would say "to the face." But a face is a big place, and the part you look at may depend on the language you grew up with.
For anyone who makes talking videos for viewers in Japan and abroad, this small habit matters. A tour guide who recorded a welcome message in Japanese may want an English version too. With an ai lip sync video, her face can deliver that new line, her lips moving in time with a recorded English voice track. Whether viewers even look at those lips is a question researchers in Japan have put to the test.
The Mouth Quietly Helps the Ear
We tend to think of listening as a job for the ears alone. In practice, the eyes pitch in. On a busy train platform or in a crowded bar, a glance at someone's lips can help you catch words the noise would otherwise swallow.
The best-known example of this teamwork is a lab trick called the McGurk effect. A recording of someone saying "ba" is paired with video of a mouth saying "ga." Many listeners then report hearing a third sound, such as "da." The ears and eyes disagree, and the brain settles the argument by blending them.
This is where Japan comes in. Earlier studies found that the illusion is much weaker for Japanese listeners than for English listeners. Something about how Japanese speakers combine sight and sound seemed to be different, but the reason was not clear.
What the Kumamoto Team Measured
To look closer, researchers at Kumamoto University and their colleagues ran four experiments with young adult native speakers of English and Japanese. Volunteers watched short clips of two women, one Japanese and one English speaker, saying the syllables "ba" and "ga." They pressed a button to say which one they heard. The team timed their answers, recorded their brain waves and tracked where their eyes went.
The results pointed in opposite directions. For the English speakers, seeing the matching mouth movement shortened their response time. For the Japanese speakers, the researchers observed the opposite effect. Adding the moving mouth made them slower, not faster, and the brain-wave readings followed the same pattern.
The eye tracking, done with 13 English speakers and 16 Japanese speakers, helped explain why. The English group's gaze was drawn to the mouth, especially in the moment before any sound came out. The Japanese group showed no such pull. The authors suggest the English speakers used those early lip movements to get ready for the sound. The Japanese speakers did not.
It helps to keep the scale in mind. These were small groups of young adults naming single syllables in a lab, not friends chatting over dinner. The study does not show that Japanese people ignore lips in daily life. It shows a clear difference in how two groups used the mouth during one simple task.
A Habit Learned Early, Not a Switch
One of the most interesting results came last. The team asked some Japanese volunteers to focus on the speaker's mouth on purpose. They did look at it longer. But looking longer did not make the mouth change what they heard.
So the difference is not only about where the eyes happen to land. The authors point to language as one possible reason. Earlier research they cite sorted English consonants into five or six groups that lipreaders can tell apart. Japanese consonants fall into only three. If the mouth carries less useful information in a language, children who grow up with it may learn to lean on the voice.
That fits another earlier finding the paper describes. At age six, Japanese and English speaking children were swayed by the mouth to a similar, small degree. After that, the pull grew with age for English speakers but stayed the same for Japanese speakers.
Culture may also play a part. The paper mentions studies of facial expressions in which Eastern viewers paid more attention to the eyes and Western viewers to the mouth. The authors are careful to say their experiments were not designed to prove a cause.
Masks, Subtitles and Anime Voices
Seen through this research, a few everyday things in Japan look a little different. Face masks are a common sight during cold and hay fever seasons. They cover the very part of the face many English speakers watch. It is tempting to connect the two, but the study did not test masks, so any link is only a guess.
Movies offer another angle. Foreign films in Japanese cinemas are often released in two versions, one subtitled and one dubbed, and viewers choose the one they like. Anime has its own tradition, too. Voice actors commonly record their lines while watching the animation. They fit their timing to mouths that often use just a few simple shapes.
Audiences everywhere enjoy both, which is a useful reminder. People adapt to many kinds of mouth-and-voice matching, and what feels natural is partly learned.
Making Talking Videos for Both Audiences
If you make talking clips for viewers in Japan and in English-speaking countries, the research points to a few simple habits.
Put the audio first. If many Japanese viewers lean on the voice, a muddy recording will hurt more than a less-than-perfect picture. Record in a quiet room, speak clearly and keep music out from under the speech.
Keep the mouth timing natural anyway. English speakers in the study looked at the mouth early, before the first word. They may be quick to notice lips that run out of step. A front-facing shot with soft, even light makes the mouth easy to read for those who do watch it.
Add captions. Subtitles help both groups, along with anyone watching on a silent train. They also give viewers a second chance to catch a word that the audio or the mouth did not make clear.
Respect the person on screen. Only upload material you have the right and permission to use. If it includes another person's image or voice, make sure you have the consent required for your intended use.
Conclusion
Where our eyes go when someone speaks may be partly learned. In this study, English speakers were already watching the mouth before a sound came out, while Japanese speakers did not show that early focus on the lips. A good bilingual video leaves room for both.















