iFLYTEK has launched a digital human creation platform that can generate a realistic digital human from a single photograph. Previously, making a comparable digital human usually required about two hours of video recording. The shift sharply lowers the barrier to producing digital humans.

Starting from one photo

The system integrates iFLYTEK's Spark voice large language model for natural dialogue and draws on vertical-industry knowledge bases to reduce hallucination, while supporting customizable personas with long-term memory. The digital humans can produce natural facial expressions, gestures and lip-sync, perform more complex actions such as picking up objects, drinking, singing and dancing, and respond in real time.

On commercial deployment, an official in charge of the relevant iFLYTEK business said the application scenarios span brand livestreaming, customer service and unmanned retail. The market is expanding fast: in 2025, China's core digital human market reached 48.06 billion yuan, about 7.12 billion US dollars, and is projected to grow to 93.56 billion yuan by 2030.

A hot market, a cold rollout

Competition is intense. Tencent's game-streaming arm unveiled a real-time multimodal digital human model supporting around-the-clock livestreaming, while Baidu launched a digital human video podcast solution said to cut production costs by 74 percent. Even so, the industry shows a gap between a hot market and a cold rollout, with many digital humans still stuck at script-following displays unable to sustain long real-time interaction.

Pan Helin, a member of the expert committee under the Ministry of Industry and Information Technology, said that as large models are adopted, digital humans are moving from a premium tool for a few enterprises toward a more standard business asset. With the technical barrier falling, the real competition may lie in who can make digital humans genuinely usable and durable.