TY - GEN
T1 - Prompt-to-Animation
T2 - 18th Annual ACM SIGGRAPH Conference on Motion, Interaction, and Games, MIG 2025
AU - Durupinar, Funda
AU - Normoyle, Aline
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/12/2
Y1 - 2025/12/2
N2 - Expressive facial animation depends on models that can convey subtle, context'dependent emotions. Procedural methods often rely on manually-tuned heuristics, while data'driven techniques are constrained by the diversity of their training data. This paper explores the zero-shot potential of large language models (LLMs) to generate facial animations. Using the cognitively-grounded Ortony-Clore-Collins (OCC) model as a framework, we designed 110 text-based scenarios and evaluated the ability of different LLMs - Gemini-2.5 Pro, GPT-4o, and Llama 3.1-8b - to generate corresponding facial animations using the Facial Action Coding System (FACS). A perceptual user study confirmed that animations from Gemini-2.5 Pro were highly recognizable, with participants successfully matching facial expressions to the correct context at rates significantly above chance. A quantitative analysis of the generated Action Units (AUs) indicated both consistency within each OCC emotion category and diversity across different scenarios. Further analysis revealed that the emotional expressions group into six clusters: happiness, sadness, anger/disgust, fear/surprise, shame, and neutral. This work demonstrates a viable, lightweight pipeline connecting textual narrative directly to motion generation without requiring custom training or large-scale motion capture datasets.
AB - Expressive facial animation depends on models that can convey subtle, context'dependent emotions. Procedural methods often rely on manually-tuned heuristics, while data'driven techniques are constrained by the diversity of their training data. This paper explores the zero-shot potential of large language models (LLMs) to generate facial animations. Using the cognitively-grounded Ortony-Clore-Collins (OCC) model as a framework, we designed 110 text-based scenarios and evaluated the ability of different LLMs - Gemini-2.5 Pro, GPT-4o, and Llama 3.1-8b - to generate corresponding facial animations using the Facial Action Coding System (FACS). A perceptual user study confirmed that animations from Gemini-2.5 Pro were highly recognizable, with participants successfully matching facial expressions to the correct context at rates significantly above chance. A quantitative analysis of the generated Action Units (AUs) indicated both consistency within each OCC emotion category and diversity across different scenarios. Further analysis revealed that the emotional expressions group into six clusters: happiness, sadness, anger/disgust, fear/surprise, shame, and neutral. This work demonstrates a viable, lightweight pipeline connecting textual narrative directly to motion generation without requiring custom training or large-scale motion capture datasets.
KW - Emotions
KW - Facial Action Coding System
KW - Facial animation
KW - Large Language Models
KW - OCC Model
UR - https://www.scopus.com/pages/publications/105024975571
UR - https://www.scopus.com/pages/publications/105024975571#tab=citedBy
U2 - 10.1145/3769047.3769055
DO - 10.1145/3769047.3769055
M3 - Conference contribution
AN - SCOPUS:105024975571
T3 - Proceedings MIG 2025 - 18th ACM Conference on Motion, Interaction and Games
BT - Proceedings MIG 2025 - 18th ACM Conference on Motion, Interaction and Games
A2 - Sumner, Robert W.
A2 - Zund, Fabio
A2 - Jorg, Sophie
A2 - Pelechano, Nuria
PB - Association for Computing Machinery, Inc
Y2 - 3 December 2025 through 5 December 2025
ER -