Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations

Yi Fei Cheng, Fan Yang, Iremsu Bas, Koichiro Niinuma, Narishige Abe, David Lindlbauer.
Published at EMNLP 2026
Teaser image

Abstract

As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we evaluate six simulation approaches on a 43-hour multi-camera office dataset, investigating how trace-derived conditioning mechanisms, including persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors, influence behavioral fidelity. We find that fidelity varies across metrics and temporal scales: statistical priors improve alignment with real activity distributions and local transitions, but also produce over-segmented routines and reduced within-person variability. These findings underscore the need for holistic evaluation across temporal granularities and at both the individual and population levels.

Bibtex

@inproceedings {Cheng2026EMNLP, 
 author = {Cheng, Yi Fei and Fan, Yang and Bas, Iremsu and Niinuma, Koichiro and Abe, Narishige and Lindlbauer, David}, 
 title = {Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations}, 
 year = {2026}, 
 booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing}, 
 publisher = {Association for Computational Linguistics}, 
 keywords = {Computational behavior modeling; user simulation; multimodal perception}, 
 location = {Budapest, Hungary}, 
 series = {EMNLP '26} 
 }