Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
Yi Fei Cheng,
Fan Yang,
Iremsu Bas,
Koichiro Niinuma,
Narishige Abe,
David Lindlbauer.
Published at
EMNLP
2026
Abstract
As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we evaluate six simulation approaches on a 43-hour multi-camera office dataset, investigating how trace-derived conditioning mechanisms, including persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors, influence behavioral fidelity. We find that fidelity varies across metrics and temporal scales: statistical priors improve alignment with real activity distributions and local transitions, but also produce over-segmented routines and reduced within-person variability. These findings underscore the need for holistic evaluation across temporal granularities and at both the individual and population levels.
Bibtex
@inproceedings {Cheng2026EMNLP,
author = {Cheng, Yi Fei and Fan, Yang and Bas, Iremsu and Niinuma, Koichiro and Abe, Narishige and Lindlbauer, David},
title = {Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations},
year = {2026},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
publisher = {Association for Computational Linguistics},
keywords = {Computational behavior modeling; user simulation; multimodal perception},
location = {Budapest, Hungary},
series = {EMNLP '26}
}