Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Apple research quantifies what LLMs actually do when they mimic humans — and finds we can control it.

Why it matters
Researchers lack empirical methods to evaluate and govern human-like behaviors in LLMs (emotions, relationship-building, boundary-setting). Apple's multi-dimensional analysis of 21,000 LLM outputs offers practitioners and labs actionable frameworks for deciding when anthropomorphism helps or hurts.
The key facts
11 to know21,000 LLM outputs analyzed
Multi-dimensional evaluation: prevalence, effects, controllability
Methods: LLM-as-judge and human evaluation combined
Behaviors examined: emotional expression, relationship-building, refusal/boundary-setting
Focus: system prompt controllability of anthropomorphic outputs
Source: Apple Machine Learning Research
21,000+ evaluations across human-like behaviors
Multi-dimensional analysis: prevalence, effects, controllability
Methods: LLM-as-a-judge + human evaluation
Focus areas: thought/emotion expression, relationship-building, refusals, boundary maintenance
Source: Apple Machine Learning Research (peer-reviewed research, not product launch)
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite their prevalence, researchers and practitioners lack methods and empirical…