What behavioral data reveals about AI value
MIT found 95% of enterprises deployed AI with no measurable return. Ninety's support team discovered why: they were measuring activity, not outcomes.

Why it matters
Enterprise AI deployments fail because organizations measure adoption (logins, prompts, licenses) rather than work change (friction, rework, process elimination). Behavioral data—what customers and employees actually do—connects AI output to business outcomes. This matters for CIOs trying to justify AI spending and for practitioners building agents that need to prove ROI.
The key facts
12 to knowMIT study: $35B–$40B AI commitments in U.S. enterprise; 95% reported no measurable return
Ninety case: AI support agents given behavioral session context (prior customer attempts, failed steps) improved ticket resolution from 73% to 79%
Distinction: provisioning signals (seats, licenses) and engagement signals (logins, prompts) show availability/activity; displacement signals (human steps removed, handoffs eliminated) and friction signals (rework, escalation, abandonment) reveal economic value
Measurement challenge: prompt volume and time-in-app are ambiguous proxies—fewer interactions can indicate better resolution, not lower adoption
Behavioral context fed to agents before conversation begins, not after; escalated tickets inherit same context rather than requiring manual reconstruction
Ninety expanded behavioral measurement beyond product use to include external AI model queries (ChatGPT, Gemini) on help content, identifying weak source material before customers arrive at support
MIT research: $35–40B committed to enterprise AI; 95% report no measurable return
Ninety case: ticket resolution improved from 73% to 79% (+6 points) when AI agents received behavioral session summaries before conversation
Distinction: provisioning signals (seats, licenses) and engagement signals (logins, prompts) measure availability and activity; displacement (handoffs eliminated, manual work gone) and friction (rework, escalation, abandonment) measure actual value
At Ninety, behavioral context reduced agent handoffs: agents no longer open separate session-replay tools or manually reconstruct customer context while customer waits
Ninety extended measurement beyond product: tested how public LLMs (ChatGPT, Gemini) answer product questions using company help-center content, scored answer quality to improve source material
Governance risk: monitoring individuals by prompt count can corrupt behavior; better unit of analysis is the workflow, not the worker
Go to the source
CIOcio.com
Publisher excerpt: Taylor Paletta was glancing at two versions of the same customer-service interaction on her monitor. As director of support and digital success at the enterprise software company, Ninety, she could read the support chat in one window while watching the customer’s product session in another. The…