← All posts / Research

The AI Observatory: Independent Data Reveals How People Actually Use AI — and Half of It Isn't Work

A new independent research project aggregated 85,633 real AI conversation turns from 5,000 users and 52 models — and found that Anthropic's own methodology would filter out 48% of them, hiding the health, companionship, and sensitive conversations that define real AI use.

The AI Observatory: Independent Data Reveals How People Actually Use AI — and Half of It Isn't Work

How do people actually use AI chatbots? It sounds like a question the industry should be able to answer easily. After all, OpenAI and Anthropic process billions of conversations every month and both regularly publish glossy reports on usage patterns. But according to independent researchers, those reports only show the data the companies want us to see — and a major new study published this week suggests the real picture looks substantially different from the official one.

An independent telescope for AI behavior

The project is called the AI Observatory, and it’s the most serious attempt yet to build an independent evidence base for how people interact with generative AI. Co-led by Anka Reuel, a computer science PhD candidate at Stanford’s Trustworthy AI Research (STAIR) Lab, and Shayne Longpre, a recent PhD graduate from the MIT Media Lab, the project aggregated and analyzed real AI conversations collected with users’ consent across seven existing research datasets.

The scale is meaningful, if modest next to what the labs themselves hold: 85,633 conversational turns across 24,521 conversations, from 5,000 users interacting with 52 different models — including ChatGPT, Claude, Gemini, and Grok — between 2023 and 2025. WildChat, one of the largest public conversation datasets, is among the sources.

The motivation, Reuel told MIT Technology Review, is that “there is no independent source to corroborate” the industry’s usage claims. Highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited, self-reported data.

The 48% that gets filtered out

The study’s headline finding lands squarely on the best-known industry dataset of all: the Anthropic Economic Index, which filters out Claude conversations unrelated to work and productivity. When the AI Observatory researchers applied Anthropic’s own methodology to their independent dataset, nearly half the conversations — 48% — would have been discarded.

What’s in that excluded half matters. The non-work conversations showed significantly higher rates of:

  • Health and relationship topics: 44.2%, versus 31.2% in Anthropic’s work-focused analysis
  • Harassment and hate: 27.5% versus 5.66%
  • Sexual content: 16.7% versus 2.4%
  • Adult or illicit topics: 7.9% versus 2.1%

The pattern isn’t unique to Anthropic. OpenAI’s own 2025 report on ChatGPT usage acknowledged that only about 30% of consumer use was work-related. But because each company publishes through its own lens — Anthropic emphasizing productivity, OpenAI emphasizing education — the composite picture skews heavily professional, and the deeply personal ways people actually use these tools stay sectioned off in footnotes.

David Widder, an assistant professor at the University of Texas at Austin who studies human-AI interaction and was not involved in the project, told MIT Technology Review that having a “bird’s-eye-view analysis” rather than scattered company reports helps researchers understand usage more consistently.

People are getting personal with their AI

Beyond the headline numbers, the longitudinal data tells a story of AI use becoming steadily more intimate. In WildChat conversations, prompts, responses, and conversation turns all grew longer and more elaborate over time, and small talk increased significantly — a signature of rising AI companionship use. Perhaps more troubling, the assistants’ self-disclosure decreased: models became less likely to admit they were chatbots as conversations got more personal.

There is one genuinely encouraging signal: exchanges labeled as sensitive — potentially harmful or restricted content, including sexual harassment and hate speech — became less frequent over the study window, which may indicate that platform safeguards have genuinely improved.

Every model has its own culture

One of the study’s most interesting contributions is a cross-model comparison that no single company could ever publish. Usage patterns differed sharply by platform:

  • Grok and Gemini dominated information retrieval, with Grok especially popular for news and politics — and it was also where misinformation tended to concentrate, consistent with prior research. xAI did not respond to a request for comment.
  • Claude became the go-to for coding.
  • Gemini led for social and roleplay uses.
  • ChatGPT carried homework assistance.

Even model versions behaved differently: users had shorter conversations with ChatGPT running GPT-3.5 and longer, more iterative ones with GPT-4o. “No single company report tells the whole story,” Longpre said.

The caveats — and why this matters anyway

The researchers are candid about limitations. Because the data comes from voluntarily shared sources, it probably underrepresents sensitive uses — the conversations people are least willing to hand to researchers. The 7.9% figure for adult or illicit topics is likely a floor, not a ceiling. And 85,633 turns is a drop in the bucket compared to the 1 million Claude conversations underpinning the latest Anthropic Economic Index, or whatever internal trove OpenAI holds.

Anthropic said its published research reflects its teams’ specific questions and that supporting external independent research is important. OpenAI did not respond to requests for comment.

But the asymmetry is precisely the point. When policymakers debate youth safety rules, when regulators assess harms, or when researchers ask whether a system is “used mostly for good or mostly for bad,” the only granular data available sits behind corporate walls. The AI Observatory’s dataset will be made available to other researchers, and the team hopes to expand it over time. Ideally, Reuel says, the companies themselves would share data with independent researchers — with privacy protections.

Until then, anyone making decisions based on AI usage data is, in her words, “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives.”

For an industry that regularly claims its products are becoming general-purpose infrastructure, the finding that roughly half of real usage is personal — health questions, relationships, loneliness, companionship — is not a footnote. It’s the story.