← All posts / Research

400,000 Reddit Posts, One Algorithm: AI Surfaces GLP-1 Side Effects Clinical Trials Missed

Penn researchers used LLMs to mine 400,000 Reddit posts from nearly 70,000 GLP-1 users, surfacing menstrual changes, chills, hot flashes and fatigue that rarely appear on drug labels — a template for AI-driven pharmacovigilance.

400,000 Reddit Posts, One Algorithm: AI Surfaces GLP-1 Side Effects Clinical Trials Missed

When a drug goes from niche prescription to mainstream phenomenon almost overnight, the formal safety apparatus has a timing problem. Clinical trials — the gold standard — are slow by design, powered by thousands of participants, not the tens of millions now taking the medication. A new University of Pennsylvania study, published in Nature Health and surfacing in a fresh wave of coverage this week, shows what the opposite approach looks like: an AI-driven sweep of more than 400,000 Reddit posts from nearly 70,000 users of GLP-1 drugs, analyzed at a scale no human team could ever attempt.

The result is a genuinely new picture of patient experience with semaglutide (Ozempic, Wegovy, Rybelsus) and tirzepatide (Mounjaro, Zepbound) — one where the familiar gastrointestinal effects sit alongside signals that rarely make it onto drug labels or into regulators’ adverse-event dashboards: menstrual irregularities, chills, hot flashes, and unexplained fatigue.

What the study actually did

The Penn Engineering team — first author Neil Sehgal, a doctoral student in Computer and Information Science, with senior author Sharath Chandra Guntuku, co-author Lyle Ungar, and Penn Center for Weight and Eating Disorders investigator Jena Shaw Tronieri — applied what Guntuku calls “computational social listening” to five years of Reddit conversations about GLP-1 medications.

The hard part was never collecting posts; it was translating them. Patients do not describe symptoms in standardized medical terminology. One person writes that they feel unusually cold, another mentions constant chills, a third describes “freezing all the time” — and a clinician would file all three under the same MedDRA (Medical Dictionary for Regulatory Activities) category. Historically, mapping informal social media language onto standardized medical codes required so much manual annotation that large-scale analysis was impractical.

Large language models such as GPT and Gemini changed that equation, Sehgal notes, making it possible to process and categorize vast amounts of text with a level of standardization that was previously unattainable. The team used LLMs to classify the informal symptom descriptions into standardized categories across the full corpus.

Validation first: the known signals showed up

A critical strength of the study is that the method found what it was supposed to find before anyone trusted it with anything novel. Roughly 44% of users in the sample described at least one side effect, and gastrointestinal problems topped the list — exactly matching the nausea and digestive issues already well documented for semaglutide and tirzepatide.

“Some of the side effects we found, like nausea, are well known, and that shows that the method is picking up a real signal,” Guntuku said. In other words, the pipeline passed its sanity check against established pharmacology before the interesting findings got any attention.

The underreported signals

Beyond the known effects, three categories stood out as frequent enough to warrant closer investigation yet underrepresented in clinical trial reporting:

  • Reproductive symptoms. Nearly 4% of users who reported side effects described menstrual changes — bleeding between periods, heavy bleeding, and irregular cycles. That figure would be even higher in a female-only sample, Sehgal points out. The researchers frame it as “a signal worth investigating,” not a established causal link.
  • Body-temperature disruption. Chills, feeling unusually cold, hot flashes, and fever-like symptoms appeared repeatedly in the corpus.
  • Fatigue. Tiredness was the second most frequently reported complaint in the Reddit data — despite relatively few clinical trials reporting fatigue often enough to hit established reporting thresholds.

The hypothalamus may explain why the menstrual and temperature findings are biologically plausible rather than statistical noise. The hypothalamus, a small region of the brain, helps regulate hunger, hormones, reproduction, and body temperature — and GLP-1 drugs are thought to work partly by engaging it. “That doesn’t mean the medications are necessarily causing these symptoms, but it could suggest that reports of menstrual changes and body temperature fluctuations are worth studying more systematically,” Tronieri said.

Association, not causation — and the researchers say so

The study’s authors are unusually careful about what their data can support. “We can’t say that GLP-1s are actually causing these symptoms,” Sehgal said. The findings are associations in what people discussed online, not proof of drug effects. The sample skews younger, more male, and more American than the overall GLP-1-using population. And self-reported posts carry their own distortions.

That candor is part of what makes the work useful. The team positions it not as a replacement for trials but as a fast, hypothesis-generating complement: “Clinical trials are the gold standard, but by design, they are slow,” Guntuku said. “This is not a replacement for trials, but it can move much faster, and that speed matters when a drug goes from niche to mainstream almost overnight.”

Why now: LLMs unlock what annotation couldn’t

Ungar has been here before — he participated in one of the earliest efforts to mine user-generated internet content for adverse drug effects back in 2011. The concept is old; the economics were impossible. Before LLMs, mapping hundreds of thousands of informal posts to MedDRA codes would have required enormous human annotation effort, capping corpora at a fraction of this size.

“Online patient communities work a lot like a neighborhood grapevine,” Ungar said. “People who are living with these medications are swapping notes with each other in real time, sharing experiences that rarely make it into a doctor’s office visit or an official report.” LLM classification finally makes the grapevine searchable at population scale.

The bigger play: an early-warning system for fast-moving products

The forward-looking implication extends beyond GLP-1s. Loosely regulated or unregulated products — injectable peptides, supplements, wellness formulations — can spread through Reddit, TikTok, and other platforms faster than conventional research can register their existence. AI analysis of those conversations could become an early detection system for emerging health concerns, providing some of the earliest indications of unexpected effects.

The Penn team’s next steps reflect that ambition: broadening beyond Reddit and beyond English-language communities to test whether the same patterns hold across different populations and platforms. Disclosures note the study had no outside funding; Tronieri reports an investigator-initiated grant from Novo Nordisk and consulting fees from Currax Pharmaceuticals, while the other authors report no conflicts.

The takeaway

This study is less a story about Ozempic than a story about method. When a blockbuster drug reaches millions of people in months, the collective conversation those patients have online becomes the largest de facto post-market surveillance dataset in existence — invisible until something can read all of it. LLMs can. The signals they surface are leads, not verdicts, but leads are precisely what clinical research runs on. Pharmacovigilance has historically been limited to what patients report to doctors, manufacturers, and regulators; the grapevine just became admissible evidence.