How Reliable Is ChatGPT'sRelationship Advice?
Researchers tested ChatGPT against human judgment on 13,138 real relationship problems. Its advice rankings barely matched, and in the hard cases it never gave the same answer twice.
People ask ChatGPT about their relationships constantly, so Hou, Leach and Huang (2024) tested how good its answers are. They collected 13,138 Reddit posts about intimate relationship problems, had ChatGPT rank the advice given in response, and compared its rankings with human judgments. Then they asked the same questions again to see whether it would agree with itself.
Key takeaways
- ChatGPT's rankings of relationship advice barely tracked human judgment: Kendall's tau-b between 0.069 and −0.188, effectively no agreement.
- In "low disparity" cases, where the options were close and the judgment subjective, ChatGPT never produced the same ranking twice.
- It is least reliable exactly where you most want a second opinion: the messy, close-call situations.
- Used as an editor, brainstorm partner or devil's advocate, it is still useful. Used as a judge, it is a coin with a confident voice.
What the study measured
Alignment with people
The researchers compared ChatGPT's ranking of advice options with human rankings across several testing scenarios. Agreement was very weak throughout, with correlation values (Kendall's tau-b) from 0.069 to −0.188. Knowing what ChatGPT preferred told you almost nothing about what people preferred.
Agreement with itself
They then gave ChatGPT identical scenarios repeatedly. In clear-cut cases it was fairly stable. In the low-disparity cases, where advice options were close in quality and the call was subjective, it never returned identical rankings across attempts, and severe disagreements between runs were common.
"In response to people's growing interest in using ChatGPT as a relationship advisor, our research evaluates ChatGPT's proficiency in discerning relationship advice. Specifically, we investigate its alignment with human judgements."— Hou, Leach & Huang, ICWSM 2024
Why this happens
Two reasons, neither surprising. Language models produce a plausible answer, not a stable one; ask twice and the sampling changes. And relationship problems as people describe them are one-sided by construction: the model only ever hears the narrator's version, so its judgment is a judgment of the framing. It cannot see whether the "cold" partner has been fielding twelve messages a day.
The second problem is the fixable one, and it is why analysis of the actual conversation is a different kind of tool from advice about a description of it.
How to use ChatGPT for relationship advice anyway
- Give it the exchange, not your summary. Paste the messages, with names removed, rather than "he's being distant." Your summary is the verdict you already reached.
- Ask for options and reasons, not a ruling. "Give me three readings of this and what would make each one likely" gets the model doing what it is good at.
- Ask twice. If the answers differ, you have learned that the situation is a close call, which is itself useful.
- Use it to draft, not to decide. Wording a hard message is a language task; deciding whether to send it is not.
- Keep it out of anything involving safety. Coercion, threats or fear are not advice questions, and a model's confident tone is a liability there. Our guide to toxic text patterns covers what to look for and where to get help.
This is also the difference between a general chatbot and a chat analyzer. MosaicChats works from the whole conversation rather than one person's account of it, and its assistant Myrah answers questions with that record in front of it, so "why did things get colder in March" is answered from the messages, not from your memory of them.
Ask about the conversation, not a summary of it
Upload the chat and MosaicChats shows how its tone moved week by week and who carried it. Myrah answers your questions from that record.
Analyze your chatFrequently asked questions
Is ChatGPT good at relationship advice?
Not reliably. In a 2024 study of 13,138 Reddit relationship posts, ChatGPT's rankings of advice barely correlated with human judgments (Kendall's tau-b from 0.069 to -0.188), and in close-call cases it never gave the same ranking twice.
Why does ChatGPT give different answers to the same relationship question?
Language models sample a plausible answer each time rather than computing a fixed one, and relationship dilemmas rarely have a single right answer for the model to converge on. The study found inconsistency was worst in low-disparity cases, where the advice options were close in quality.
What is the safest way to use AI for relationship advice?
Give it the actual messages rather than your summary, ask for several readings with reasons, ask twice and compare, and use it to draft wording rather than to make decisions. Keep anything involving safety or coercion with a person.
Is a chat analyzer more reliable than ChatGPT for this?
It answers a different question. ChatGPT judges your description of a situation; a chat analyzer such as MosaicChats measures the conversation itself, such as reply times, who initiates and how tone changes over weeks. Measurements are reproducible; a verdict on a one-sided story is not.
Related reading
References & Sources
- Hou, H., Leach, K., & Huang, Y. (2024). ChatGPT giving relationship advice – How reliable is it? Proceedings of the International AAAI Conference on Web and Social Media, 18(1), 610–623. Source