AI tools increasingly promise to decode messages, score romantic interest and tell you whether a date “really likes you.” A peer-reviewed study published in Scientific Reports on May 11, 2026 gives those claims a serious test. It also gives them a serious limit.
The researchers found that ChatGPT and Claude 3 could detect some verbal indicators of attraction from complete speed-date transcripts. Their predictions were associated with later ratings and matching decisions. The associations were modest, and some of the cues used by both AI and human observers were not valid indicators of the actual outcome.
What the 2026 study tested
The research reused recordings from eight speed-dating events held at a large Midwestern university in 2007. The original sample included 187 undergraduates, 93 women and 94 men, with a mean age of 19.6. Each four-minute date involved a woman and a man.
Professional transcribers produced the conversations. The research team matched 964 transcripts to the participants’ later reports and decisions. Each transcript averaged 864 words.
| Part of the study | What the researchers used |
|---|---|
| Language model | GPT-4-0613 rated the complete written transcript three times. Claude 3 was tested in supplementary analyses. |
| Human comparison | Eighteen judges watched the videos. Four other judges read a random subset of 500 transcripts. |
| Subjective outcomes | Participants rated liking, likelihood of a follow-up date, commonality, personality similarity and connection. |
| Behavioural outcome | Both participants privately chose whether to exchange contact information. Twenty-three per cent of the dates produced a mutual match. |
This matters because the model was not handed a list of isolated signs. It saw the full back-and-forth of a short conversation. It was then compared with a private, later decision rather than a reviewer’s opinion about whether the date looked successful.
The result, in proportion
| Prediction | Association with the measured outcome | Fair reading |
|---|---|---|
| ChatGPT reading transcripts and predicting mutual contact exchange | r = 0.12 | A small positive association. |
| Human judges reading transcripts and predicting mutual contact exchange | r = 0.13 | Similar to ChatGPT when given the same information. |
| Human judges watching videos and predicting mutual contact exchange | r = 0.31 | Better than either transcript-only judge, but still imperfect. |
| ChatGPT predicting participants’ subjective ratings | r = 0.13 to 0.23 | Modest associations across five reported experiences and intentions. |
In mixed-effects models that accounted for participants going on several dates, ChatGPT’s prediction remained associated with the outcomes. The fixed effect explained a modest share of variance, with marginal R2 values of approximately 0.01 to 0.04.
“As good as humans” does not mean accurate
ChatGPT was on par with a small group of people who read the same transcripts for one outcome. Both performed weakly. The comparison does not show that ChatGPT can reliably classify a new date, produce an individual probability or replace the answer of either person involved.
AI and people shared some bad guesses
The researchers examined which conversational features influenced ChatGPT’s predictions and whether those features predicted actual matching. The overlap was imperfect. ChatGPT and human observers often relied on similar cues, but some cues were weakly related, unrelated or pointed in the opposite direction from the later decision.
Positive tone illustrates the distinction. It correlated r = 0.28 with ChatGPT’s prediction, but only r = 0.05 with actual matching. A conversation sounding positive helped the model form a confident impression. It was far less informative about whether both people exchanged contact information.
That is the central risk in an AI “interest detector.” A polished explanation can reveal which cues the model associates with attraction without showing that those cues are valid for the person being discussed.
What newer speech research adds
A July 25, 2026 preprint examined a different dataset: Japanese speed-dating conversations involving 147 adults aged 19 to 60. The researchers combined a transcript-only language model with a supervised model using speech features.
The combined system improved pairwise ranking accuracy over the transcript-only model in all four tested settings. Gains in participant-level Pearson correlations varied, however, and none remained statistically significant after correction. The authors’ conclusion was conditional: speech can add information in some settings, not that audio creates a universal attraction detector.
The preprint is useful because it changes the question. Conversation text may not contain every relevant feature. Tone and speech patterns can add signal. They still do not turn prediction into access to another person’s private state.
What the study does not establish
- It does not show that an AI tool can diagnose attraction from a few screenshots or selected messages.
- It does not establish accuracy across ages, cultures, sexual orientations, relationship stages or ordinary dates.
- It does not show that the studied model versions behave like every current or future system.
- It does not convert mutual contact exchange into a complete measure of attraction.
- It does not measure consent. Attraction, a second date and consent to a particular act are different questions.
The participants consented to research recording and analysis. That does not create permission to upload another person’s private messages or recordings to an AI service for a secret assessment. Privacy remains part of the decision.
A better use of AI when you are uncertain
AI can be more useful as a check on your own reasoning than as a machine that assigns someone else a hidden feeling. Ask it to separate observation from inference, remove pressure from an invitation or identify where you are treating ambiguity as certainty.
| Useful question | Question to avoid |
|---|---|
| “Separate the facts I observed from the explanations I invented.” | “Give me the percentage chance she secretly likes me.” |
| “Help me make this invitation specific and easy to decline.” | “Tell me which signs prove I should escalate.” |
| “What response would respect a no, silence or continued uncertainty?” | “Explain how to overcome her hesitation.” |
The bottom line
The 2026 study is evidence against two extreme claims. AI was not reading random noise. It detected a small amount of real signal in complete speed-date conversations. It also did not come close to reliably knowing what one person wanted.
Use the result the way the size permits. A model’s interpretation can be one hypothesis. The other person’s voluntary response to a clear, low-pressure invitation is better information.
Research cited
- Matz, S. C., Peters, H., Cerf, M., Grunenberg, E., Eastwick, P. W., Back, M., & Finkel, E. J. (2026). “Large language models can detect verbal indicators of romantic attraction.” Scientific Reports, 16, 21441. Published May 11, 2026.
- Kikuchi, Y., Hayashi, T., Kimura, R., Inoue, N., Ishii, R., & Okada, S. (2026). “Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating.” arXiv:2607.23037. Preprint posted July 25, 2026; conference record doi:10.1145/3776574.3831151.
Sources and reported values were checked on August 10, 2026. Correlation describes association across observations. It is not an individual accuracy percentage.
From prediction to a real answer
Read the first three chapters free
You Can’t Read Her examines missed flirting, failed romantic prediction and the difference between a private guess and a voluntary response.