Can a test predict a breakup? Why there is no honest percentage
August 11, 2026 · 14 min read
No compatibility test can calculate the probability of a breakup, and LoveScore does not try. A number only earns the word “probability” after someone has followed a large group of couples who took that exact test, recorded over years what actually happened to them, and then checked the arithmetic: among the couples told “68%”, roughly 68 in a hundred should have separated. Questionnaires do not do that work. What a test can do is describe the picture two people painted in their answers today — LoveScore reads 47 answers from each partner into nine relationship dimensions, an overall Love Score, your strengths as a couple and your friction points, and not one of those is a forecast. The wording is blocked on our side as well: “odds of divorce”, “you will break up” and “doomed” are patterns in the stop-list every generated report is checked against, and a text carrying one of them never reaches a reader.
If you searched for that percentage, though, the number was probably not the point. Behind “breakup probability calculator” there is usually a simpler and heavier question: is this bad? The question is legitimate, and we are not going to talk you out of it. But an honest answer to it does not look like a percentage. It looks like a short list of specific places worth talking about. Below: what those calculators are actually doing, what a scale you can measure looks like instead, and what the report puts in place of a number.
What a breakup probability calculator is actually doing
Search results for this are full of tools that will hand you a figure in fifteen seconds. The machinery behind most of them is simpler than it looks. You answer a handful of questions, the answers are turned into a score the way any quiz score is built, and then the score is relabeled. “You scored low on trust” becomes “68% chance of breaking up”. Nothing was validated between those two sentences. The only thing that changed is the word in front of the number, and the word is what makes it worth clicking.
The difference between the two sentences is not politeness of phrasing. The first describes answers, and you can check it: ask the same questions next month and see whether it moves. The second is a claim about the future that nobody stands behind. And a percentage about the future has a property a score does not have — it cannot be wrong in any way you would ever notice. If you stay together, well, it said 68%, not 100%. If you separate, it was right all along. A statement that survives both outcomes is not a measurement of anything.
The swap has a price, and the service showing you the number is not the one who pays it. A percentage reads as a sentence handed down. It does not invite a conversation; it invites you to either accept it or argue with a figure, and neither of those is something two people can do together. That is why there are no predictions in LoveScore — not through good manners, but by construction. The ban lives in an automatic check on the generated text rather than in an author's intentions.
What a scale you can actually measure looks like
A measurable scale differs from a forecast in three ways, and all three can be checked from the outside, before you answer a single question.
- It names what is being measured. Not “your compatibility” in general, but a specific construct. Trust, for instance, is defined as belief in a partner's honesty, willingness to be unguarded around them, and the ability to lean on their word without checking behind it. From a definition like that it is clear which questions belong on the scale and which do not.
- The result repeats. The same answers give the same score, to the point. In LoveScore a deterministic algorithm does the arithmetic; the language model comes in afterwards and writes prose over numbers that are already fixed. It cannot move them, and it cannot invent one — that is enforced by a separate check.
- You can see what the number depends on. Which questions feed which dimension, which dimensions weigh more than others in the overall score, how two sets of answers become one couple score. All of it is written out step by step in the methodology, rather than summarized as “our unique algorithm”.
None of the three turns a scale into a prediction — and that is not a shortcoming of the method, it is where the method ends. A scale says: this is how you described your life today. What happens next depends on what the two of you do, not on what a screen showed you. If you want the longer checklist for telling an honest instrument from a random number generator, we wrote one: how to choose a compatibility test.
What the report shows instead of a percentage
Love Score is an index of alignment, not a forecast
The overall score is assembled from eight of the nine dimensions with fixed weights — the same weights for every couple. Trust and communication carry the most: trust is a precondition for everything else, and communication is the only dimension that describes not a property of a person but a process, the one a couple uses to work through disagreement. Emotional maturity, empathy, responsibility and peace of mind follow. The lightest are the two match scales — the balance of togetherness and space, and the picture of the future — where an aligned couple reaches the top without needing much weight to get there. The ninth dimension, partner responsiveness, stays out of the overall score entirely; why, further down. The scale is calibrated so that a couple answering dead center on every question lands at exactly 50: not “half a relationship”, but a neutral point of reference.
One property of the formula matters especially in a conversation about worry. The score is non-compensatory at the top: if any one of the critical dimensions has dropped — trust, communication, peace of mind, or, for couples past the first months, the picture of the future — the overall score cannot rise higher than the weakest of them allows. Strong marks elsewhere do not buy it back. That exists precisely so an average cannot paint a comfortable picture over one serious gap. But a ceiling is not a prediction either. It says there is something in the answers that the rest does not compensate for, not that the relationship is over. What a particular figure does and does not mean is unpacked in how to read your Love Score.
Friction points: the closest thing to a “red flags test”
The metric nearest to what people mean by risks is captioned, on screen, Friction points — and the caption is the whole claim it makes. It is built from nine components — a shortfall of trust, conversations that never reach a decision, both partners reacting fast, control and separation anxiety, a picture of the future that does not line up, a mismatch on personal space, contribution spread unevenly, two different pictures of the same conversations, one partner carrying the empathy. A tenth, a blind spot, was designed and is switched off: it rests on a measure we have not calibrated yet, and we are not going to score people on an uncalibrated scale.
The value is deliberately not shown as a percentage — the % sign is forbidden for this metric in the configuration itself. What the screen carries is three things and no more: the caption, a value written as so many out of 100, and a band label, which is what leads. The bands run from “no pronounced risks visible” through “systemic bottlenecks” to “a critical configuration”. The top band does not mean it is over. It means enough divergences have stacked up that working through them alone, one at a time, is not a realistic plan. Read the whole block that way and it stays useful: a friction point is a topic worth discussing before it discusses you, not a rating of your relationship and not a verdict on it. The difference from a percentage is substantial. Not “you will break up over money”, but “your answers about who carries the household diverge more than anything else, and that is where a conversation would start”.
Strengths are counted separately, not as “100 minus the bad”
The other half of the picture is your strengths as a couple: ten load-bearing supports, eight of which only count when the quality is present in both partners. The remaining two — a matched balance of togetherness and space, and a matched picture of the future — work differently. There the support is credited for the answers converging, and only if the topic is live for at least one of the two: agreeing about something neither of you has ever thought about is not a resource. That logic is worked through in how similar should couples be. Crucially, this metric is computed independently of friction, not as what is left over from it. Which is why the ordinary and very common configuration is a couple with plenty of supports and noticeable friction at the same time. That is not a contradiction in the report. It is how living relationships are built: people who hold each other firmly in one place can grind for years in another.
Scenarios are about readiness, not about whether it happens
A separate block covers six life steps: moving in together, getting married, having children, surviving a renovation, starting a business together, everyday life together. A scenario score does not show the likelihood of the event. It shows how closely two sets of answers converge around that step today. This is the easiest place to mistake a description for a forecast, so, plainly: a low score on getting married does not mean you will not get married. It means your answers about a shared future currently diverge, and that is better known before the conversation than halfway through it. How the underlying dimension works is on the Future Vision page — there is no “right” answer there, and a couple that agrees on wanting neither marriage nor a shared household scores just as high as a couple planning a family.
Why there is no “red flags” button
“Relationship red flags test” is searched about as often as anything involving breakup odds, and our answer to it is an awkward one: there is no such button, and the phrase itself is on the stop-list — “red flag” and “green flag” alike. The reason is not squeamishness. It is what a label does to an observation. “Trust is low here” describes answers, and something can be done with it. “Red flag” is a verdict glued to a person, and two people can barely discuss it at all: one is left defending, the other accusing.
The ban on clinical language is stricter still. Narcissist, sociopath, psychopath, codependency, toxic relationship, the words disorder, diagnosis and pathology — the model writing the report is not allowed to say any of it, and the validator catches every one of those words. Attachment-style labels are blocked for the same reason: “you have an anxious attachment style” sounds like knowledge and works like a diagnosis. The test does not know your partner. It knows 45 substantive answers they gave on one particular day, plus two attention checks. You do not get a diagnosis out of material like that — not about a person and not about a couple. Where jealousy is concerned the report says only what it can actually see, which is how much control and separation anxiety showed up in the answers; the difference between that and a label is the subject of healthy vs. unhealthy jealousy.
“Who loves who more” is a question with no scale
The test does not measure love, and it certainly does not rank two people against each other. The nearest thing it has is Partner Responsiveness — the one scale where you answer not about yourself but about how much attention and interest you feel coming from the other person. That is exactly why it is deliberately kept out of the overall Love Score: otherwise one partner's answers would enter the couple's score twice, once as a self-assessment and once as a rating of the other.
The dimension works by comparison rather than by ranking. How a person rates their own engagement — through empathy, through conversation, through holding steady in conflict — is set beside how the person next to them sees it. A match confirms the self-perception. A gap points either to a blind spot, where you consider yourself more attentive than it feels from the other side, or to underrating yourself, where your partner values your engagement more highly than you do. Neither of those means “loves me less”. Both mean the two of you talk about each other's contribution less often than would be useful.
The same logic runs through the archetypes: there are twelve, and none of them is better or worse than another. An archetype describes which dimension stands out most in a person, not their place in a league table of the couple.
What the test refuses to do
We state the limits of the method out loud, because a test that pretends to have none is lying about the thing that matters most.
- It does not predict an outcome. No score has been validated against what later happened to any couple. Love Score is an index of how your answers align with the model, not a forecast.
- It does not compare you to other couples. “Higher than most couples”, “above average” and the rest are on the stop-list too. No population base for such comparisons exists, and they sound convincing anyway — which is the dangerous combination.
- It does not rate your partner or diagnose anyone. The report speaks about the couple, and about each person through their own answers. It does not hand down a judgment on a human being.
- It does not invent numbers. Every number in the generated text has to be present in the core calculation, and that is checked automatically. If the generated text fails the check, a deterministic template built from the same numbers goes out in its place.
- It does not build a report out of unreliable answers. The questionnaire carries attention checks and watches the pace and the sameness of responses. Where there is moderate doubt, exact scores give way to ranges and the text becomes more careful; where the answers are clearly unreliable, no report is produced at all. An honest report cannot be made from random input.
If you came here worried
The worry that sends someone looking for a breakup percentage almost always points at something specific: a conversation being postponed, a promise nobody went back to, a subject you both walk around. A percentage would not have lifted it even if such a number existed. A high one would have soothed you briefly, a low one would have frightened you for a long time, and neither would have told you where to start.
What a test can give you instead is an address. Both partners answer 47 questions separately, from their own devices, in about seven minutes each, and the report is calculated once both are in. The free part shows the Love Score, your strengths, your friction points, both archetypes and three scenarios; the full report with the written analysis is a one-time $2.99. Its real value is not in the scores, though. It is that afterwards a couple has an opening line — not “we need to talk”, but “look, our answers about free time diverge more than anything else. How do you see it?”. Money, plans and trust usually go unspoken not because people do not want to talk about them, but because it is unclear which edge to pick them up by.
And one thing worth saying plainly, since it is why some people are reading this at all. If being around your partner is sometimes frightening, if you find yourself managing your words to avoid setting something off, if control has become the texture of ordinary days — that is not a question a compatibility test answers. The report assumes two people who are able to talk to each other calmly. When that is not the case, what helps is not a questionnaire result but a conversation with someone outside it; free and anonymous support lines exist in most countries, and “I think I might be exaggerating” is a perfectly good reason to call one. A low score is not a diagnosis and a high one is not a guarantee: relationships move with what two people do next, not with what a screen showed them. If the subject is heavy and the conversation between you keeps failing, seeing a professional is a normal step — the same kind of practical step as seeing a doctor about a pain that will not go away.
Frequently asked questions
Can a test predict a breakup?
No. For a number to be a probability, someone would have to follow a large group of couples who took that exact test for years and check the predictions against what actually happened to them. Questionnaires do not do that work — they describe the picture your answers paint today. In LoveScore, phrasings like "odds of divorce" and "you will break up" are on a stop-list that every generated report is checked against.
Are breakup probability calculators accurate?
They are not measuring what the label says. The percentage is usually built from the same answers that would give an ordinary quiz score, then relabeled as a prediction, because a prediction is more compelling. A figure like that cannot be wrong in any way you would notice: stay together and it only said 68%, separate and it was right.
Does LoveScore show relationship red flags?
Not as a label: "red flag" and "green flag" are on the stop-list, along with clinical words like narcissist, codependency or the name of any diagnosis. What the report has instead is a metric captioned Friction points, showing where the two sets of answers diverge most. It is built from nine components, from a shortfall of trust to contribution spread unevenly.
What does a low Love Score mean — that we are going to break up?
No. Love Score is an index of how closely your answers align with the model, not a forecast: no score has ever been validated against what later happened to a couple. A low result means there is something in the answers that the rest does not compensate for. That is a list of topics for a conversation, not a verdict on the relationship.
What should I do if the result frightened me?
Remember that you are looking at a description of today's answers, not at the future. Take one friction point — whichever lands hardest — and talk it through together, starting from yourself rather than from a grievance. If the conversation will not hold, or if the relationship does not feel safe, the sensible next step is not retaking the test but talking to a professional.
Keep reading
How to choose a compatibility test: a 7-point checklist
Are compatibility tests accurate? Seven signs of an honest couples test, from behavior-based questions to a published method — and what filters out the fakes.
July 27, 2026 · 9 min read
How does AI calculate compatibility? Where the numbers come from
In LoveScore a deterministic algorithm turns 47 answers into a score, and the language model only writes the text on top. Here is where each part works.
August 11, 2026 · 17 min read
AI compatibility by birth date: is AI astrology accurate?
What an AI birth date compatibility calculator actually computes, why adding a language model adds no data, and how to tell a measurement from a label.
August 11, 2026 · 18 min read
Ready to find out your Love Score?
47 questions, about 7 minutes, each partner on their own device. The basic report is free.
Take the test for freeLoveScore is an entertainment and educational service for adults 18+. Articles and reports are not a psychological consultation or a diagnosis, and they do not replace working with a specialist.