WhatsWrapped

Who is more into whom: the four numbers behind the asymmetry score

The asymmetry score weighs day-openers at 30%, double-texts at 25%, reply speed at 25% and volume at 20%. Here is why that order, why ±0.15 is called dead even, and what the score ignores on purpose.

One question, four inputs, and a deliberate order

The asymmetry score is a single number between -1 and +1 that estimates who is doing more of the reaching in a two-person chat. It is built from four measures with fixed weights: who starts the day (30%), who double-texts (25%), who replies faster (25%), and who writes more (20%).

The weights are not a tie-break between four equally good signals. They rank the inputs by how much intent each one carries. A behaviour that requires someone to act with no prompt at all counts for more than a behaviour that only happens because the other person already spoke. Volume, the number most people expect to dominate, sits last for exactly that reason.

It is worth naming what the number is not measuring before going input by input. It measures initiative, not affection, and it has no opinion about whether the imbalance is a problem. A parent who always messages first is not more in love than their kid; they are just the one who opens the thread.

Why the first message of the day is worth 30%

Three of the four inputs are reactive. You cannot reply quickly to a message that never arrived, and you cannot out-write someone in a conversation that nobody started. The day's first message is the only input that is generated from a cold start, after hours of silence, with no social obligation pushing it out. That is why it carries the largest single weight.

It also has a useful property: it is naturally rate-limited. There is roughly one day-opener per day per chat, so a single chatty evening cannot inflate it. Word count has no such ceiling, which is why one dramatic 400-word voice-note transcript or one long rant can visibly move volume-based metrics and barely touch this one.

The honest limitation is that day-openers are sensitive to schedules rather than feelings. Whoever wakes first, whoever is in the earlier timezone, and whoever has a job that allows a phone at 8am will tend to win this input. If your chat is long-distance or your sleep patterns differ by two hours, treat a day-opener skew as partly structural. The related concept pages on who starts the conversation and on session boundaries explain how a new day's opener is separated from a continuation of last night's thread.

Double-texting at 25%, and why a run counts once

A double-text here is a run of two or more consecutive messages from one person with no reply in between. The whole run counts once, no matter whether it is two messages or eleven.

That decision matters more than it sounds. If every message in a run counted separately, the metric would stop measuring persistence and start measuring typing style. People who break one thought into six bubbles would look desperate; people who write one dense paragraph would look aloof. The behaviour the score is actually interested in is the moment someone decides to send again without having been answered, and that decision happens once per run.

Counting runs rather than messages also stops one bad night from swallowing the score. Sending fourteen messages into silence is a single data point, the same as sending two. Whether that pattern is worth worrying about is a question about the conversation, not the arithmetic.

Reply speed at 25%: the input nobody plans

Reply speed sits level with double-texting because it captures something people rarely curate: what they drop when a message lands. You can decide to play it cool about texting first. Answering in ninety seconds at 3pm on a Tuesday is closer to a reflex.

It is also the noisiest of the four. Sleep, shifts, meetings and flights all produce long gaps that have nothing to do with interest, which is why the recap treats fast in-conversation replies, overnight gaps and general reply waiting time as separate concepts rather than pretending one average explains a person. The asymmetry score compresses that texture into one comparison, so read it alongside the reply-time and burst-reply views rather than instead of them.

Note the interaction with the previous input: someone who replies slowly creates the silence that the other person then double-texts into. The two inputs often move together, and that is not double counting, it is the same dynamic seen from both ends.

Volume last, at 20%

Message count and word count are the loudest numbers in any chat recap, and the weakest evidence of who is chasing whom. High volume can mean pursuit. It can equally mean someone is a narrator, has a long commute, dictates voice notes, or simply talks in more words than their partner does in any medium.

It is still in the mix because the extreme case is real: a chat where one person writes almost everything is genuinely lopsided, and a score that ignored volume would miss it. Capping the weight at 20% is the compromise. A verbose person who never texts first, never double-texts and answers slowly will not be labelled the pursuer on word count alone.

If volume is the number you actually care about, the message count and word count concepts cover it on its own terms, including how a per-active-day view changes the picture for chats with long gaps.

The ±0.15 dead-even band

Anything inside ±0.15 is reported as dead even, and no winner is named. The reason is that most chats land there. A gap between 0.07 and 0.04 is not a finding about a relationship, it is the residue of one dead phone battery, one week of holiday, or one long thread that happened to fall inside the export.

Small scores are also unstable across time windows. Shift the wrapped period by a month and a 0.05 can change sign; a 0.6 almost never does. Reporting a direction that flips depending on where the export starts would make the number feel precise while being close to arbitrary.

There is a design argument too. This is a stat people screenshot and send to the other person in the chat. A tool that manufactures a winner out of noise is handing someone an argument based on rounding, so the band exists to say plainly that the two of you are reaching for each other at about the same rate.

What the score leaves out on purpose

It has no view on tone. A fast "k" scores exactly like a fast paragraph, and warmth, sarcasm, apology and irritation are invisible to it. Sentiment scoring on private messages is unreliable across languages, slang and inside jokes, and a wrong warmth verdict is worse than no warmth verdict.

It ignores who says the emotionally heavy things first. Nothing in the score tracks who said "I love you" first, who says it more, or who asks the questions. Phrase-level features like top expressions and signature words exist elsewhere in the recap, and they deliberately do not feed this number, because content of affection is not the same as initiative.

It does not know about per-message length as an intensity signal, and it cannot see anything that happened off WhatsApp. Calls, other apps, living in the same flat and spending Saturday together all suppress texting for the person who is closest to you. That is the single biggest reason a score should never be read as a measure of who cares more.

And it cannot see restraint. Someone who deliberately gives the other person space during a hard week looks, to four counters, exactly like someone who lost interest. The concepts of chat silence and ghosting are described separately precisely because a gap needs context that arithmetic does not have.

Reading your own number without over-reading it

Split direction from magnitude. Direction says who is reaching more; magnitude says how confident that reading is. Above roughly 0.4 the pattern usually survives changing the time window. Just outside the dead-even band it often does not.

Then look at which inputs produced it. A +0.5 built almost entirely on day-openers is a story about schedules and morning habits. A +0.5 where all four inputs point the same way is a pattern that shows up in every part of the chat and is harder to explain away. The component breakdown is the useful part; the headline number is just the summary.

Finally, check the window. A recap covering a period that includes a two-week trip, a breakup, or the first month of a new chat is describing that period, not the relationship. Re-run a different window and see what holds. The version worth sharing is the one that survives the question "what else was going on then".

Try it with your chat

Frequently asked questions

What exactly do the four weights add up to?+

Who starts the day counts 30%, who double-texts 25%, who replies faster 25%, and who writes more 20%. The result is a single figure from -1 to +1 showing which side is reaching more.

Why does writing more count for the least?+

Volume reflects style as much as interest. Some people narrate their day in paragraphs regardless of who they are talking to. It stays in the calculation because an extremely one-sided chat is real, but at 20% a talkative person is not automatically labelled the one doing the chasing.

If one person sends ten messages in a row, does that count ten times?+

No. A run of two or more messages with no reply in between counts once, however long it is. Otherwise the metric would measure how people break up their sentences rather than their willingness to write again unanswered.

Why does my recap say dead even instead of naming someone?+

Because the score landed inside ±0.15, where most chats sit. Differences that small usually come from artifacts like a dead battery or a holiday week, and they can flip sign if the time window shifts, so no winner is declared.

Does the score take tone or affection into account?+

No. It does not read tone, sentiment, or who says the emotionally significant things first, and it cannot see calls, other apps or time spent together. It only compares four behavioural counters, which is why it should be read as a description of initiative, not of feeling.

How high does the score need to be before it means something?+

Direction becomes reasonably stable when the magnitude is well above the dead-even band, around 0.4 and up, and it is stronger still when all four inputs point the same way. Just outside ±0.15 with only one input driving it, the pattern often disappears in a different time window.

Related terms