theattractivenesstest.com

Comparison

The ChatGPT attractiveness test: what it is really doing

Asking a chatbot to rate your face out of 10 became a mass habit before anyone checked what was under it. The short version: it is generating a plausible-sounding number, not measuring one.

Generated versus computed

A language model reading your photo does not locate your eyes, measure the distance between them, or compare that distance to anything. It converts the image into a representation, then produces the text most likely to follow your request. When you ask for a rating out of 10, the most likely text is a rating out of 10 - typically between 6.5 and 8, with a compliment attached.

Nothing in that process is a measurement. There is no reference population, no defined quantity being measured, and no reason for two runs to agree. Users notice this immediately: the most common complaint about chatbot face ratings is that the score changes when you ask twice.

Side by side

 Chatbot ratingLandmark measurement
Same photo twiceDifferent answerIdentical to the decimal
What it measuresNothing - text is generated478 landmark coordinates
Reference populationNone29,927 measured scans
Shows its workingNoEvery sub-score printed
Photo leaves deviceYes, uploadedNo, stays in the browser
BiasTrained to flatter and hedgeProportion targets are culturally loaded

The last row matters. A measurement test is not free of bias - the proportion targets it scores against carry a specific aesthetic history. The difference is that the bias is inspectable: you can read the formula and see exactly which assumption produced which number.

Where a chatbot is genuinely better

At describing rather than scoring. Ask a model what it notices about a face - features, colouring, the impression it gives - and you get something a landmark mesh cannot produce, because geometry is blind to skin, hair, expression and style.

The failure is specifically in the number. Description is what language models do; measurement is not, and dressing a generated impression up as a score out of 10 gives it a precision it does not have.

Questions people ask

Is the ChatGPT attractiveness test accurate?
It is not a test and it is not measuring anything. A language model looks at your photo and produces text that fits the pattern of a rating - it has no landmark detection, takes no measurements, and holds no reference distribution to compare you against. The number it gives is generated, not computed, which is why the same photo can come back as a 6 in one chat and an 8 in another.
Why does ChatGPT give different scores for the same photo?
Because the output is sampled from a probability distribution over text rather than derived from the image. Rephrase the prompt, start a new chat, or simply ask again and the answer moves. A measurement pipeline behaves the opposite way: identical input, identical output, every time, to the decimal.
Why does ChatGPT usually give a flattering number?
Models are trained to be helpful and to avoid causing harm, and telling a stranger their face is unattractive is squarely in the territory both of those pull away from. The result is a strong upward bias plus a lot of hedging. That is a reasonable design choice for a chatbot. It also makes the number meaningless as a measurement.
Will ChatGPT refuse to rate my face?
Often, yes. Appearance rating sits close to several policy boundaries, so refusals and heavily qualified answers are common, and behaviour differs across models and versions. Much of the prompt-engineering advice circulating online is aimed at working around those refusals, which should itself tell you something about the reliability of what comes out the other side.
What does a geometric test do instead?
It places 478 landmarks on your face, computes distances and angles between them, converts those into symmetry and proportion sub-scores, and compares each against a fixed distribution of previously measured faces. Every step is arithmetic on coordinates. It is a much narrower claim than 'how attractive are you', and it is one the method can actually support.
Is it safe to upload my photo to a chatbot for this?
It is a real upload to a third party, subject to that provider's retention and training policies, which change over time. A photo of your face is biometric data in several jurisdictions. A test that runs in your own browser never transmits the image at all, which sidesteps the question entirely.

Run the measured version

Same photo, same numbers, every time - and the photo never leaves your browser.

Start the test