ΨEmpiralis
← Back to Learn

What Actually Makes a Test Scientific

Between a carefully developed questionnaire and a random internet quiz often lie years of research and thousands of participants. A look behind the scenes -- in plain language, with real examples.

Two things that look alike but aren't

A serious questionnaire and an entertainment quiz look identical at first glance: questions, answer scales, a result at the end. That's exactly why it's so hard for most people to tell them apart -- the difference doesn't show up in the design, but in a process you never see as a user: Where did the questions come from? Were they tested on real people? And how does anyone know your result actually means anything?

This article walks through exactly those questions, step by step, covering the core methods of test psychology with concrete examples from tests you'll also find on Empiralis.

Step 1: Where do the questions actually come from?

In a scientific questionnaire, not a single question falls out of thin air. It usually starts with a clear theoretical definition of what's actually being measured, followed by a systematic literature review and often an expert panel that drafts a large initial pool of candidate questions and critiques it collectively.

The PHQ-9, one of the most widely used depression questionnaires in the world, is a good illustration: its nine questions map one-to-one onto the nine diagnostic criteria for a depressive episode as defined in the American diagnostic manual, the DSM. Every single question can be traced directly back to a criterion clinicians already use in practice -- none of it is invented or just "sounds plausible."

A random internet quiz, by contrast, usually has exactly one source: the imagination of whoever wrote it in an afternoon. No theory, no expert panel, nothing to check it against.

Step 2: Are the questions actually tested on real people?

Before a questionnaire is ever published, it typically goes through several pilot studies with hundreds to thousands of participants. Every single question is statistically checked: does it actually distinguish between people with high and low levels of the trait being measured? Or does everyone answer it the same way regardless of how they feel -- in which case it gets cut.

This so-called item analysis is one reason the final version of a test often has far fewer questions than the original pool. Every question that survives has genuinely earned its place.

Reliability: Does the test even give consistent results?

Picture a tape measure that gives you 31, 37, and 26 inches on three consecutive measurements of your waist. It doesn't matter which of those numbers might happen to be correct -- a tape measure that swings that wildly is useless as a measurement tool. That's exactly what reliability checks for in psychological questionnaires: do the individual questions in a test consistently hang together, and does the test give similar results when taken again? The standard metric is Cronbach's alpha, a value between 0 and 1 -- values above .80 are generally considered good.

We go into this in more depth in our article "Reliability and Validity, Explained Simply," including the actual values we document for each test on its info page.

Validity: Does the test measure the right thing?

Reliability alone isn't enough. A tape measure can be extremely precise and still be completely unsuited to taking someone's temperature -- it's reliable, but not valid for that purpose. Psychological tests check validity in several ways. One central approach is criterion validity: do the test's results correlate with an already-established, independent measure the way theory predicts they should? A burnout questionnaire, for example, should correlate with actual sick days -- if it doesn't, it's questionable whether it's really capturing burnout rather than just general low mood.

Equally important is discriminant validity: the expectation that a test should NOT correlate with things that have nothing to do with the trait being measured. A good depression questionnaire shouldn't meaningfully correlate with someone's shoe size -- if it did, something would be wrong with the measurement. This kind of check is entirely absent from made-up online quizzes, since it would require systematically collecting data in the first place.

Norming: What does "above average" actually mean?

When a test tells you that you're "more resilient than 73% of people," that claim is only meaningful if there's a real reference sample behind it -- a sufficiently large, ideally representative group of people who've already taken the same test, with documented results. Only against that can a single result be meaningfully placed.

What matters isn't just the size of that sample, but who's in it: a group of psychology students at one university produces different reference values than a cross-section of the general population across ages and backgrounds. Legitimate test authors disclose exactly where their norms come from. When an internet quiz hands you a suspiciously precise number, like an exact "IQ of 127," without ever naming the comparison group behind it, that precision is pure assertion.

Peer review: the check before anything is even published

Before a newly developed questionnaire appears in an academic journal, independent experts from the same field anonymously review whether the methodology, sample, and analysis hold up -- this process is called peer review. It's no guarantee against every flaw, but it's a real, effective first hurdle against obviously deficient methods.

What makes established tests truly robust, though, is replication: when independent research teams, in different countries, with different samples, over many years, keep arriving at similar results. The WHO-5 well-being index, for instance, has been used worldwide in dozens of languages since the 1990s and keeps getting re-validated. An internet quiz answers to no one and is never independently checked a second time.

One example from start to finish: how an idea becomes a validated questionnaire

The PHQ-9 is a good case study for walking through the whole path at once. It started in the 1990s with an observation: primary care doctors rarely have time for lengthy psychiatric interviews, yet depression was going undetected far too often. A research team led by Robert Spitzer, Kurt Kroenke, and Janet Williams developed a short self-report instrument whose questions, as described above, were derived directly from the official diagnostic criteria.

The questionnaire was then tested on several thousand patients in primary care and gynecology clinics across the US, and its results were compared against structured clinical interviews conducted by trained professionals -- the gold standard for diagnosing depression at the time. Only once it was shown that the short self-report reliably matched the far more time-consuming expert assessment was the PHQ-9 published in an academic journal. It has since been replicated in countless further studies around the world and is now a standard tool in primary care. From the first idea to broad acceptance took several years -- not a weekend project.

How to spot a legitimate test yourself

Does the site cite a concrete original source -- ideally a study with named authors, a publication year, and a journal?

Are reliability or validity figures mentioned at all, or is the claim of rigor just an unsupported assertion?

Is it disclosed where any comparison values or percentages actually come from?

Does the site communicate realistic limits -- for instance, that an online test doesn't replace a diagnosis -- or does it promise strikingly precise, final judgments?

Does the site only reference itself, or can the original source be independently found and read?

Our commitment

Every test on Empiralis links to its scientific original source and openly documents its license status on that test's info page. Where an original version wasn't freely available or its use was legally unclear, we deliberately left that test out rather than reconstruct it from memory or use an unverified translation. That discipline is exactly what separates a scientifically grounded self-test from a random quiz -- and it's why we talk about methods, not just results, at this length in the first place.

Advertisement