Methodology

How the Compass works

The constructs it measures, where they come from, how scoring works, and exactly how far along its validation is. Published because you should be able to check the thing before you trust it.

Section 1

What it measures

Four blocks. Two of them are rated on a scale and produce scores. Two of them are described by you and produce no score at all, which is a distinction the report repeats because it changes how much weight each part can carry.

Why do I collect?

Six motives. You choose which apply, put them in your own order, then break the top three down by facet. Nothing here is scored. This block is a stated preference, not a measurement, and the report says so where it appears.

NostalgiaConnection to your own past and the people in it
FandomAttachment to specific players, teams, or characters
CompletionSatisfaction from finishing sets, runs, checklists
Long-Term InvestmentCards as assets worth holding, and the patience that implies
The ChaseChance, pack opening, the pull itself
CommunityPeople, trading, shows, shared experience

Each motive carries six facets, thirty-six in total, which let the report describe what specifically appeals rather than only which category does.

What do I collect?

Six fixed-choice questions covering category, sports, games, what you look for, the shape of your collection, and your state. Descriptive only, unscored, and every option is a fixed choice. There is no field anywhere in the assessment that accepts typing.

How do I approach it?

Four tendencies, sixteen rated statements. Higher is not better or worse; it is simply more of that tendency.

ResearchChecking before acting, and knowing what you hold
BudgetingSetting a limit, staying inside it, and being accountable for it
Heat of the MomentBuying fast under hype, a live break, or right after a result
Risk AppetiteWillingness to accept an uncertain outcome for upside

Research and Budgeting are measured separately from Heat of the Moment because they fail in different ways. A collector can be excellent at homework and still lose money at 11pm in a live break, and combining them into one score hides exactly the person most at risk.

Where do I need to watch out?

Three more, fourteen rated statements. This is the part of the report that is free for everyone, and it is never placed behind a payment.

Learning from MistakesWhether a costly pattern changes or repeats
Fits My LifeWhether the hobby sits alongside the rest of your life or competes with it, in time and in plans
Hard to Step BackControl, chasing, concealment, using the hobby to escape

The last two replace the terms "harmonious" and "obsessive" passion from the research literature. The constructs are unchanged. The labels are plain because people read these about themselves, and "obsessive" reads as a diagnosis to a non-specialist. This instrument does not diagnose.

On the report, two of these three are reversed. Learning from Mistakes and Fits My Life are worded so a high score is a good thing, while Hard to Step Back runs the other way. Shown side by side, two bars would point one direction and one the other. So the report reverses the first two, renames them to match, and prints the underlying score beneath each one.

Section 2

Where the constructs come from

Every construct here is drawn from published research. The items that measure them are original to this instrument and have not been tested. That distinction matters: a construct having support in the literature says nothing about whether these particular statements measure it well.

ConstructSource
Six collecting motivesParrott (2026), Communication & Sport, on why people collect sports cards
Completion as a distinct driverMcIntosh & Schmeichel (2004); Carey (2008)
Nostalgia as a social emotion, and its effect on pricesWildschut et al. (2006); Aquino & Gershenson (2024)
Passion that fits a life versus passion that controls itVallerand et al. (2003); Vallerand (2010), the dualistic model of passion
Card pack spending and problem gamblingXiao et al. (2025), Psychology of Addictive Behaviors
Undisclosed pull odds in physical card productsLo, Xiong & Xiao (2026), Journal of Behavioral Addictions
Collecting versus hoardingNordsletten & Mataix-Cols (2012)
Financial risk toleranceGrable & Lytton (1999); Kuzniak et al. (2015)

The items themselves are original wording, informed by these constructs. No item is copied from a published scale.

Section 3

How scoring works

Fixed arithmetic. The same answers always produce the same report, every time, for everyone.

The rated scales

Each of the seven rated constructs is the mean of its items on a 1 to 5 scale, after reverse-keyed items are flipped. Six of the seven use four statements; Hard to Step Back uses six, because it covers four distinct facets and needed the room. Nine of the thirty items are reverse-keyed, so agreeing with everything does not produce a high score across the board.

Bands cut at 2.5 and 3.5, and the combined exposure reading uses the same two cuts. These are absolute, not percentiles: with no validation sample there is no comparison group, so a high band means you agreed with those statements and nothing more. The report states this on the page rather than in a footnote.

Mixed readings. Several scales contain statements worded in opposite directions, so agreeing with a statement and with its opposite is possible. When that happens the average lands mid-scale, and reporting it as a confident position would describe something you never said. The report still shows the score and the bar, because the number is real, but it labels that scale Mixed signals instead of naming a band, uses the middle description, and says why. Answering both directions the same way is not necessarily careless. It often means the honest answer depends on the situation, and that is worth seeing rather than averaging away. Giving the same answer to nearly every statement in the whole questionnaire is treated differently and more bluntly, because it weakens everything in the report.

One predicted cross-loading, recorded before the data exists. Two statements measure things close enough that they may not separate: "I could tell someone close to me what I spend without being embarrassed" sits in Budgeting, and "I've hidden or downplayed my card spending from people close to me" sits in Hard to Step Back. One is comfort with disclosure, the other is active concealment, and they are being left on separate scales deliberately rather than merged on judgement. If the factor analysis puts them together, that is a finding about the construct and it will be reported as one. Predicting it in advance is worth more than adjusting for it afterwards, which is why it is written here rather than decided quietly.

Two other thresholds are used, and both are provisional choices rather than findings. The preparation-against-pressure grid splits at 3.0 on each axis, and the report says so when a score sits within 0.25 of that line. The gap note appears when preparation and steadiness differ by 1.5 or more. The exposure reading is marked for a closer look when more than one of the four Hard to Step Back statements is answered at Agree or above. That last one is a provisional cutoff chosen by reasoning about facet coverage, not calibrated against data, and it is the single number here most in need of a pilot sample.

The two derived readings

Preparation is the average of Research and Budgeting. Steadiness is 6 minus Heat of the Moment. Plotting one against the other gives four decision profiles, and the distance between them is often more informative than either number alone: a collector who prepares thoroughly and then abandons the plan under pressure is a different problem from one who never prepared.

A combined reading across risk appetite, pressure, budgeting, learning and stepping back produces a single exposure band. It contributes to the watch-out section and never to anything sold.

Why the ranked block has no scores

Ranking produces ipsative data: the values depend on each other, because moving one motive up necessarily moves another down. That has real consequences, and they are worth stating plainly rather than hiding. Internal consistency cannot be computed for those six motives. They cannot be factor-analyzed. Correlations among them are negative by construction.

This is deliberate rather than a shortcoming. Four rating statements per motive, unvalidated, would be weaker evidence about your primary reason for collecting than simply asking you to put them in order. The block stops claiming to measure six latent constructs and instead elicits a stated preference, which is what the report was always using.

Where the narrative comes from

Descriptions are selected by band from a fixed set written in advance. Some paragraphs appear only when a condition holds, such as a sentimental motive and the investment motive both landing in your top two. Wherever the report flags something, it prints the statements that drove it, in the instrument's wording and with the answer you gave, so nothing is asserted without its evidence. Where a reading comes from a scale average rather than any single statement, it says that instead, and where more statements qualified than fit on the page it says how many were held back.

No model writes any of this at the time you take it.

Section 4

What the report tells you

Every number in the report is explained where it appears, alongside what it is an average of. Any flag or elevated reading lists your own answers that produced it, with what you selected, so you can judge whether it reflected a real pattern or a bad week.

The report avoids clinical vocabulary entirely. It describes behaviors and patterns, never conditions. It says "some of your answers suggest," never "you have."

The report also states plainly that a different week could produce a different result, and invites you to retake it. A single reading presented as a fixed fact about a person is the failure this is designed to avoid.

Section 5

Validation status

Being straight about this matters more than looking finished.

DraftThe instrument is version 0.7 and has not completed validation.
ProvisionalScores are a structured prompt for reflection, not a measurement.
No normsYour score is not compared against other collectors. There is no sample yet.

A validated psychological instrument has been through several stages: interviews to check that people read the items as intended, a pilot large enough to test whether the scales hold together, a second independent sample to confirm it, comparison against established measures, and a retest to check stability over time.

The Compass has completed none of those. It was built from the published research described above, and the structure is defensible, but a structure built from theory is a hypothesis until data tests it.

The plan, and the honest timeline

Validation data comes from people taking the assessment while it is free. There is no paid research panel, because someone who passes a "do you collect cards?" screener is a poor stand-in for the people this is actually built for.

At realistic volumes that means roughly 12 to 18 months before there is enough data to say anything firm. This page will be updated as that changes, including if the results are unflattering.

One test will decide a central claim

The Compass treats deliberate habits and in-the-moment reactions as two separate things. That is a falsifiable bet, not a preference. If the data shows they are really one thing, the model is wrong and it will be changed, and this page will say so.

Section 6

Limitations

  • It is a short self-report questionnaire. It only knows what you told it. It does not know your finances, your circumstances, or what kind of week you have had.
  • Self-report under-detects loss of control. People in difficulty tend to answer guardedly. A reassuring result is not evidence that everything is fine, which is why the balanced reading is stated flatly rather than as congratulation.
  • Short scales are less reliable. These use four statements each, and six for Hard to Step Back, which trades some precision for a questionnaire people will actually finish.
  • Archetypes simplify. A label is a summary of two scores, preparation and steadiness, and not a category you belong to. They are named for the pattern rather than for the person, and where a score sits near a boundary the report says so instead of naming the cell confidently.
  • It cannot diagnose anything, and it is not a substitute for a professional who can.
Section 7

Data and privacy

There is no account, no email, and no name. There is nothing to link a result back to a person, because nothing identifying is ever collected.

With your permission, given on the consent screen, the anonymous answers alone are kept so the instrument can be improved and eventually validated. The privacy policy lists exactly what the record contains, field by field, and that list is the authoritative one. In short: your item responses, the context questions, your motive ranking, the scale means calculated from your answers, the instrument version, how long you took, a random identifier and a timestamp. No email, no name, no IP address stored with your answers, and no free text anywhere, because there is no field in the assessment that accepts typing. There is no person attached to a row, which is the point.

You can decline that and still take the assessment and get exactly the same report.

Your report is generated once and handed to you. It is not stored anywhere, which means it cannot be recovered without retaking the assessment. Download it before you close the tab.

Section 8

Ethics

  • Adults only. Card packs are bought by children. This questionnaire asks about spending and control, and it is written for adults.
  • Warnings are never sold. There is one report and everything in it is currently free. Whatever the price becomes, the watch-out section stays free. Depth can be sold; safety cannot.
  • Non-diagnostic language throughout. No reading asserts that you have anything. The most serious band says that more than one of your answers points the same way and lists them, and nothing more. The wording for that band has been written to be non-diagnostic and reviewed line by line by the author, who is a doctoral-level industrial-organizational psychologist, against that standard. It has not been reviewed by a licensed clinician. That is a deliberate statement rather than an omission: the four statements behind the most serious reading sit close to recognised screening content, and a clinician is better placed than I am to judge whether a threshold chosen by reasoning rather than by data is set in the right place. That review is sought before the instrument leaves pilot, and this line will say so when it happens.
  • It points outward. Anything in the exposure section is framed as a prompt for reflection and conversation, never a conclusion, and it points you toward the people best placed to help rather than toward a conclusion: someone who knows you, or a professional. It deliberately does not print a helpline number, because an unvalidated questionnaire is not a safe thing to route someone off the back of, and a wrong or stale number is worse than none.
Who built it

A doctoral-level industrial-organizational psychologist who designs assessments professionally, and who collects cards. Not a clinician. That is precisely why the exposure section points toward people who are.

Section 9

How this was built

This assessment was developed with substantial help from an AI system, Claude. That is worth stating plainly rather than leaving for someone to work out.

What the AI did

Claude was used as a thought partner throughout. It searched and summarized the research literature, drafted candidate items, proposed the scale structure, wrote first drafts of the scoring rules and the report narrative, built the software, and argued with me about design decisions. A large share of the words in the report you receive were drafted by it.

What I did

Every decision that shaped the instrument is mine, and several of them went against the first proposal I was given. A few examples, so this is not an empty claim:

  • The original structure was organized around three alliterative section names chosen because they made a nice acronym. I scrapped it, because the labels had started dictating the measurement instead of describing it. The plain Why, How, and Where questions replaced them.
  • I set the standard for how rigorous different parts need to be. The sections dealing with money and control are held to a strict bar. The sections about why you collect are allowed to ship while still provisional, because being slightly wrong about someone's motives is not the same kind of error.
  • I rejected the proposal to buy a research panel for validation data, on the grounds that people recruited through a survey panel are a poor stand-in for actual collectors, no matter what the screener says.
  • I decided what goes behind a payment and what does not. Nothing in the exposure section is ever sold.
  • I reviewed and approved every item, every scoring rule, and every line of narrative that a reader sees.

I am a doctoral-level I/O psychologist. Assessment design is what I do professionally, and I applied the same judgment here that I would apply to an instrument built for a client.

What that means for you

It should not change how much you trust the results, in either direction.

An item drafted by an AI and an item drafted by a person are in exactly the same position before validation: both are hypotheses about how people will answer, and neither is evidence until data tests them. The validation plan in Section 5 is what would make these scales trustworthy, and it applies identically regardless of who typed the first draft.

What AI assistance did change is speed and breadth. An instrument like this would normally take months of literature review and drafting. That is a real advantage, and it is also a real risk, because fluent text is persuasive whether or not it is correct. The rigour has to come from somewhere else: from the sources, from the review, and eventually from the data.

Where AI appears in the product

There is one report. It is generated in your browser from fixed arithmetic, so the same answers always produce the same report, and no model writes any of it at the time you take it.

The assessment itself uses no AI at scoring time. Your answers go through fixed arithmetic and a fixed set of rules. The same answers always produce the same report, which is a deliberate design decision: a scoring process that varies between runs cannot be checked, and would make this page a work of fiction.

Questions about any of this?

If something here does not hold up, I would rather hear it than not. The document behind this page is longer and more technical, and I am happy to share it with anyone who wants to look properly.

Take the Compass About