Post
This was clever - students were much less likely to report their data using round numbers in the 2002 data (suggesting it was manipulated).
Data Colada
@datacolada.bsky.social
· 6h
An influential paper by Ariely & Wertenbroch (2002) reported two studies with tampered data.
Post 1 of 2.
datacolada.org/138
12:02 PM · Aug 31, 2026
(I should add - I think this is one of the biggest frauds in terms of impact. I've definitely cited Ariely and Wertenbroch (2002) multiple times in the past, personally.)
Wertenbroch notes that there are several other papers that show demand for precommitment in other contexts. datacolada.org/wp-content/u...
It's funny that the non-round numbers reported are all "multiples of 5", "2" and "69".
Ok I'll admit it, being a statistician looks like a fun gig sometimes. "that's the Sex Joke number, and an authentic dataset will contain people who enter it if given the opportunity"
Right, I kinda now want to just look for missing "69s" in questions as a fraud indicator!
in 20 years, Data Colada will expose somebody for overusing 69 in their fabricated data. The fraudulent researcher is unaware that authentic datasets also have some 67s in there
Is this like how they reported Everests height as 29002 feet when the first survey was 29k on the dot because they didn't want people to think it was rounded
The article at the link is pretty damning, isn't it?
The original data makes no sense. No correlation between "I liked this task" and "I found this task interesting".
Oh yeah, Data Colada take downs are consistently absurdly well documented.
This same data anomaly was also present in the car insurance field experiment that had previously been discredited (reported odometer readings extensively used round numbers at time 1, but not time 2): datacolada.org/98
Not even the decency to use a random number generator of some well-known shape to generate fake observations from the original clumped data? smh. Even LaCour managed that* and he wasn't at Princeton [wait, I'm just now hearing...]
*To make wave 2, take wave 1 and add Normal noise with sd=4 🤪
Oh, I think the opposite.
Like the problem is the data looks *too good* and they didn't account for "students will fraudulently round the numbers (or report "69").
Like it really looks like the data was just generated from a skewed normal distribution.
Whenever I see descriptions of this stuff I think how easy it would be to fake data that actually looks legit, they aren't just frauds they're profoundly lazy too
(I get that they assumed no one would ever look at the data, but still)
Oh I think it's actually very hard to fake if people are actually looking at it. Take point #3 from the paper.
Like you don't just have to fake each variable - you have to fake the relationship between variables!
nobody reported a value of zero in the 2002 data so that's immediately sus, too.