Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

A scholar of how scientific research is conducted and of statistical errors that show up in many peer-reviewed scientific publications, Uri Simonsohn, has devoted much thought with his colleagues to the issue of "p-hacking." Simonsohn is a professor of psychology with a better than average understanding of statistics. He and his colleagues are concerned about making scientific papers more reliable. Many of the interesting issues brought up by the comments on the article kindly submitted here become much more clear after reading Simonsohn's various articles[1] about p values and what they mean, and other aspects of interpreting published scientific research.

Simonsohn provides an abstract (which links to a full, free download of a funny, thought-provoking paper)[2] with a "twenty-one word solution" to some of the practices most likely to make psychology research papers unreliable. He has a whole site devoted to avoiding "p-hacking,"[3] an all too common practice in science that can be detected by statistical tests. You can use the p-curve software on that site for your own investigations into p values found in published research.

He also has a paper on evaluating replication results[4] (an issue we discuss from time to time here on Hacker News) with more specific tips on that issue.

"Abstract: "When does a replication attempt fail? The most common standard is: when it obtains p>.05. I begin here by evaluating this standard in the context of three published replication attempts, involving investigations of the embodiment of morality, the endowment effect, and weather effects on life satisfaction, concluding the standard has unacceptable problems. I then describe similarly unacceptable problems associated with standards that rely on effect-size comparisons between original and replication results. Finally, I propose a new standard: Replication attempts fail when their results indicate that the effect, if it exists at all, is too small to have been detected by the original study. This new standard (1) circumvents the problems associated with existing standards, (2) arrives at intuitively compelling interpretations of existing replication results, and (3) suggests a simple sample size requirement for replication attempts: 2.5 times the original sample."

[1] http://opim.wharton.upenn.edu/~uws/

[2] http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2160588

[3] http://www.p-curve.com/

[4] http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2259879

AFTER EDIT: Hat tip to HN participant jasonhoyt for noticing that the URL on the thread-opening submission was not the canonical URL, and doesn't point to the latest preprint version of the article we are discussing in this thread. The canonical URL (which is generally to be preferred for a posting to HN) is

https://peerj.com/preprints/447/



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: