New research indicates that many scientific papers claiming men and women respond differently to treatments or interventions lack the statistical evidence to back up those claims. When analyzing recent studies in the behavioral and brain sciences, scientists found that fewer than one in four articles properly compared male and female responses. These findings, published in PNAS, raise concerns about the actual evidence behind medical recommendations tailored specifically for different sexes.
Public health initiatives have increasingly pushed for the inclusion of both female and male subjects in scientific research. Policies implemented by major funding organizations aim to ensure that diseases and treatment responses are understood across all populations. The underlying rationale is that biological differences between the sexes might influence how medical conditions develop or how patients react to drugs.
“Sex differences have always captured people’s attention, so they have been a focus of biomedical and behavioral research for a long time,” said Donna L. Maney, a professor of psychology at Emory University and senior author of the study. “A few years ago there were some policy changes designed to increase the number of studies that include both females and males, so I was already studying how these policies were increasing the number of sex differences being reported.”
“What I noticed was that it seemed like most studies claiming sex differences didn’t actually compare the sexes with each other – instead, they were using a flawed analytical approach well-known to be invalid,” Maney said.
This observation aligns with previous meta-research on the topic. A 2020 survey of neuroscience papers indicated that although half of the published experiments included both male and female subjects, only 15 percent actually ran statistical tests to compare them.
A subsequent 2022 survey of neuroscience and psychiatry research echoed this pattern. That study found that despite a rise in the inclusion of both sexes over a ten-year period, only 5 percent of the analyzed papers used sex as a primary variable to look for actual differences. These earlier surveys demonstrated that neuroscience studies rarely analyzed sex as a meaningful variable, even when male and female subjects were present.
This research progression prompted scientists to test whether published studies that explicitly claim to have found sex differences actually used valid statistical comparisons to support their statements. The researchers behind the new study focused exclusively on published articles that featured claims of sex-dependent effects directly in their titles. The authors searched a massive database of academic literature for terms like “sex-specific” or “gender-dependent.”
Out of more than a thousand matching papers, they selected a representative sample of 200 articles published between 2019 and 2023 in fields related to the brain and behavior. Half of the selected studies involved human participants, while the other half relied on nonhuman animals, predominantly rodents like mice and rats. The researchers read each paper to see how the original authors analyzed their data to reach their conclusions about sex differences.
“To test whether an effect actually differs between sexes, you have to compare the effect itself between the males and the females statistically,” Maney explained. “That is, to claim a sex difference, you must compare the sexes with each other.”
“Surprisingly often, that comparison isn’t being made,” Maney said. “In our sample of 200 papers with a claim of a sex-specific effect in the title, fewer than one in four contained a statistical comparison that supported the claim.”
“Some of the papers actually did contain the appropriate comparison across sex, but the difference was not statistically significant,” Maney said. “That means that the researchers put a claim in the title of their paper that was inconsistent with the analysis they presented in the paper. Others wrote in the paper that they did the comparison but they didn’t report the results.”
“The magnitude of the problem surprised us,” Maney added. “More than half the papers contained no statistical comparisons between females and males, yet they claimed a sex difference in the title.”
Many of these studies relied on a common statistical error. “In this approach, let’s say you are testing the effects of some kind of treatment,” Maney said. “It’s well-known that if you divide your sample into subgroups and test each group independently, you are likely to miss real effects simply because you have fewer subjects within each subgroup than in all of them put together.”
“Dividing into subgroups makes it harder to detect a real effect, so you end up missing things,” Maney continued. “When you divide the sample into males and females and test the effect separately within each group, you can easily miss the effect in one of the sexes, and end up concluding the treatment worked only in one sex. It’s a well-known statistical error that a lot of people have written about and advised against, but the error is still commonly made.”
The use of proper statistical tests varied depending on the specific field of study. Psychology papers had the highest rate of appropriate evidence, with about 39 percent of the articles using the correct tests to back up their titles. Neuroscience articles had the lowest rate, with only about 18 percent providing appropriate statistical support for their claims.
Studies involving human participants provided valid evidence about 34 percent of the time, compared to just 15 percent for studies using nonhuman animals. “We were also struck that using an invalid approach did not prevent papers from getting into highly ranked journals, and they were still cited just as often as papers that provided support for their claims,” Maney told PsyPost. “So this isn’t a problem confined to obscure papers or inexperienced researchers.”
“To me, this finding could suggest that sex differences are so popular that claiming one might be more important for publishing and getting citations than showing evidence for it,” Maney said. “But our data don’t speak to that directly.”
Indeed, the researchers found that neuroscience articles claiming sex-dependent effects in their titles increased more than fivefold over a twenty-year period.
“That matters because these findings can influence how scientists think about female and male biology and, potentially, how treatments are developed or recommended,” Maney explained. “In our sample, we saw calls to treat men and women differently when it comes to things like anxiety and depression, suicide prevention, diagnosing mental health conditions, and so on, all without having compared the sexes with each other at all in their study.”
“One important point is that we didn’t show that all these claims of sex-specific effects are false,” Maney emphasized. “For most of the papers we classified as inappropriate, the necessary comparison either wasn’t performed or wasn’t reported. Some of those claims might well hold up if the data were reanalyzed. Less than 9% of the papers contained reported statistical evidence that clearly contradicted the claim in the title.”
“Our point is not about whether sex differences exist,” Maney added. “It’s about what counts as evidence for them. If we claim that an effect differs between the sexes, we need to actually test whether it differs between the sexes.”
“One next step is to go back to the published studies and reanalyze the underlying data when possible, so that we can learn how often these claims hold up when the sexes are compared appropriately,” Maney said. “The current study tells us that the evidence needed to evaluate many of the claims wasn’t reported; it doesn’t tell us what the result of the missing comparison would have been. So we are working our way through all those papers to find out whether the claims hold up, using the data from the papers themselves.”
“More broadly, we want to understand how researchers make decisions about analyzing and interpreting sex-related variation and to develop better tools for doing that work,” Maney said. “For example we developed a web site called SexDifference.org, where you can enter some of the information reported in a paper to find out whether it supports a claim of a sex difference and how large that difference is.”
The lack of preregistered hypotheses among the sampled papers complicates the issue further. Only two out of the 200 studies showed evidence of registering their study designs ahead of time. In many cases, researchers might analyze their data after an experiment, notice a variation between males and females, and elevate that exploratory observation to the main finding of the paper.
“I think it’s important to distinguish sex inclusion from rigorous sex-differences research,” Maney said. “The push to include females in biomedical research addressed a real historical problem. But once both sexes are included, we still have to use appropriate methods to determine what the data tell us about sex. Our results suggest that the next phase of sex-inclusive research needs to focus not only on representation, but also on the rigor of the analyses and the conclusions we draw from them.”
Future scientific endeavors might need to shift their focus away from binary comparisons entirely, exploring continuous or multivariate biological characteristics instead. “Ultimately, I would like to see the field move beyond simply asking whether females and males are different and toward asking more informative questions about the sources of variation among individuals,” Maney concluded.
The study, “Claims of sex-dependent effects in the behavioral and brain sciences are infrequently supported by valid statistical comparisons,” was authored by Madeline T. Olivier, Andrew W. Brown, Simon Chung, Colby J. Vorland, and Donna L. Maney.
Leave a comment
You must be logged in to post a comment.