(Roughly) Daily

Posts Tagged ‘scientific papers

“Science is a cooperative enterprise spanning the generations… a community of minds, reaching back to antiquity and forward to the stars”*…

The sharing of experimental results and the underlying data is critical to the advance of science. Indeed, when I had the chance to do a scenario planning exercise with a collection of the leading research university librarians in the U.S. a couple of decades ago, the biggest threat/fear they surfaced was the concern that the free and open exchange of ideas and data, as manifest formally in scientific publication and informally in the collegial cooperation among scientists, would be occluded by an increasing proprietary embrace of knowledge.

73% of geneticists surveyed in an article in the 23/30 January 2002 issue of the Journal of the American Medical Association agreed that although keeping data private may help the individual researcher, data hoarding is detrimental to the progress of science Still, sadly, that threat has grown since the turn of the millennium.

By way of current (and dramatic) example: as Celina Zhao reports, more than half of AI “unicorns” have never published a paper or preprint…

Today’s biggest artificial intelligence (AI) startups make no shortage of bold promises. Their technologies, some boast, will revolutionize software development, drug discovery, and scientific research.

Yet a new preprint posted on 16 July on bioRxiv suggests many of these firms barely participate in one of science’s most fundamental practices: publicly documenting discoveries in scientific literature so other researchers can evaluate and build on them. More than half of AI unicorns—private companies valued at more than $1 billion—have never played a leading role in publishing a scientific paper or preprint, according to the new analysis. Collectively, they accounted for just one in every 1000 AI papers published in 2025.

“For a field that is supposedly reshaping science and is so advanced in terms of scientific potential, not having any scientific documentation seems like a very weird paradox,” says paper co-author John Ioannidis, a metascientist at Stanford University [see here]. “How can you judge that what they say is real, validated, and reproducible?” The scarcity of publications, others say, also makes it harder to assess AI’s social impacts, including energy use and safety.

But University of Alberta AI ethicist Mohamed Abdalla says the findings reflect the incentives facing commercial AI developers, rather than solely a failure to uphold scientific norms. “It’s not the company’s job to advance science, right?” he says. “The company’s job is to advance money.”

Ioannidis has long studied how unicorns, particularly in biotech, engage with the scientific literature. (In 2015, he was the first to publicly scrutinize the lack of peer-reviewed studies produced by Theranos, the blood testing startup that proved to be based on fraudulent data.) He wondered whether AI unicorns would show similar patterns.

To find out, he and his team first identified all 317 unicorn AI companies that have existed from 1998 to 2025. Then, they searched for publications affiliated with these startups—including journal articles, conference papers, reviews, and preprints. They selected those where a company researcher played a leading role as a first or last author, indicating the startup had made a substantial contribution to the work. The final data set included 2077 final publications, comprising 1389 peer-reviewed papers and 688 preprints.

More than half of the startups had never produced a single qualifying paper, the analysis revealed. Scientific influence proved even more concentrated, with the top 5% of firms accounting for greater than 90% of all citations. OpenAI alone was responsible for nearly 40% of all citations in the data set, followed by the Chinese computer vision company Megvii and the platform Hugging Face. And even at the most prolific companies, much of the output came from the same small group of repeat authors. For example, despite OpenAI employing roughly 4500 people, only eight researchers had authored five or more qualifying papers.

The findings are unsurprising to some AI researchers given how the industry is structured. For example, unlike the pharmaceutical industry, where published discoveries can be protected by patents, AI companies have learned they often gain little from publicly disclosing technical advances, says Nur Ahmed, an AI researcher at the University of Arkansas. Google’s landmark 2017 paper on the transformer—the architecture that underpins today’s large language models—has become a classic cautionary example, Abdalla adds. Although Google patented aspects of the technology, “I don’t think anybody’s paying Google for that,” he says.

Startups also operate on much faster timelines than academia, where peer review can lumber on for months or even years. That’s why many AI companies have embraced what Avijit Ghosh, an AI policy researcher at Hugging Face, calls the “blogification” of research: announcing new models and releasing code or data sets through blog posts and technical reports rather than scientific journals. The new analysis didn’t track those outputs, he points out.

For Ghosh, the debate shouldn’t center on publishing in journals versus blogs. What matters is whether companies are releasing enough code, data sets, or model weights (the numbers that determine how a model interprets and responds to a prompt) for others to independently verify and build on their work, he says.

The preprint also found that firms based in China consistently published more papers than their counterparts based in the United States. Whereas leading U.S. frontier labs have increasingly kept the details of their most capable models secret or “closed sourced,” leading Chinese companies have embraced “open-source” models. Moonshot AI, one of the Chinese startups included in the study, recently unveiled Kimi K3—one of the strongest open models to date—and publicly released its model weights through Hugging Face today.

But whether models are open or closed, the rapid pace toward increasingly powerful generalist AI worries Emma Pierson, a computer scientist at the University of California, Berkeley. She argues AI research—whether published freely or kept secret—risks accelerating models that pose serious societal and safety concerns, including supercharging cyberattacks. “If we were racing forward on cancer-curing AI, I would be like, ’Fantastic, full steam ahead,’” she says. “But that’s not what we’re racing toward, right?”…

The secretive unicorns: “AI’s top startups are barely publishing their research,” from @science.org.

By way of example? In order to have a broader footprint in AI for (default proprietary) scientific discovery, Google moves away from a successful AI effort (that did publish): “Google DeepMind dismantles Nobel-winning AlphaFold team in strategy shift” (gift article from the FT). One wonders: when these LLMs run out of published papers on which to train, where (and how) will they source the knowledge they need to stay useful?

* Neil deGrasse Tyson

###

As we share and share alike, we might recall that it was on this date in 1887 that Chester A. Hodge of Beloit, Wisconsin received patent No. 367,398 for ‘spur rowel’ barbed wire (consisting of spur shaped wheels with 8 or 10 points mounted between 2 wires).  It was one of many patents for barbed wire (e.g., here), which spread across the American West rapidly (thanks, in no small measure to the guy featured in the almanac entry here)– and (by protecting farmers from foraging free-ranging cattle) paved the way for the expansion of wheat (and other kinds of) farming… even as it spelled the doom of a commons– the open range.

Close-up view of coiled barbed wire, showcasing its intricate twists and pointed spikes.
Roll of modern agricultural barbed wire (source)

“I don’t think academic writing ever was wonderful”*…

Academic writing is famously abstruse. But, Stefan Washietl, founder of Paperpile, reminds us, it isn’t always so. As Rob Beschizza observes

Stefan Washietl collected the shortest scientific papers. Some are unvarnished mathematical proofs, some are humor to amusing or incisive ends, others are clever-dickery that shoves the conclusion into the abstract. All are wonderful!…

Accessible academia: treat yourself to “The Shortest Papers Ever Published,” from @washietl and @paperpile via @Beschizza in @BoingBoing.

* Stephen Jay Gould

###

As we go for the gist, we might send voluminous birthday greetings to Constantine Samuel Rafinesque; he was born on this date in 1783. An autodidact naturalist, traveler, and writer who, in spite of work of variable reliability, substantially expanded knowledge via his extensive travels, collecting, cataloging, and naming huge numbers of plants and some animals. Among these are many new species he is credited with being the first to describe.

Years ahead of Charles Darwin’s theory of evolution, Rafinesque conceived his own ideas. He thought that species had, even within the timeframe of a century, a continuing tendency for varieties to appear that would diverge in their characteristics to the point of forming new species. Accordingly, he was over-enthusiastic at distinguishing what he called new species.

Rafinesque wrote prolifically, and often self-published. His work varied from brilliant insightfulness to carelessness, and raised the eyebrows– and sometimes the ire– of his scientific contemporaries. Indeed, he so incensed John James Audubon with his belief that Audubon has included unnamed species in his sketches of birds, that Audubon pranked him, feeding him sketches of imaginary fish… which Rafinesque believed and included in his writings, where (for 50 years or so) they remained as part of the scientific record.

source

“With my tongue in one cheek only, I’d suggest that had our palaeolithic ancestors discovered the peer-review dredger, we would be still sitting in caves”*…

As a format, “scholarly” scientific communications are slow, encourage hype, and are difficult to correct. Stuart Ritchie argues that a radical overhaul of publishing could make science better…

… Having been printed on paper since the very first scientific journal was inaugurated in 1665, the overwhelming majority of research is now submitted, reviewed and read online. During the pandemic, it was often devoured on social media, an essential part of the unfolding story of Covid-19. Hard copies of journals are increasingly viewed as curiosities – or not viewed at all.

But although the internet has transformed the way we read it, the overall system for how we publish science remains largely unchanged. We still have scientific papers; we still send them off to peer reviewers; we still have editors who give the ultimate thumbs up or down as to whether a paper is published in their journal.

This system comes with big problems. Chief among them is the issue of publication bias: reviewers and editors are more likely to give a scientific paper a good write-up and publish it in their journal if it reports positive or exciting results. So scientists go to great lengths to hype up their studies, lean on their analyses so they produce “better” results, and sometimes even commit fraud in order to impress those all-important gatekeepers. This drastically distorts our view of what really went on.

There are some possible fixes that change the way journals work. Maybe the decision to publish could be made based only on the methodology of a study, rather than on its results (this is already happening to a modest extent in a few journals). Maybe scientists could just publish all their research by default, and journals would curate, rather than decide, which results get out into the world. But maybe we could go a step further, and get rid of scientific papers altogether…

A bold proposal: “The big idea: should we get rid of the scientific paper?,” from @StuartJRitchie in @guardian.

Apposite (if only in its critical posture): “The Two Paper Rule.” See also “In what sense is the science of science a science?” for context.

Zygmunt Bauman

###

As we noodle on knowledge, we might recall that it was on this date in 1964 that AT&T connected the first Picturephone call (between Disneyland in California and the World’s Fair in New York). The device consisted of a telephone handset and a small, matching TV, which allowed telephone users to see each other in fuzzy video images as they carried on a conversation. It was commercially-released shortly thereafter (prices ranged from $16 to $27 for a three-minute call between special booths AT&T set up in New York, Washington, and Chicago), but didn’t catch on.

source

“If I have seen further, it is by standing on the shoulders of giants”*…

 

The discovery of high-temperature superconductors, the determination of DNA’s double-helix structure, the first observations that the expansion of the Universe is accelerating — all of these breakthroughs won Nobel prizes and international acclaim. Yet none of the papers that announced them comes anywhere close to ranking among the 100 most highly cited papers of all time.

Citations, in which one paper refers to earlier works, are the standard means by which authors acknowledge the source of their methods, ideas and findings, and are often used as a rough measure of a paper’s importance. Fifty years ago, Eugene Garfield published the Science Citation Index (SCI), the first systematic effort to track citations in the scientific literature. To mark the anniversary, Nature asked Thomson Reuters, which now owns the SCI, to list the 100 most highly cited papers of all time. (See the full list at Web of Science Top 100.xls or the interactive graphic [above].) The search covered all of Thomson Reuter’s Web of Science, an online version of the SCI that also includes databases covering the social sciences, arts and humanities, conference proceedings and some books. It lists papers published from 1900 to the present day.

The exercise revealed some surprises, not least that it takes a staggering 12,119 citations to rank in the top 100 — and that many of the world’s most famous papers do not make the cut…

Read more (and find an enlargeable version of the infographic above) in Nature’s “The top 100 papers.”

* Isaac Newton

###

As we iterate “ibid.,” we might send send leak-less birthday greetings to a man who facilitated the writing of several of these papers:  George Safford Parker; he was born on this date in 1863. While working as a telegraphy instructor in Janesville, Wisconsin, he became dismayed by the unreliability of his students’ pens.  He experimented with ways to prevent ink leaks; and in 1888, founded the Parker Pen Company.  The next year he received his first fountain pen patent.  By 1908, his factory on Main Street in Janesville was reportedly the largest pen manufacturing facility in the world.  Parker eventually became one of the world’s premier pen brands, and one of the first brands with a global presence.

 source

 

Written by (Roughly) Daily

November 1, 2014 at 1:01 am