Digital sequence information and the Cali Fund: paying for genes that live in databases

· 8 min read · by Joaquín Ferreyra

A rack of small sample tubes beside a pressed dried leaf on a laboratory bench in low evening light.

For most of the history of bioprospecting, the valuable thing a drug or seed company took from a biodiversity-rich country was physical: a soil microbe, a leaf, a sea sponge. Increasingly it doesn't need the sample. It needs the sequence, and the sequence is usually already online and free to download. That shift quietly broke the main international system for sharing the benefits of genetic resources. At the UN biodiversity summit in Cali, Colombia, in late 2024, governments agreed on a patch: a global fund that companies are asked, but not obliged, to pay into.

What follows is a question-and-answer guide to how the problem arose, what was decided and how to judge whether it works.

What is "digital sequence information"?

A deliberately vague placeholder. Negotiators under the Convention on Biological Diversity started using "digital sequence information", or DSI, in the mid-2010s because they could not agree on a precise term. At a minimum it covers DNA and RNA sequences. Depending on whom you ask, it can also cover protein sequences, data on gene expression and metabolites, and the annotations that tell a researcher what a sequence does.

Most public DSI sits in a small number of databases. The largest is a three-way arrangement, the International Nucleotide Sequence Database Collaboration, linking GenBank in the United States, the European Nucleotide Archive and the DNA Data Bank of Japan. The three exchange data daily and make it available to anyone, without conditions.

Why did it break the Nagoya Protocol?

The Nagoya Protocol, adopted in 2010 and in force since October 2014, is built on bilateral deals. A company or researcher who wants genetic resources from a country asks permission (prior informed consent) and negotiates terms (mutually agreed terms), which can include royalties, technology sharing or joint research. The logic depends on a physical handover. There is a moment of access that the provider country can control.

Sequencing removed that moment. Once a sequence is published, a user anywhere can download it, analyze it and, now that gene synthesis is cheap and routine, build the DNA from scratch without ever touching the original organism. There is no permit to request and often no record of who used what. Provider countries, mostly in the global South, called it a loophole. User countries and most of the scientific community answered that tracking every download would cripple open science, which runs on frictionless sharing.

In practice the use is mundane and constant. A plant breeder looking for drought tolerance can search public databases for versions of a gene in wild relatives of a crop, then edit the matching gene in an elite variety. An enzyme company can mine sequences from microbes collected by someone else's expedition years earlier. Neither needs to contact the country where the organism was found.

The standoff lasted most of a decade. The way out was to stop trying to trace individual sequences and to collect money instead from the companies that profit from sequence data in general.

What was agreed before Cali?

At the Montreal summit in December 2022, the one that adopted the Kunming-Montreal Global Biodiversity Framework, governments agreed in principle to set up a multilateral benefit-sharing mechanism for DSI, including a global fund. Target 13 of the framework also calls for a significant increase in benefits shared from genetic resources and DSI by 2030. Cali was left the hard part: who pays, how much, and where the money goes.

What did COP16 decide?

The Cali meeting ran from late October into the first days of November 2024, finishing in overtime. Its DSI decision created what is now called the Cali Fund. The main elements:

  • Who is expected to pay. Large entities that use DSI and benefit commercially from it. "Large" is defined by thresholds on assets, sales and profits, so that small firms fall outside.
  • Which sectors. The decision lists sectors that rely directly or indirectly on DSI, including pharmaceuticals, nutraceuticals, cosmetics, animal and plant breeding, biotechnology, laboratory equipment for sequencing, and information services associated with DSI, artificial intelligence included.
  • How much. An indicative rate of 1% of profits or 0.1% of revenue, the figures most widely reported from the decision.
  • How binding. Not very. Companies "should" contribute. There is no penalty for not doing so, though governments are invited to take measures that encourage payment.
  • Where the money goes. To biodiversity work in countries, with at least half earmarked for the self-identified needs of Indigenous peoples and local communities, paid either directly or through governments.
  • What doesn't change. Access to public databases stays free, and users are not asked to trace which sequence went into which product.

How big is 1%?

A quick hypothetical shows the design. Take an invented company with $10 billion in revenue and $1 billion in profit. One percent of profit is $10 million. So is 0.1% of revenue. The two yardsticks give the same answer whenever the profit margin is exactly 10%.

Move away from that point and they diverge. A thin-margin business, say a contract laboratory earning 3% on sales, would owe far less on the profit option. A drug company earning 25% would owe far less on the revenue option. If companies can use either yardstick, as the "or" suggests, each will use the cheaper one, which says the drafters cared more about getting firms to participate than about the size of the pot.

Note what the rate is pegged to: the company's profits or revenue, not the slice of business traceable to particular sequences. That was the price of avoiding sequence-by-sequence tracking. It also makes the number look bigger to a finance director, which helps explain the cool reception from industry groups.

Back-of-envelope estimates of what the fund could raise if most eligible companies paid have run large. The operative word is "if".

Who actually profits from DSI?

This is where corporate concentration comes in. The users that capture most of the commercial value from sequence data are not spread evenly. They include the largest drug companies and the largest seed companies (the seed industry consolidated sharply in the 2015–2018 merger wave and now leans heavily on genomic tools in breeding), and, increasingly, technology companies that train AI models on public biological data. Protein-structure and protein-design models are built on decades of openly deposited sequences and structures. The value they generate goes to whoever has the computing power, the labs and the patent attorneys to turn a prediction into a product.

Patents tighten the loop. Information downloaded for free can end up inside a patent claim, and fights over patents on transgenic soybeans showed long ago how broad such claims can get. Free in, proprietary out: that is the business model the whole benefit-sharing debate is really about.

Why are scientists nervous?

Open-science advocates broadly welcomed the fact that Cali imposed no download tracking and no access fees. Their worries concern what comes next.

  • National rules. Some countries already treat genetic information as covered by their access laws; Brazil's 2015 biodiversity law is the usual example. If more countries add DSI to national permitting, researchers could face a patchwork of obligations for data that sits in shared international databases.
  • Database pressure. The decision encourages databases to support the mechanism, for instance by improving information on where samples came from. Better metadata is good science in itself, but some database managers worry about being turned into enforcement tools.
  • Tension with Indigenous data rights. Not everyone wants data to be open. Indigenous data governance frameworks, such as the CARE principles published by the Global Indigenous Data Alliance in 2019, hold that communities should have a say over how data from their lands and knowledge are used. "Open by default" and "consent first" pull in different directions, and Cali did not resolve that.

How does this fit with other treaties?

DSI turns up wherever genetic resources are governed. The High Seas Treaty adopted in June 2023 includes benefit-sharing for marine genetic resources, and their sequence information, from areas beyond national jurisdiction. The FAO's treaty on plant genetic resources has its own long-running argument over whether and how to bring DSI into its multilateral system. And the pandemic agreement negotiations at the World Health Organization have been working through a pathogen-specific version of the same puzzle: how to keep virus sequences flowing fast while making sure the countries that share them get access to vaccines and treatments. These forums are under no obligation to agree with one another, so companies could face several overlapping requests.

What should we watch?

  • Whether anyone pays. A voluntary fund lives or dies on early contributions from recognizable names. Formal arrangements for receiving money were being put in place in early 2025; silence from the largest DSI users after that would tell its own story.
  • Whether governments add teeth. Parties could tie contributions to permits, public procurement or market access, turning "should" into something closer to "must", one country at a time.
  • The United States. The US is not a party to the Convention on Biological Diversity, and many of the largest DSI users are American. Their exposure runs through their operations in countries that are parties.
  • The review. The arrangement is meant to be kept under review, and the next full biodiversity COP, in Armenia in 2026, is the first real chance to revisit the rates and the voluntary design.
  • The money's path. At least half is earmarked for Indigenous peoples and local communities. How directly it reaches them, rather than pooling in ministries, will matter more to the fund's legitimacy than its headline total.

The decision texts are on the Convention on Biological Diversity's website for anyone who wants the fine print. The short version: Cali found a way to ask companies to pay for sequence data without tracking a single download. Whether they answer is, for now, up to them.

Joaquín Ferreyra

Written by Joaquín Ferreyra

Joaquín tracks mergers, market shares and the business models behind agribusiness, pharma and big tech. He likes a good concentration ratio and distrusts any market where four companies sell most of everything.