Junk DNA: An Archive Nobody Curated

Bar chart of the human genome: endogenous retroviruses at about 8 percent against 1.5 percent protein-coding, the non-coding majority once called junk DNA

An onion has a genome roughly five times larger than yours.

Not five times more genes — five times more DNA. The marbled lungfish has about forty times more. Some ferns and amoebae are larger still. Meanwhile a pufferfish gets by with roughly one-eighth of your genome and is a perfectly functional vertebrate with essentially the same body plan.

If every base pair were doing essential work, this makes no sense. An onion is not five times more complex than a person. The gap is what gave us the phrase junk DNA – and the argument about whether the label was ever right.

This is the C-value paradox, and the evolutionary biologist Ryan Gregory turned it into a challenge known as the onion test: whatever function you propose for all that extra DNA in humans, explain why an onion needs five times as much of it.

It is the single hardest question in the junk DNA debate, and it is the one almost never addressed.

What Is Junk DNA?

“Junk DNA” refers to portions of the genome that do not code for proteins and were long assumed to have no function. The term was coined by geneticist Susumu Ohno in 1972.

The proportions are stark:

  • Protein-coding sequence: about 1.5% of the human genome
  • Transposable elements: close to half — LINEs, SINEs, DNA transposons
  • Endogenous retroviruses: roughly 8% — sequence of viral origin
  • Pseudogenes: thousands of broken copies of once-working genes
  • The remainder: introns, repeats, regulatory regions, and large stretches of unclear status

So roughly 98.5% of your genome does not encode a protein. The question is what, if anything, it does instead.

The ENCODE Fight

In 2012 the ENCODE consortium — a large international project mapping functional elements in the human genome — published results announcing that around 80% of the genome is biochemically functional.

The press coverage was unambiguous: junk DNA is dead.

The backlash from evolutionary biologists was immediate and fierce, and the dispute turned entirely on one word.

ENCODE defined “functional” as showing reproducible biochemical activity — being transcribed into RNA, bound by a protein, or carrying a particular chemical mark. That is a measurable, defensible technical criterion.

Critics argued it was the wrong criterion. Dan Graur and colleagues published a rebuttal whose title conveys the temperature of the exchange, and the substance of the objection was this: transcription is not the same as function. Cells transcribe a great deal of RNA that is immediately degraded. Proteins bind DNA non-specifically all the time. A sequence being chemically busy does not mean the organism needs it.

The analogy offered: a car engine produces heat and noise. Both are reproducible outputs. Neither is what the engine is for.

Graur’s counter-argument used a different definition — selected-effect function, meaning a sequence is functional if it is maintained by natural selection, which can be tested by looking at how strongly it is conserved across species. By that measure, the fraction of the human genome under purifying selection comes out at roughly 8-15%.

Where the Field Actually Landed

The honest answer is that both extremes were wrong, and the resolution is genuinely interesting.

Much non-coding DNA is definitely functional. This is not in dispute and was known before ENCODE:

  • Regulatory elements. Promoters, enhancers, silencers and insulators control when and where genes are expressed. Enhancers can sit hundreds of thousands of bases away from the gene they regulate.
  • Non-coding RNAs. MicroRNAs regulate translation. Long non-coding RNAs do many things — XIST, which silences one X chromosome in females, is a large non-coding RNA and is absolutely essential.
  • Structural sequence. Telomeres cap chromosome ends; centromeres are where spindle fibres attach. Neither codes for protein; both are indispensable.
  • Introns. Spliced out of transcripts, but they enable alternative splicing — one gene producing multiple proteins — which is a major source of human proteomic complexity.

And much of it is almost certainly not. Millions of copies of degraded transposable elements, pseudogenes accumulating mutations without consequence, and vast expanses of repeats that vary enormously in quantity between closely related species. The onion test applies with full force here.

The current consensus, roughly: the genome is not mostly junk, and it is not mostly functional. A substantial minority is doing identifiable work, a substantial fraction is inert, and the boundary is unresolved and depends on which definition of “function” you adopt.

Why the Genome Looks Like This

The more useful question is not how much is junk but why a genome would accumulate so much of it.

Transposable elements are the main answer, and they are essentially genomic parasites. A retrotransposon copies itself and inserts the copy elsewhere. That serves the element’s propagation, not the organism’s. Most copies degrade into inert sequence over time, but they are not removed, because deletion is not automatic.

Whether they accumulate depends on population size. In large populations, selection is efficient and slightly deleterious insertions get removed. In small populations, genetic drift dominates and mildly harmful sequence can drift to fixation. This is Michael Lynch’s argument: large genomes are partly a consequence of small effective population size rather than a functional achievement.

That framework predicts exactly the pattern observed — bacteria under intense selection with compact genomes, and multicellular eukaryotes with small populations carrying enormous amounts of repetitive sequence.

Blueprint, or Archive?

The persistent framing is of DNA as a blueprint or a program — and a lot of the confusion comes from that picture.

A blueprint has no wasted lines. A program has no dead code that persists for a hundred million years. Both are designed, and design implies efficiency.

A genome is not designed and does not behave that way. It is better understood as an archive that has never been curated: layers of working sequence, deprecated sequence, viral insertions that were domesticated into essential functions, viral insertions that were never useful, duplications, broken copies, and repeats — all accumulated over billions of years, with deletion happening only when it happens.

That framing also makes sense of the genuinely remarkable findings. Syncytin, the retroviral envelope gene that builds the placenta, is exactly what an uncurated archive produces: a parasite arrives, most copies decay, and occasionally one gets repurposed into something the organism cannot live without.

You do not get that from a blueprint. You get it from accumulation plus selection acting on whatever happens to be lying around.

What the Label Cost

Ohno’s term did real damage, in a way worth noting.

For decades, non-coding regions were under-studied because the name signalled they were not worth studying. Research effort concentrated on the 1.5% that made proteins. When regulatory sequence turned out to matter enormously — and when genome-wide association studies found that the large majority of disease-linked variants fall outside protein-coding regions — the field had to go back and look at territory it had labelled as waste.

That is a lesson about naming. Calling something junk is a claim about its function, made before the function is known, and it is very effective at preventing anyone from checking.

It is the same failure that read Roman lime clasts as sloppy mixing for a century, and that filed Göbekli Tepe as a medieval cemetery for thirty years. The observation was recorded accurately. The label attached to it stopped anyone from looking again.

Frequently Asked Questions

What is junk DNA?

Portions of the genome that do not code for proteins and were assumed to lack function. The term was coined by Susumu Ohno in 1972. About 98.5% of human DNA does not code for protein.

Is junk DNA really junk?

Partly. Much non-coding DNA has identified functions — regulatory elements, non-coding RNAs, telomeres, centromeres and introns. A large fraction, particularly degraded transposable elements and pseudogenes, shows no evidence of function.

What did the ENCODE project find?

ENCODE reported in 2012 that around 80% of the genome is biochemically functional, defining function as showing reproducible biochemical activity such as being transcribed or protein-bound. Critics argued this conflates activity with function.

How much of the human genome is functional?

It depends on the definition. By biochemical activity, ENCODE’s figure was around 80%. By selected-effect function — sequence conserved by natural selection across species — estimates run to roughly 8-15%.

What is the onion test?

A challenge posed by Ryan Gregory: since the onion genome is about five times larger than the human genome, any proposed function for all our non-coding DNA must also explain why an onion needs five times as much.

Why do some organisms have much larger genomes than humans?

Largely because of transposable element accumulation, which is influenced by effective population size. In small populations, genetic drift allows mildly deleterious insertions to persist, while large populations remove them more efficiently.

What percentage of human DNA codes for proteins?

About 1.5%. Transposable elements make up close to half, endogenous retroviruses roughly 8%, and the remainder consists of introns, regulatory sequence, pseudogenes and repeats.

Where do disease-linked genetic variants occur?

The large majority of variants identified by genome-wide association studies fall outside protein-coding regions, in sequence that was previously dismissed as junk.

The Onion Is Still Unanswered

The question “how much of the genome is junk?” has consumed an enormous amount of argument, and it may be the wrong question.

Your genome is not a document someone wrote. It is a deposit — four billion years of accumulated sequence, with no editor, no version control and no deletion policy. Viral insertions sit next to genes they were later recruited to serve. Broken copies of working genes sit next to the working originals. Half of it consists of elements that copied themselves in for their own reasons.

Some of that got picked up and put to work, and one such fragment builds the organ you grew in. Most of it did not, and simply stayed, because nothing removes it.

An onion has five times as much. Nobody has explained why, and until somebody does, every confident claim about what the other 98.5% is for is running ahead of the evidence in exactly the way the word “junk” did in the other direction.


Continue Reading