DNA — deoxyribonucleic acid — is the molecule that carries the instructions for building and running a living organism. It is a chemical archive: roughly 3.2 billion letters of code in a human cell, written in an alphabet of just four letters, packed into a nucleus around six thousandths of a millimetre wide.
Everything that follows is a consequence of one structural fact: DNA is two strands wound around each other, and each strand can be used as a template to rebuild the other. That single property is what makes heredity, repair and the copying of a cell into two daughter cells possible.
The structure of DNA
The sugar-phosphate backbone
Each strand is a repeating chain of two alternating components: a five-carbon sugar called deoxyribose and a phosphate group. They are joined by phosphodiester bonds, forming a rigid scaffold that runs the length of the strand. This is the “rail” of the ladder.
The four bases
Attached to each sugar is one of four nitrogenous bases. This is where the information lives:
| Base | Abbreviation | Pairs with |
|---|---|---|
| Adenine | A | Thymine (T) |
| Thymine | T | Adenine (A) |
| Guanine | G | Cytosine (C) |
| Cytosine | C | Guanine (G) |
A, T, G and C are the entire alphabet. Every protein in your body is ultimately specified by a sequence of those four letters.
The double helix
In 1953 James Watson and Francis Crick, building on X-ray diffraction work by Rosalind Franklin and Maurice Wilkins, described DNA as two antiparallel strands twisted into a right-handed helix. “Antiparallel” means the strands run in opposite directions — one runs 5′ to 3′, the other 3′ to 5′ — which is why the bases face inward and pair so precisely.
Base pairing follows strict rules. A pairs with T through two hydrogen bonds; G pairs with C through three. The G-C pairing is therefore slightly stronger, which is one reason DNA with a higher G-C content needs more heat to separate.
The two strands are held together by hydrogen bonds — individually weak, but collectively strong enough to keep the molecule stable while still being able to be “unzipped” when the cell needs to read it.
The nucleotide
The repeating unit is a nucleotide: one phosphate, one deoxyribose sugar, one base. A strand is simply nucleotides linked in sequence, and the order of those nucleotides is the genetic code.
How two metres fit inside a nucleus
A single human cell contains about 3.2 billion base pairs of DNA. Laid end to end, the DNA in one cell would stretch roughly two metres. The nucleus it must fit inside measures about six micrometres across — around 300,000 times smaller.
It gets there through layered folding:
- DNA wraps around protein spools called histones, forming structures often called “beads on a string”
- That fibre coils into a thicker chromatin fibre
- Chromatin loops and compacts further during cell division into the familiar X-shaped chromosomes
Humans have 46 chromosomes — 23 pairs, one set inherited from each parent. The packing is not merely storage: how tightly DNA is wound determines which genes can be read. Tightly wound DNA is generally silent; loosely wound DNA is accessible.
What a gene actually is
A gene is a stretch of DNA that contains the instructions for making a functional product — usually a protein. The human genome contains roughly 20,000 protein-coding genes, a number that surprised researchers when the human genome was completed in 2003, since it is not dramatically more than far simpler organisms.
The flow of information is often summarised as the central dogma of molecular biology:
DNA → RNA → protein
Transcription
An enzyme called RNA polymerase unzips a section of DNA and builds a complementary messenger RNA (mRNA) copy. The mRNA is a working duplicate — it can leave the nucleus, while the original DNA stays protected inside it.
Translation
In the cytoplasm, a ribosome reads the mRNA three letters at a time. Each three-letter group is a codon, and each codon specifies one amino acid. There are 20 amino acids commonly used in proteins, and the genetic code that maps codons to amino acids is nearly universal across all life — the same codon means the same amino acid in a bacterium and in a human.
The ribosome assembles the amino acids into a chain, which then folds into its working three-dimensional shape. Folding matters as much as sequence: a protein with the right amino acids in the wrong shape does not function.
Introns and splicing
Human genes are not continuous. They contain coding sections (exons) interrupted by non-coding sections (introns). Before translation, the introns are cut out and the exons joined together. Alternative splicing lets one gene produce several different proteins, which is a major reason 20,000 genes can generate a far larger proteome.
How DNA copies itself
DNA replication is semiconservative — each new double helix contains one original strand and one newly synthesised strand. This was demonstrated experimentally by Meselson and Stahl in 1958 and is why the mechanism is so accurate: the old strand serves as the template and the proofreading check for the new one.
The process:
- Helicase unzips the double helix at a replication fork
- Single-strand binding proteins keep the separated strands from re-joining
- Primase lays down a short RNA primer to start synthesis
- DNA polymerase adds complementary nucleotides, reading 3′ to 5′ and building 5′ to 3′
- The leading strand is built continuously; the lagging strand is built in fragments (Okazaki fragments) and then joined
- Proofreading enzymes correct mismatched bases as they go
Fidelity is extraordinary — roughly one error per billion base pairs after proofreading and repair, a far better rate than any human engineering system manages at comparable scale.
Mutations: what changes, and what it does
A mutation is a change in the DNA sequence. It can happen during replication, or be caused by mutagens such as ultraviolet light, tobacco smoke, or certain chemicals.
The main types
- Substitution — one base replaced by another. Sometimes silent (the codon still codes for the same amino acid), sometimes changing one amino acid
- Insertion or deletion (indel) — bases added or removed. If not a multiple of three, this shifts the entire reading frame — a frameshift, which usually destroys the protein’s function from that point onward
- Repeat expansion — short sequences repeated many more times than usual
- Chromosomal changes — large-scale deletions, duplications, or rearrangements
Most mutations are neutral or nearly so. Some are harmful. A small number are beneficial in a particular environment — this is the raw material of evolution, and it is why genetic variation matters for a species’ survival.
Mutation and cancer
Cancer is fundamentally a disease of accumulated DNA damage. Mutations in oncogenes push cells to divide; mutations in tumour suppressor genes remove the brakes. A single mutation is rarely enough — most cancers require several to accumulate, which is why they typically develop over years.
For how this is being exploited therapeutically, see our piece on the new frontier in cancer research in India.
What the other 98% of DNA does
Only about 1.5% of the human genome codes directly for protein. For decades the remainder was dismissed as “junk DNA”, a label that has aged badly.
Much of it is now understood to have function:
- Regulatory sequences — promoters and enhancers that switch genes on and off, often located far from the gene they control
- Introns and untranslated regions that affect how much protein a gene produces
- Structural elements that help package chromosomes
- Repetitive elements — transposons and their remnants make up a large share, some now co-opted for regulatory roles
- Non-coding RNA genes that regulate other genes without ever making a protein
Some sequences appear to be genuinely without function — relics of ancient viral insertions and evolutionary baggage. The honest summary is that “junk DNA” was wrong as a blanket claim and the full picture is still being assembled.
Epigenetics: changes that do not alter the letters
Epi-genetics means changes in gene activity that do not change the underlying sequence. The two best-studied mechanisms are DNA methylation (a methyl group attached to cytosine, usually silencing the gene) and histone modification (altering how tightly DNA is wrapped).
These marks can be influenced by diet, stress, ageing and environment, and some can be passed to offspring. Crucially they are generally reversible, which is what distinguishes them from mutations — and what makes them pharmacologically interesting.
What DNA is used for
Forensics and identity
Short tandem repeat (STR) profiling compares variable repeat regions across several loci. The probability of two unrelated people sharing a full profile is typically in the billions to one, which is why DNA evidence is so powerful — and why contamination handling matters as much as the analysis itself.
Ancestry and genealogy
Population genetics compares allele frequencies between groups to estimate geographic origin and relatedness. Mitochondrial DNA, inherited maternally without recombination, is particularly useful for tracing deep maternal lineages.
CRISPR gene editing
CRISPR-Cas9 uses a guide RNA to direct the Cas9 enzyme to a specific DNA sequence, where it makes a cut. The cell’s own repair machinery then either disables the gene or installs a supplied replacement sequence. Jennifer Doudna and Emmanuelle Charpentier received the 2020 Nobel Prize in Chemistry for developing it.
The technique is transformational for research and is producing approved therapies for certain genetic conditions — but editing human germline cells remains widely prohibited and scientifically controversial.
mRNA vaccines
Understanding how DNA encodes protein is what made mRNA vaccines possible: instead of delivering a weakened pathogen, they deliver the genetic instructions for a single viral protein, which your own cells briefly produce to train an immune response. For the wider picture of how that immune training works, see how do vaccines work.
DNA as evidence of evolution
Comparing DNA sequences between species gives a quantitative measure of relatedness. Humans and chimpanzees share approximately 98.7% of their coding DNA; the figure shifts depending on what is measured, but the ordering of relationships from such comparisons matches the pattern built independently from fossils and anatomy.
Shared “broken” genes — the same inactive gene disrupted in the same way in both humans and chimps — are particularly strong evidence, since the odds of two species independently breaking a gene in an identical position are negligible.
DNA also records events that leave no fossils: population bottlenecks, migrations, and interbreeding with archaic humans such as Neanderthals, whose sequences survive in modern non-African populations.
Frequently asked questions about DNA
Does DNA determine everything about you?
No. DNA provides a predisposition, not a verdict. Gene expression is shaped by epigenetics, environment, diet, chance and development. Identical twins with identical genomes develop differences in traits, disease risk and even appearance over a lifetime — the clearest demonstration that sequence is necessary but not sufficient.
How much DNA do we share with a banana?
The often-quoted figure is around 60%, but it needs heavy qualification. It refers to specific conserved genes that perform basic cellular functions shared across all eukaryotic life, not overall similarity. Any two eukaryotes share the core machinery of cell division and metabolism; that reflects common ancestry, not close relationship.
Can you run out of DNA?
No. DNA is replicated before every cell division, so each daughter cell receives a full copy. Telomeres — repetitive sequences at chromosome ends — do shorten with each division in most cells, acting as a rough division counter. Once they become critically short, the cell enters senescence. Telomerase lengthens them again in stem cells, gametes, and unfortunately in most cancer cells.
Is genetic testing worth it?
It depends entirely on what you want from it. Testing for a known pathogenic variant in a family with a clear hereditary condition has high value. Broad consumer ancestry reports are entertainment with a scientific basis. Predictive results for common diseases carry modest risk differences and are easily misread without context. A genetic counsellor exists precisely to handle the interpretation, and for serious results they are the right first call rather than the last.
Why is junk DNA no longer the right term?
Because function has been found in a substantial share of what was dismissed, and because absence of evidence for function is not evidence of absence — proving something has no function is far harder than proving it does. Some of the genome almost certainly is non-functional, but the label implied a confidence the data never supported.
This article explains general biology concepts and is not medical or genetic counselling advice.















